Summary
What you’ll impact
The Data Platform Engineer at the organization will own the end-to-end autonomy data pipeline, handling flight data collection, storage, cataloging, and query infrastructure. This greenfield role involves building scalable cloud storage, data processing, and tooling to make large-scale multi-sensor logs searchable, replayable, and usable for testing and machine-learning workloads.
Responsibilities
What you'll do
- Build the path that brings flight data back from the field, including triggered capture on the vehicle, prioritized upload so the highest-value flights return first, and resumable transfer with integrity verification.
- Turn raw logs into a usable corpus: decode, time-align multi-sensor and video streams, validate, and quarantine malformed data before it reaches downstream users.
- Design the catalog and tag model that index the corpus, and stand up the cloud storage and database that hold it
- Build the query layer so an engineer can retrieve every flight matching a condition, for example loss of target lock at terminal stage under high glare within the last 90 days, and get playable video back in seconds.
- Serve logs to the evaluation harness with stable ordering, exact time alignment, and reproducible results across runs, so a regression job can run over thousands of flights at once.
- Build versioned, immutable datasets from catalog queries, with lineage recorded so any model training set can be rebuilt exactly months later.
Requirements
What you’ll bring
- 5+ years building production data or backend infrastructure, including at least one system you owned end to end from initial design through ongoing operation
- Direct experience with large-scale log or sensor data: multi-terabyte and growing, with video and multiple synchronized sensor streams (rosbag, MCAP, HDF5, Parquet, or equivalent formats), rather than row-oriented business data
- Designed and owned a data schema, index, or catalog that other engineers queried daily, and lived with the consequences of that design, including at least one migration
- Strong Python, plus SQL and working ownership of a relational database (PostgreSQL or equivalent) used in production
- Practical experience with cloud object storage and compute (Azure, AWS, or GCP) and the ability to provision and operate it independently, without a dedicated platform or DevOps team
- Experience with distributed or parallel batch processing and job orchestration (Spark, Ray, Dask, Airflow, Dagster, or equivalent) across large volumes of recorded dat
- Working knowledge of time synchronization and alignment across sensor streams, and of deterministic, reproducible processing of recorded data
- A track record of building internal tooling that other engineers adopted, including at least one case where you changed the design based on how it was actually being used