Search
Data Platform Engineer – Data Operations (all genders)

Data Platform Engineer – Data Operations (all genders)

locationMunich, Germany
PublishedPublished: Published today
Full time
About Us

STARK is a new kind of defence technology company revolutionizing the way autonomous systems are deployed across multiple domains. We design, develop and manufacture high-performance unmanned systems that are software-defined, mass-scalable, and cost-effective. This provides our operators with a decisive edge in highly contested environments.

We're focused on delivering deployable, high-performance systems — not future promises. In a time of rising threats, STARK is bolstering the technological edge of NATO Allies and their Partners to deter aggression and defend Europe — today.

About the team

The Data Operations team owns the entire data lifecycle behind STARK's AI stack: collection, acquisition, generation, curation, and management. We run our own data-collection campaigns across Europe, evaluate new sensors and platforms, and build the internal data platform that turns raw recordings into ready-to-use datasets. Everything we produce feeds directly into the perception and autonomy systems deployed on STARK's platforms — a real data advantage is built, not bought. The team is scaling up right now: real scope, direct impact, no legacy.

Your mission Data is the fuel of STARK’s AI stack — you build the engine that makes it usable. You own the software backbone of our data platform: the metadata systems, ETL pipelines, data contracts, catalogs, databases, and internal tools that let engineers find, understand, validate, and reuse terabytes of multi-sensor field data in minutes, not days. You treat data context as a product: structured, searchable, version-aware, documented, and traceable from raw recording to processed asset, annotation delivery, dataset, and downstream ML workflow. Today, much of this is manual, scattered, or implicit — your job is to automate it away, support labeling efforts with the right data tooling, and turn operational data into reliable systems. Responsibilities
  • Design, implement, and maintain our metadata database and data catalog (datasets, recordings, sensors, labels, lineage)

  • Build and operate ETL/ingest pipelines that bring field recordings, synthetic data, and external deliveries into our cloud storage (GCP)

  • Own the data management and labeling lifecycle end-to-end: coordinate and communicate with external labeling companies and data subcontractors, track deliveries, run QA reports, and build the operational workflows they work in

  • Develop internal enabling tools for the whole AI organization: dataset search and filtering, APIs/backend, dashboards, and self-service data access

  • Run data migrations and indexing jobs; keep the catalog consistent and fast as data volume grows

  • Handle admin support and user access management — and then automate these support tasks so they stop being manual work

  • Establish good engineering hygiene in a young codebase: tests, typing, docs, logging, CI/CD

  • Shape the long-term architecture and vision of the data platform together with the team

Qualifications
  • Strong Python

  • Solid SQL/PostgreSQL, including schema design

  • Experience with data modeling and metadata systems

  • Experience designing and operating ETL/data pipelines

  • Docker and CI/CD basics

  • Hands-on with object storage (GCS, S3, or similar)

  • Good software engineering hygiene: tests, docs, typing, logging

  • Organized and pragmatic: you can prioritize between a quick fix and a proper solution, and you know when each is right

  • Not allergic to support tasks — but technical enough to automate the support away

  • Comfortable coordinating with external vendors and non-technical stakeholders

Nice to have
  • Familiarity with ML datasets and labeling workflows (images, video, lidar; annotation formats like COCO)

  • Experience with synthetic data generation or GenAI-assisted data workflows (auto-labeling, data augmentation, foundation-model-based curation)

  • Experience with GCP services beyond storage (BigQuery, Cloud Run, IAM)

  • Experience with data versioning / dataset tooling (DVC, LakeFS, FiftyOne, or similar)

  • Experience in a startup environment — comfortable with ambiguity and changing priorities

  • Exposure to robotics data formats (ROS bags, MCAP, PX4 logs)

Image gallery

Video gallery

Consent to this service