ITKarrierenDein Stack. Deine Möglichkeiten.Merkliste 0
Menü
← Zur Jobsuche

Senior Machine Learning Data Processing Engineer

Veröffentlicht am

Arbeitsort
10115 Berlin, Berlin, Deutschland

Stellenbeschreibung

Our client is a non-profit organisation advancing research and technical solutions for safe-by-design AI systems.

They are seeking a Senior ML Data Processing Developer to build, scale and optimise large-scale machine learning data pipelines. You will work across data engineering, data curation and machine learning, transforming web-scale data into high-quality datasets for next-generation AI models.

This role goes beyond traditional data engineering, with a strong focus on data quality, algorithmic filtering, model-based evaluation and novel data-processing techniques.

Key Responsibilities

  • Design, build and scale end-to-end pipelines for web-scale ML training data.
  • Develop data-processing systems for deduplication, quality scoring, filtering, toxicity removal, PII scrubbing, metadata extraction and custom transformations.
  • Build data-quality tooling including heuristic filters, ML classifiers, LLM-based evaluators and human-in-the-loop workflows.
  • Implement robust data versioning, lineage, provenance, monitoring and quality controls.
  • Develop leakage-detection mechanisms to prevent evaluation contamination.
  • Work with Research and Engineering teams to define evolving data requirements and identify large-scale datasets.
  • Conduct dataset coverage analysis and develop targeted data-acquisition strategies.
  • Collaborate with Legal and Governance teams on data licensing, compliance and privacy requirements.
  • Build tools that enable researchers to efficiently explore and understand datasets.
  • Improve the scalability, reliability, quality and cost-efficiency of data-processing workflows.

Requirements

  • Degree in Computer Science, Software Engineering or a related field.
  • 5+ years' experience in data processing, ML engineering or NLP.
  • Experience working with massive unstructured text datasets at trillion-token scale.
  • Strong experience with distributed processing frameworks such as Spark, Ray or Flink.
  • Experience with PII scrubbing, content-safety filtering and evaluation-contamination prevention.
  • Strong Python skills and experience writing production-grade data-processing systems.
  • Experience with pipeline orchestration tools such as Airflow, Prefect or Dagster.
  • Ability to collaborate across Research, Engineering and Legal/Governance teams.

Nice to Have

  • Experience developing ML models for data-quality tasks, including classifiers and LLM-based evaluators.
  • Familiarity with LLM inference optimisation tools such as vLLM or SGLang.
  • Experience with Docker, Kubernetes and infrastructure-as-code.
  • Familiarity with ML experiment tracking tools such as Weights & Biases.
  • Experience with data licensing or web-scale data acquisition.
  • Contributions to open-source data processing or NLP projects.