← Zur Jobsuche
Senior Machine Learning Data Processing Engineer
Veröffentlicht am
- Arbeitsort
- 10115 Berlin, Berlin, Deutschland
Stellenbeschreibung
Our client is a non-profit organisation advancing research and technical solutions for safe-by-design AI systems.
They are seeking a Senior ML Data Processing Developer to build, scale and optimise large-scale machine learning data pipelines. You will work across data engineering, data curation and machine learning, transforming web-scale data into high-quality datasets for next-generation AI models.
This role goes beyond traditional data engineering, with a strong focus on data quality, algorithmic filtering, model-based evaluation and novel data-processing techniques.
Key Responsibilities
- Design, build and scale end-to-end pipelines for web-scale ML training data.
- Develop data-processing systems for deduplication, quality scoring, filtering, toxicity removal, PII scrubbing, metadata extraction and custom transformations.
- Build data-quality tooling including heuristic filters, ML classifiers, LLM-based evaluators and human-in-the-loop workflows.
- Implement robust data versioning, lineage, provenance, monitoring and quality controls.
- Develop leakage-detection mechanisms to prevent evaluation contamination.
- Work with Research and Engineering teams to define evolving data requirements and identify large-scale datasets.
- Conduct dataset coverage analysis and develop targeted data-acquisition strategies.
- Collaborate with Legal and Governance teams on data licensing, compliance and privacy requirements.
- Build tools that enable researchers to efficiently explore and understand datasets.
- Improve the scalability, reliability, quality and cost-efficiency of data-processing workflows.
Requirements
- Degree in Computer Science, Software Engineering or a related field.
- 5+ years' experience in data processing, ML engineering or NLP.
- Experience working with massive unstructured text datasets at trillion-token scale.
- Strong experience with distributed processing frameworks such as Spark, Ray or Flink.
- Experience with PII scrubbing, content-safety filtering and evaluation-contamination prevention.
- Strong Python skills and experience writing production-grade data-processing systems.
- Experience with pipeline orchestration tools such as Airflow, Prefect or Dagster.
- Ability to collaborate across Research, Engineering and Legal/Governance teams.
Nice to Have
- Experience developing ML models for data-quality tasks, including classifiers and LLM-based evaluators.
- Familiarity with LLM inference optimisation tools such as vLLM or SGLang.
- Experience with Docker, Kubernetes and infrastructure-as-code.
- Familiarity with ML experiment tracking tools such as Weights & Biases.
- Experience with data licensing or web-scale data acquisition.
- Contributions to open-source data processing or NLP projects.