data engineer for healthcare data
Veröffentlicht am
- Arbeitsort
- 10115 Berlin, Berlin, Deutschland
Stellenbeschreibung
Описание: Statista is a global business data platform that provides reliable, easy-to-use data, analytics products, and services to support fact-based decision-making worldwide. Its Healthcare Platform transforms international hospital data into structured, queryable healthcare data assets.
Задачи:
Build, optimize, and operate reliable ELT pipelines using Python, SQL, and Prefect/Airflow; Ingest data from heterogeneous international sources, APIs, databases, and lakehouse storage including S3 and Apache Iceberg; Drive entity resolution and master data management for international hospital entities; Map raw source data to canonical structures and maintain standardized vocabularies such as ICD/OPS and specialty taxonomies; Establish data contracts and schema management with Pydantic and dbt contracts; Ensure dataset reproducibility, data lineage tracking, and automated validation across platform pipelines; Optimize data storage, query execution, and compute costs across AWS and Snowflake; Implement automated testing and deployment workflows for data pipelines using GitHub Actions and Terraform; Partner with Analytics Engineers, Data Scientists, and domain experts to deliver
documented, research-grade, production-ready datasets.
Требования:
Advanced Python and analytical SQL for complex data ingestion across diverse file formats, REST APIs, databases, and cloud lakes; Hands-on experience with Prefect, Airflow, or Dagster; Practical experience with entity resolution or record linkage frameworks; Experience with schema management and data contracts using Pydantic, dbt contracts, or JSON Schema; Deep hands-on experience in an AWS production environment, including S3 and ECS/EC2; Experience with cloud data warehouses, including Snowflake; Proven experience automating data pipeline deployments, integration tests, and validation workflows via GitHub Actions; 3+ Years of experience in data engineering building production pipelines and data platforms; Experience taking a core platform or product through build, launch, and iteration within one organization; Bachelor's or Master's degree in Computer Science, Data Science, Software
Engineering, or a related quantitative field; Strong analytical and systems mindset; Ability to transform messy, heterogeneous international data into clean, well-governed, highly structured data assets; Fluent English; Highly structured, curious, detail-oriented, and collaborative working style; Nice to have: Knowledge of ontology or semantic frameworks, international medical classifications or vocabularies, metadata registries and lineage catalogs, Terraform, healthcare domain experience, German.
Условия:
Work from abroad up to 30 calendar days a year; Hybrid work and flex-time; International team and social events; Subsidized urban mobility and access to fitness and wellness options; Free access to Langdock; Career and training opportunities; Attractive locations and modern offices; Mental health support with OpenUp; Some benefits apply only to the German entity and to Junior-level roles or above.