Senior AI Platform Engineer (Python, AWS, Data Pipelines)
Die ganze Ausschreibung von Aether Biomedical
Darum lohnt es sich
Location: remote with the first day onboarding in Warsaw and occasional visits once per quarterRate: 170 pln/h on b2b Project Overview We are building CaaS (Content as a Service) a platform that transforms publisher content (PDF textbooks and Excel manifests) into structured, enriched, AI-ready data.The platform processes content once and exposes it through a unified service layer used by multiple downstream applications.
Key use cases • RAG-based Teacher Assistant • Editorial tooling • Future AI-powered student-facing products The goal of this role is to design, build, and maintain a scalable data and AI platform that ingests, processes, enriches, and serves content reliably across multiple environments and consumers.
Responsibilities Data Engineering & PipelinesBuild and maintain multi-stage data ingestion pipelinesDesign and implement idempotent, restartable batch processing workflowsUse S3 as core storage layer for raw and processed dataImplement pipeline stages including:Content ingestion and book identity assignmentPDF-to-markdown conversion (AI OCR)Table of contents and structure extractionHierarchical chunkingEmbedding generationAI / LLM ProcessingUse LLMs and OCR models to extract structured data from PDFsDesign prompts and context strategies for consistent outputsGenerate structured metadata and enrich content for downstream use casesData Storage & ConsistencyMaintain PostgreSQL (Aurora) as system of recordDesign and maintain SQL schemas and versioned migrationsEnsure data consistency across:S3PostgreSQL (Aurora)Vector database (Weaviate)Implement reconciliation logic across distributed systemsRetrieval & Vector SearchWork with Weaviate for vector search and semantic retrievalSupport RAG-based applicationsDesign data organization strategies (by subject, country, and client)APIs & IntegrationBuild REST APIs using FastAPIExpose content as a service for multiple downstream applicationsIntegrate with internal and external systemsEngineering PracticesWrite strongly typed Python code (mypy)Follow CI/CD processes with automated checks (ruff, pytest)Work across dev / staging / production environmentsDebug distributed data inconsistencies Key Requirements • Must-have • Strong Python development experience (production systems) • AWS experience (S3, Glue, Aurora) • Experience with data pipelines (ETL / batch processing) • Strong SQL and PostgreSQL experience • Experience with schema design and migrations • Nice-to-have • Experience with LLMs in production (OCR, content processing, enrichment) • Prompt engineering / context engineering • Experience with vector databases (Weaviate, Pinecone, Qdrant, pgvector) • Knowledge of embeddings, semantic search, and RAG • Experience with FastAPI • Experience with Airflow / MWAA • Experience building data platforms serving multiple consumers
Bereit?
Bewerbung wird direkt an Aether Biomedical übergeben — kein Konto nötig.