Senior AI Platform Engineer (Python, AWS, Data Pipelines)

Jetzt bewerben bei Aether Biomedical Sichere Bewerbung über StudySmarter
Aus der Stellenanzeige

Die ganze Ausschreibung von Aether Biomedical

Darum lohnt es sich

Location: remote with the first day onboarding in Warsaw and occasional visits once per quarterRate: 170 pln/h on b2b Project Overview We are building CaaS (Content as a Service) a platform that transforms publisher content (PDF textbooks and Excel manifests) into structured, enriched, AI-ready data.The platform processes content once and exposes it through a unified service layer used by multiple downstream applications.

Key use cases • RAG-based Teacher Assistant • Editorial tooling • Future AI-powered student-facing products The goal of this role is to design, build, and maintain a scalable data and AI platform that ingests, processes, enriches, and serves content reliably across multiple environments and consumers.

Responsibilities Data Engineering & PipelinesBuild and maintain multi-stage data ingestion pipelinesDesign and implement idempotent, restartable batch processing workflowsUse S3 as core storage layer for raw and processed dataImplement pipeline stages including:Content ingestion and book identity assignmentPDF-to-markdown conversion (AI OCR)Table of contents and structure extractionHierarchical chunkingEmbedding generationAI / LLM ProcessingUse LLMs and OCR models to extract structured data from PDFsDesign prompts and context strategies for consistent outputsGenerate structured metadata and enrich content for downstream use casesData Storage & ConsistencyMaintain PostgreSQL (Aurora) as system of recordDesign and maintain SQL schemas and versioned migrationsEnsure data consistency across:S3PostgreSQL (Aurora)Vector database (Weaviate)Implement reconciliation logic across distributed systemsRetrieval & Vector SearchWork with Weaviate for vector search and semantic retrievalSupport RAG-based applicationsDesign data organization strategies (by subject, country, and client)APIs & IntegrationBuild REST APIs using FastAPIExpose content as a service for multiple downstream applicationsIntegrate with internal and external systemsEngineering PracticesWrite strongly typed Python code (mypy)Follow CI/CD processes with automated checks (ruff, pytest)Work across dev / staging / production environmentsDebug distributed data inconsistencies Key Requirements • Must-have • Strong Python development experience (production systems) • AWS experience (S3, Glue, Aurora) • Experience with data pipelines (ETL / batch processing) • Strong SQL and PostgreSQL experience • Experience with schema design and migrations • Nice-to-have • Experience with LLMs in production (OCR, content processing, enrichment) • Prompt engineering / context engineering • Experience with vector databases (Weaviate, Pinecone, Qdrant, pgvector) • Knowledge of embeddings, semantic search, and RAG • Experience with FastAPI • Experience with Airflow / MWAA • Experience building data platforms serving multiple consumers

Bereit?

Bewerbung wird direkt an Aether Biomedical übergeben — kein Konto nötig.

Jetzt bewerben
Weitere Stellen bei diesem Arbeitgeber

Aether Biomedical hat 10 weitere offene Stellen:

Alle 11 Stellen bei Aether Biomedical ansehen →
Ähnliche Stellen

Wenn dir dieser Job gefällt, schau dir auch an:

Weiter stöbern:

Kostenfrei starten