Senior Data Engineer
Core
Design, build, and operate data infrastructure for a clinical trial design platform that extracts insights from protocols and regulatory documents to accelerate trial design decisions.
Role type
Senior Data Engineer (Clinical AI/ML Infrastructure)
Builds
Ingestion pipelines, relational and graph data stores, vector stores, and data-serving APIs for LangGraph-based agentic AI workflows.
Domain
Clinical trials, Biomedical data, AI/ML infrastructure
Deliverable
production ML models
Required skills
Python, PDF parsing, PostgreSQL/Amazon Aurora, Graph databases (Neptune/Neo4j), AWS services, FastAPI, RAG pipeline design, Workflow orchestration (Airflow/Prefect), Infrastructure as Code (Terraform/CDK)
Preferred skills
Biomedical knowledge graphs, Clinical data standards (MeSH/MedDRA), Apache Spark, dbt
Technologies
Python, PyMuPDF, pdfplumber, unstructured.io, Amazon Aurora, GraphDB, Neptune, Neo4j, AWS (S3, Lambda, Step Functions, SQS/SNS, IAM), OpenSearch, pgvector, Pinecone, FastAPI, LangChain, LangGraph, Airflow, Prefect, Terraform, AWS CDK, Docker, Git
Responsibilities
Build ingestion pipelines for clinical trial protocols and regulatory documents; Design and implement data models in relational and graph databases; Develop embedding and vectorization pipelines for RAG-based retrieval; Build and maintain ETL/ELT workflows; Implement data quality validation for clinical data; Build data serving APIs; Set up data lineage tracking and audit trails.
Seniority
Senior, hands-on IC