Platform & Reliability Engineer
Core
Own the platform's reliability, uptime, and observability to ensure 99.9% availability for enterprise document ingestion and querying.
Role type
Senior IC Platform & Reliability Engineer
Builds
AWS infrastructure, observability pipelines, incident response processes, and deployment safety mechanisms
Domain
Enterprise document management, deep-tech, compliance (GDPR, ISO 27001/42001, HIPAA)
Deliverable
production ML models
Required skills
AWS infrastructure, SLO/SLI definition, incident response, capacity planning, compliance implementation, CI/CD pipeline management, post-mortem analysis
Preferred skills
Agent-based development (Claude Code), error budget management, high-volume data ingestion optimization
Technologies
AWS, Claude Code
Responsibilities
Define and publish error-budget policies, manage on-call rotation and alerting, lead incident response and post-mortems, plan capacity for high-volume ingestion, ensure compliance with security standards, optimize deployment pipelines for safety and speed
Seniority
Senior, hands-on IC