CareerPlanSign in

Research Engineer / Performance Engineer, RL Distributed Systems

San Francisco, CA | New York City, NY | Seattle, WA💼 Full-time💰 $500,000–$500,000🗓 2026-09-29 → 2026-10-01

Core

Design, build, and operate distributed systems that run reinforcement learning at scale, ensuring correctness under failure and maximizing compute efficiency.

Role type

Senior IC Research Engineer (Distributed Systems)

Builds

Distributed infrastructure for RL training, sampling, and environment execution

Domain

AI/ML infrastructure, Distributed Systems

Deliverable

production ML models

Required skills

Python, Rust/C++/Go, distributed systems fundamentals, fault tolerance, autoscaling, observability, debugging complex failures

Preferred skills

ML training/inference infrastructure, Kubernetes, sandboxed execution, high-performance networking (RDMA), async Python (Trio/asyncio), RL workloads

Technologies

Python, Rust, C++, Go, Kubernetes, RDMA, Trio, asyncio

Responsibilities

Design schedulers for heterogeneous clusters, build failure detection and recovery systems, scale environment execution, design autoscaling policies, build diagnostics systems, trace data corruption bugs, design operational interfaces for automated tools

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 901,000+ jobs from 20+ sources.