混元训练 Infra 工程师-Dataloader/Checkpoint 方向-(北京/深圳/上海/杭州)
Core
Designing and optimizing distributed data loading frameworks and checkpoint management systems for large-scale AI model training.
Role type
Senior IC distributed systems engineer (AI infrastructure)
Builds
High-throughput data loading pipelines, scalable checkpoint storage solutions, and optimized training infrastructure.
Domain
Artificial Intelligence / Distributed Systems / High-Performance Computing
Deliverable
production ML models
Required skills
Python, C++, Linux kernel internals, IO models, PyTorch, distributed training principles, object storage, distributed file systems, GPU/CPU optimization, performance bottleneck analysis
Preferred skills
Experience with dynamic sampling, incremental updates, compression strategies, and cold/hot data tiering
Responsibilities
Develop multi-source data loading frameworks to optimize preprocessing pipelines and address IO bottlenecks; Design high-throughput storage and loading schemes for checkpoints with version management and backup recovery; Monitor throughput, latency, and memory metrics to ensure training continuity under extreme conditions; Collaborate with cross-functional teams to align on business requirements and establish technical best practices.