CareerPlanSign in

混元训练 Infra 工程师-Dataloader/Checkpoint 方向-(北京/深圳/上海/杭州)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Designing and optimizing distributed data loading frameworks and checkpoint management systems for large-scale AI model training.

Role type

Senior IC distributed systems engineer (AI infrastructure)

Builds

High-throughput data loading pipelines, scalable checkpoint storage solutions, and optimized training infrastructure.

Domain

Artificial Intelligence / Distributed Systems / High-Performance Computing

Deliverable

production ML models

Required skills

Python, C++, Linux kernel internals, IO models, PyTorch, distributed training principles, object storage, distributed file systems, GPU/CPU optimization, performance bottleneck analysis

Preferred skills

Experience with dynamic sampling, incremental updates, compression strategies, and cold/hot data tiering

Responsibilities

Develop multi-source data loading frameworks to optimize preprocessing pipelines and address IO bottlenecks; Design high-throughput storage and loading schemes for checkpoints with version management and backup recovery; Monitor throughput, latency, and memory metrics to ensure training continuity under extreme conditions; Collaborate with cross-functional teams to align on business requirements and establish technical best practices.

Sourced via tencent · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.