AI管线数据专家/架构师(深圳/北京)
Core
Design and implement high-efficiency data processing platforms for pre-training and post-training pipelines, optimizing compute and storage for large model training.
Role type
Senior IC data pipeline architect (LLM training infrastructure)
Builds
End-to-end data pipelines, compute/storage platforms, and dataset management systems for large language model training
Domain
Artificial Intelligence / Large Language Models / Data Infrastructure
Deliverable
production ML models
Required skills
Python, C++, Java, distributed batch processing frameworks, data pipeline architecture, storage optimization, data version control, observability
Preferred skills
None stated
Technologies
Spark, Ray, Flink, Dask, DVC, MLflow, Delta Lake, Data Catalog
Responsibilities
Design and implement data processing platforms for pre- and post-training pipelines; optimize platform-level compute efficiency through operator orchestration; build and optimize storage solutions for massive datasets including tiering, caching, and lifecycle management; write technical documentation and performance evaluation reports
Seniority
Senior, hands-on IC