系统分析工程师-AI业务
Core
Analyze performance bottlenecks in ByteDance's large-scale AI clusters (LLMs, search, voice/vision) and design/optimize AI infrastructure architecture including GPU selection, servers, interconnects, and super-nodes.
Role type
Senior IC AI Systems Performance Engineer
Builds
AI infrastructure (servers, interconnects, super-nodes), performance analysis tools, and platforms
Domain
AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
Linux kernel, containerization, C/C++, Python, GPU architecture, high-speed interconnects, collective communication, LLM model structures, recommendation system models, PD separation, KV Cache, model parallelism, AI inference engines, training frameworks, operator fusion, communication library acceleration, compilation optimization, performance profiling tools
Preferred skills
GPU/CPU performance tuning, AI cluster optimization
Responsibilities
Analyze performance bottlenecks in large-scale AI clusters; Build AI benchmarks and abstract models for cluster simulation; Develop performance analysis tools and platforms; Support end-to-end AI server product design from requirements to optimization
Seniority
Senior, hands-on IC