SRE高级/工程师/架构师/负责人
Core
Design and maintain highly scalable, high-availability distributed systems for big data and computing core services, focusing on reliability, cost efficiency, and stability.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Automated operational solutions, monitoring platforms, and high-availability architecture upgrades for large-scale clusters.
Domain
Big Data & Computing Systems
Required skills
Linux internals, network IO, storage, Python/Go/Java/Shell, distributed systems, big data technologies, system design, algorithmic thinking
Preferred skills
Nginx, Kubernetes, Docker, OpenStack, Hadoop, Spark, Flink, performance bottleneck analysis
Responsibilities
Build automated solutions for large systems, monitor system availability and performance metrics, optimize service reliability and cost, design automation platforms for rapid iteration, troubleshoot business issues and optimize service governance.
Seniority
Senior, hands-on IC