Alibaba Group
Site Reliability Engineer
- LocationUnited States
- TypeFull-time
- Posted2026-09-14
- Valid through2026-09-28
- BudgetCNY 60000–90000 / month
- Long-termYes
About this task
We are seeking a technically skilled Site Reliability Engineer to help build and maintain a highly available, high-performance AI inference and model service platform. The role focuses on monitoring and alerting, incident response, customer issue resolution, troubleshooting, and operational automation.
Key Responsibilities:
1. Oversee the deployment, operation, maintenance, and continuous improvement of the website and platform, including initial construction and subsequent operational changes.
2. Monitor the platform and system applications, rapidly diagnose and resolve network-, service-, and hardware-level failures, and help meet SLA targets.
3. Design and optimize monitoring metrics, log collection, and alerting strategies to improve system observability.
4. Participate in emergency responses to online incidents, conduct root cause analyses, and drive long-term solutions to prevent recurrence.
5. Investigate and resolve customer-reported API quality-of-service issues, including latency, performance, and optimization problems. Collaborate with development teams to identify issues in application clusters, edge networks, or infrastructure.
6. Develop Python or Go tools and scripts to automate deployment, scaling, fault recovery, and other operational workflows.
7. Build automated diagnostic toolchains to accelerate issue resolution and improve customer satisfaction.
Requirements:
1. At least three years of experience in SRE, DevOps, or backend development, with expertise in distributed-system operations.
2. Experience with cloud computing, AI infrastructure, or Alibaba Cloud is preferred.