HirePortal

Site Reliability Engineer / DevOps — Retail Engineering

  • Apple
  • Shanghai, China
  • CNY 500,000 – CNY 800,000

Summary

Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.


Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here.
Retail and Marcom Engineering, an IS&T team, builds and operates the systems and experiences that connect Apple's products with its customers. The team owns the technology behind both Apple's online and physical stores and drives the interactive marketing experiences and tools that keep creative operations moving.
Together, those functions deliver the technology behind every product story Apple tells and every purchase a customer makes.

Description

You will drive the reliability, deployment, and scalability of compute platforms across on-premises and hybrid cloud environments. Collaborating closely with cross-functional technical and business partners, you will build Infrastructure as Code, optimize container orchestration, and streamline CI/CD delivery pipelines. You will champion automation and operational excellence, ensuring high availability, robust security standards, and proactive observability across large-scale distributed systems.

Preferred Qualifications

In-depth understanding of SRE principles, including error budgeting, SLO/SLI/SLA definition, and advanced observability practices (Prometheus, Splunk, Grafana, OpenTelemetry).
Advanced programming skills in Java, Python, or Go, with hands-on experience across relational, NoSQL, or OLAP databases and event-driven streaming architectures (Kafka, RabbitMQ).
Track record of managing production on-call rotations, critical incident triage, root cause analysis (RCA), and post-incident reviews (PIR).
Solid knowledge of enterprise security standards, cryptography, authentication protocols (OAuth, SAML, SSO), and compliance governance.

Minimum Qualifications

Bachelor’s degree in Computer Science or equivalent field with 7+ years of experience, or Master’s degree with 5+ years of experience.
7+ years of experience in Site Reliability Engineering with a strong focus on building, scaling, and operating large-scale distributed platform services, and Java applications.
Strong technical grasp of Open Source technologies designed for large-scale data processing.
Proven expertise in designing, analyzing, and troubleshooting complex distributed systems.
Proficiency in at least one modern programming or scripting language (Python, Java, Go, Bash, Ansible, or similar).
Practical experience designing and deploying end-to-end observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry, ELK, etc.).
Demonstrated troubleshooting and problem-solving skills across production software and infrastructure environments.

Skills

  • Kubernetes
  • Infrastructure as Code
  • CI/CD
  • Cloud Computing
  • Python
  • Go
  • Monitoring and observability

Related jobs

AppleApply for this job