SRE Software Engineer
Software Engineering
Austin, TX, USA · Ontario, Canada
Posted on Aug 29, 2026
People at Apple don't just build products, they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it. The Apple Services Engineering (ASE) team builds and provides systems and infrastructure that power Apple's services (such as iCloud, Apple Music, Apple Intelligence, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. Our services have to scale globally, stay highly available, and "just work." If you love designing, engineering, and running systems and infrastructure that will help millions of customers, then this is the place for you!
The ASE Compute team is looking for a Site Reliability Engineer to deploy and manage a large Kubernetes platform that Apple's services run on, partnering with engineering teams across the company to solve complex problems using both open-source and in-house tooling. You will contribute to the development of our controllers and namespace management infrastructure, working alongside senior engineers to strengthen the reliability of our Kubernetes services. You will learn to write well-tested code, participate in design reviews, and gradually take ownership of features. You'll have the opportunity to engage with the upstream community, gain hands-on experience with production-scale systems, and build the technical foundation to support service teams across Apple. The role also offers room to build AI-assisted tooling that accelerates triage, operational workflows, and infrastructure automation for the whole team.
- Deploy, configure, and maintain large-scale, multi-tenant Kubernetes environments
- Write and maintain operational tooling to improve reliability and reduce manual intervention
- Implement and maintain reliability standards for the platform: SLOs, error budgets, alerting philosophy, upgrade and rollout strategy, and the run-books that follow from them.
- Contribute to CI/CD pipelines, revision control workflows, and configuration management practices
- Take on-call, troubleshoot production issues, and follow up on post-incident action items
- Help enforce security best practices, OS hardening, and compliance standards across the fleet
- Hands-on experience in Linux systems administration and containerization with enterprise distributions such as RHEL, Oracle Linux, or CentOS
- Proficiency in Python or Go for scripting and tooling
- Solid understanding of Linux fundamentals: file systems, process management, user and group administration, and package managementWorking knowledge of networking concepts including TCP/IP, DNS, DHCP, and basic firewall configuration
- Experience with version control systems such as Git and configuration management (Puppet, Ansible, or equivalent)
- Strong written and verbal communication skills
- Site Reliability Engineering, DevOps, or Infrastructure focused experience
- Experience with third-party cloud platforms (AWS, GCP, or Azure)
- Experience with containerization and orchestration technologies such as Docker or Kubernetes
- Familiarity with bare-metal provisioning and lifecycle management at datacenter scale
- Understanding of cloud-native observability (Prometheus, Thanos, Splunk, or similar)
- Familiarity with CI/CD pipelines and DevOps practices
- Knowledge of OS security hardening, encryption, and regulatory compliance frameworks