Site Reliability Engineering Lead

Graphcore · Austin, Texas · Posted 2026-09-29

Apply on Graphcore's site

About Graphcore

How often do you get the chance to build a technology that transforms the future of humanity?

Graphcore products have set the standard in made-for-AI compute hardware and software, gaining global attention and industry acclaim. Now we are developing the next generation of artificial intelligence compute with systems that will allow AI researchers to develop more advanced models, help scientists unlock exciting new discoveries, and power companies around the world as they put AI at the heart of their business.

Graphcore recently joined SoftBank Group – bringing large and ongoing investment from one of the world’s leading backers of innovative AI companies.

Job Summary

We are seeking an experienced Site Reliability Engineering leader to build and lead a new SRE organization responsible for the production operation of a rapidly scaling AI supercomputing platform. The environment combines highly customized compute, high-performance networking, storage and supporting infrastructure, and will grow through multiple phases of deployment.

This is a rare opportunity to establish the reliability function for a new platform from the ground up. The platform and its operational model are being developed in parallel and will ultimately support a 24x7x365 production service with stringent availability requirements.

You will take the SRE organization from initial formation through production launch, stabilization and scale. This includes hiring and developing the team, defining the operating model, establishing production readiness and incident-management practices, and ensuring reliability and operability are engineered into the platform from the outset.

SRE is responsible for the operational capability required to run the platform reliably in production, while partnering with engineering teams that remain accountable for the reliability and operability of the systems they build.

This is not a purely managerial position. During the development and early production phases, the SRE Manager will be expected to work directly with engineering teams, develop a deep understanding of the platform, and participate in troubleshooting and incident response.

Over time, success will increasingly mean building the people, processes, automation, tooling, and operational discipline that allow the organization to operate effectively without depending on you for day-to-day escalation.

Responsibilities and Duties

Required Skills and Experience

Desired but Not Required

Candidates are not expected to have experience in all the areas below. Experience in several would be particularly valuable:

What Success Looks Like

In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.

About Graphcore

Graphcore develops a microprocessor designed for AI and machine learning applications.

More jobs at Graphcore

Related searches

Updated 2026-10-10.