Site Reliability Engineer, Post Training

Thinking Machines Lab · San Francisco, California; New York · $300,000 - $350,000 USD · Posted 2026-08-31

Apply on Thinking Machines Lab's site

About Thinking Machines ===========================

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role ==============

We're hiring a Site Reliability Engineer (SRE) to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.

You'll work side by side with research teams during active model runs — debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.

What You’ll Do ==============

Skills & Qualifications =======================

Minimum Qualifications ----------------------

Preferred Qualifications ------------------------

Logistics =========

More jobs at Thinking Machines Lab

Related searches

Updated 2026-10-10.