Site Reliability Engineer, Production

Thinking Machines Lab · San Francisco, California; New York · $300,000 - $350,000 USD · Posted 2026-08-31

Apply on Thinking Machines Lab's site

About Thinking Machines ===========================

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role ==============

We're looking for a Site Reliability Engineer (SRE) to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research teams to make every layer of the system more robust and resilient.

About Tinker ============

Tinker is our fine-tuning API that empowers researchers and developers to customize frontier AI to their needs — opening access to capabilities that have previously been concentrated in a handful of labs. We manage the infrastructure while allowing Tinkerers full flexibility in training open weights models with their own data, algorithms, and for their own needs. Tinker is rapidly adding new customers, features, and novel use-cases. We’re hiring to grow the platform alongside the Tinker community.

What You’ll Do ==============

Skills and Qualifications =========================

Minimum qualifications ----------------------

Preferred qualifications ------------------------

Logistics =========

More jobs at Thinking Machines Lab

Related searches

Updated 2026-10-10.