Research Engineer - Benchmarks

Firecrawl ยท Toronto Hub; San Francisco HQ ยท $250,000 - $290,000 USD ยท Posted 2026-09-29

Apply on Firecrawl's site

Research Engineer - Benchmarks ==================================

Every week, someone asks which data provider is actually best: for company records, for finding the right person, for fresh job listings. Right now the answers come from the vendors themselves. You'll build the independent version. You'll run rigorous, automated, public benchmarks of the providers in Alexandria and beyond, and publish them weekly. The same results will feed straight back into Alexandria so it learns which provider to call for which job.

This isn't our internal evals role. You're measuring the market, in public, where every number will be challenged by the vendors it ranks. You'll own it end to end: the datasets, the ground truth, the scoring, the harness, the weekly release, and the loop into Alexandria's provider selection. No one hands you a methodology. You write it, defend it, and ship it every week.

Salary Range: $250,000โ€“$290,000 USD/year (SF) / $210,000โ€“$224,000 CAD/year (Toronto)

Equity Range: Competitive equity. Details shared during the process.

Location: San Francisco, CA (SF HQ) or Toronto, ON (Toronto Hub). Hybrid, onsite 3+ days a week.

Equity Range: Competitive equity. Details shared during the process.

Location: San Francisco, CA (SF HQ). On-site, five days a week.

Job Type: Full-Time

Experience: 4+ years in ML, research engineering, or data engineering, with evaluation or benchmark work you've shipped

Work Authorization: Must be authorized to work in the United States or Canada. We're not able to sponsor US visas right now. For Canada, we'll consider sponsorship on a case-by-case basis through our Toronto Hub.

About Firecrawl ===================

Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data. It's the boring-hard problem everyone building with LLMs eventually hits, solved.

In September 2026, we raised a $75M Series B led by Smash Capital, and we're spending it building the largest repository of knowledge in the world. We hit 8 figures in ARR in year one and more than doubled it in year two. We have 187k+ GitHub stars, putting us in the top 40 repositories of all time, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started.

We're a small team punching far above our weight. Everyone here owns a real piece of the product and company, end to end, and runs it themselves. No hiding behind process or headcount.

This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web. That library is called Alexandria, and it starts now.

What You'll Do ==================

What We're Looking For ==========================

What We're Not Looking For ==============================

A Note On Pace ==================

We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings, but this role probably isn't for you.

Benefits & Perks ====================

Available to all employees --------------------------

Available to US-based full-time employees ---------------------------------------------

Available to Canada-based full-time employees -------------------------------------------------

Available to SF-based employees -----------------------------------

Available to Toronto-based employees ----------------------------------------

Interview Process =====================

  1. Application Review: Send us your work. We want the benchmark, leaderboard, eval, or dataset you built, and how you knew it was right. We care about what you've shipped, not where you went to school.
  1. Intro Chat (~25 min): A quick conversation to get to know each other. We'll cover what you've been working on, what drew you to Firecrawl, and what you want next. Time for your questions too.
  1. Technical Chat (~45 min): A real problem from our world: design a benchmark that decides whether FullEnrich or DataLegion finds the right person's current title. We'll cover where the ground truth comes from and how you'd defend the result to the vendor who loses. Come ready to think out loud.
  1. Workflow Chat (~30 min): Show us how you actually work: your AI tools, your setup, and a recent thing you shipped faster than you would have a year ago.
  1. Founder Chat (~25 min): Culture, pace, ownership, and how you like to work. Time for your questions too.
  1. Paid Work Trial (~40 hours): Ship a small benchmark end to end on a real provider category, paid at a contractor rate. It's the truest signal for both sides. Remote-friendly, and we'll flex around your current commitments.
  1. Decision: We move fast after the trial.

If you want to be the person the whole market checks before picking a data provider, you should join us.

๐Ÿ‘‰ Apply now.

About Firecrawl

The API to search, scrape, and interact with the web at scale. ๐Ÿ”ฅ

More jobs at Firecrawl

Related searches

Updated 2026-10-11.