M

Founding Applied AI Researcher - Model Evaluation & Data Strategy

Morpheus Talent Solutions San Francisco Bay Area
Visa Sponsorship
Apply
AI Summary

Join an early-stage AI research company as a founding Applied AI Researcher to design evaluations, form and test hypotheses about model failure, and define datasets, rubrics, reward signals, and quality controls. Publish research that positions the company as a research partner to the field. A prior publication record is essential.

Key Highlights
Design evaluations for generative, reasoning, tool-use, and agentic systems across modalities
Own real experimental design: hypothesize where a model breaks, build the eval to test it, and quantify what's actually failing
Build quality systems - calibration, blind review, adjudication - that hold up in non-deterministic domains
Key Responsibilities
Design evaluations for generative, reasoning, tool-use, and agentic systems across modalities
Own real experimental design: hypothesize where a model breaks, build the eval to test it, and quantify what's actually failing
Build quality systems - calibration, blind review, adjudication - that hold up in non-deterministic domains
Technical Skills Required
Python Model APIs Structured datasets
Benefits & Perks
$200K-$350K base + equity
Visa sponsorship available
Nice to Have
Experience at an AI lab, foundation-model company, or post-training team
Expert-data or human-eval program design
Multimodal/coding/agentic eval work

Job Description


Applied AI Researcher - Model Evaluation & Data Strategy


San Francisco (in-person preferred; open to remote across US, UK, Australia, and Europe) · Retained search - confidential client


The engagement

Morpheus has been exclusively retained to lead the search for a founding Applied AI Researcher on behalf of an early-stage, profitable AI research company.


About the client


An early-stage AI research company that works with leading AI labs to find where frontier models fail and build the expert human data that fixes them. They run a vetted network of 5,000+ top-1% specialists across finance, medicine, law, engineering, music, and other domains - competing on the quality of expert judgment, not the scale of cheap labeling.


Backed by a top pre-seed fund and angel investors who are founders and senior researchers at leading frontier AI labs. Already profitable.


The role


A founding, research-first seat - a genuine thought partner on evaluation and data strategy, not someone who coordinates other people's research, and not client-facing or delivery. You'll design evaluations, form and test hypotheses about model failure, and define the datasets, rubrics, reward signals, and quality controls that move performance - then work with engineers to turn them into scalable programs. Publishing is core to this role, not a perk - you'll be expected to author and present research that positions the company as a research partner to the field, so a prior publication record is essential.


What you'll do


  • Design evaluations for generative, reasoning, tool-use, and agentic systems across modalities - text, audio, vision - where technique and domain expertise differ meaningfully by modality.
  • Own real experimental design: hypothesize where a model breaks, build the eval to test it, and quantify what's actually failing.
  • Build RL environments and reward signals in close partnership with engineers.
  • Recommend SFT data, preference data, expert demonstrations, critiques, and eval sets.
  • Build quality systems - calibration, blind review, adjudication - that hold up in non-deterministic domains.
  • Run pilots that prove whether an intervention moves performance, and publish work that positions the company as a research partner to the field.


What the client is looking for


  • A track record of published research - you've authored papers at venues like NeurIPS, ICML, ICLR, ACL, or EMNLP (or comparable). This is a research seat with a mandate to publish, so a demonstrated publication record is essential.
  • Genuine experimental-design experience - you've designed studies and evals, not just executed someone else's rubric.
  • Comfort operating in non-deterministic domains and with novel, fast-moving research frameworks.
  • Agentic evaluation proficiency; RL environment experience, ideally built alongside engineers.
  • Cross-modal understanding - awareness that audio, text, and vision each demand different techniques.
  • Strong Python, model APIs, and structured datasets; solid grounding in benchmark design, human eval, rubric development, and statistical analysis.
  • Familiarity with SFT, preference optimization, RLHF/RLAIF, reward modeling, synthetic data, or LLM-as-a-judge.
  • Strong technical writing and the ability to drive ambiguous research independently.


Nice to have


Experience at an AI lab, foundation-model company, or post-training team; expert-data or human-eval program design; multimodal/coding/agentic eval work; public benchmarks or eval frameworks.


The reality - worth knowing up front


This is an early, high-momentum team that currently works a six-day week: Saturdays are fully remote and self-directed, no set hours - most people use them as a heads-down research day. Compensation is $200K-$350K base + equity; visa sponsorship available.


To apply


Apply here or message me directly. I represent this search exclusively and will share the company name, team, and full details confidentially with candidates who are a strong fit.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

coffeespace

San Francisco Bay Area
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Nxt Level

San Francisco Bay Area

Senior Talent Acquisition Sourcer (AI/ML & Engineering Leadership)

Programming
2d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

anthropic

San Francisco Bay Area

Subscribe our newsletter

New Things Will Always Update Regularly