B

Reinforcement Learning Engineer

biostack platforms • San Francisco Bay Area
Visa Sponsorship Relocation
Apply
AI Summary

Build reinforcement learning environments and post-training systems for healthcare AI at BioStack. Design tasks, rewards, verifiers, and evaluation harnesses for clinical reasoning and longitudinal patient care. Requires strong judgment around data quality and experience with RL, language model post-training, and agent environments.

Key Highlights
Build RL environments for clinical reasoning, diagnostic decision-making, and chronic disease management
Design rewards and verifiers capturing correctness across longitudinal patient care tasks
Requires 2+ years experience with reinforcement learning, post-training, or agent systems
Key Responsibilities
Build healthcare-specific RL environments including tasks, action spaces/tool interfaces, reward functions, verifiers, and evaluation harnesses
Run post-training experiments on language models and agents using techniques such as SFT, RLVR, RLHF/RLAIF, and reward modeling
Turn clinical and biomedical datasets into training environments with measurable, verifiable outcomes
Design rewards and verifiers capturing correctness across clinical reasoning and longitudinal decision-making tasks
Train and evaluate multi-step agents operating across patient histories, clinical tools, and structured/unstructured medical data
Build scalable pipelines for rollouts, training, evaluation, experiment tracking, and dataset iteration
Analyze model failures and use them to improve environments, rewards, datasets, and subsequent training runs
Technical Skills Required
Reinforcement Learning Language Model Post-training Agent Environment Design
Benefits & Perks
Base salary: $200,000–$350,000 per year
Equity: 0.1%–1.0%
401(k): Company match
Medical, dental, and vision insurance
Complimentary lunch provided daily
Unlimited budget for AI tools and software
Relocation & joining bonus available
Visa sponsorship available

Job Description


About BioStack

BioStack is building the data layer for AI-native healthcare and drug discovery. We work with leading AI labs, human data companies, and frontier biotech teams to source, structure, and deliver high-value clinical and preclinical datasets for model training, evaluation, and deployment.

We sit at the intersection of healthcare, frontier AI, and data infrastructure. Our work spans medical institutions, clinics, imaging centers, and data partners globally, turning messy real-world clinical workflows into AI-ready products that matter.


The long-term vision is to make high-quality healthcare accessible to everyone

and radically improve drug discovery by linking real-world healthcare data with genomics,

imaging, biomarkers, and experimental data. This creates a foundation for AI systems that can

learn from millions of patient journeys, understand why treatments work for some patients and

fail for others, personalize care based on clinical and genomic context, identify the right

interventions earlier, and uncover new therapeutic opportunities from the connection between

biology and real-world outcomes.


BioStack is backed by PeakXV, Y Combinator, Afore Capital, SV Angel as well as high-profile angels from OpenAI, Meta and Google DeepMind.


About the Role


As an RL Engineer at BioStack, you will build reinforcement learning environments and post-training systems for healthcare AI.

BioStack is building the data and environment layer for medical AI: sourcing high-value clinical data, turning it into model-ready workflows, and building tasks, rewards, verifiers, benchmarks, and agent environments where models can learn against meaningful and measurable outcomes.

You will work across the full RL loop — from environment and reward design to training, evaluation, and iteration. Projects may span clinical reasoning, longitudinal patient care, diagnostic decision-making, chronic disease management, and biomedical research.

This is a hands-on engineering role. You will build environments, run experiments, train models and agents, analyze failures, and improve the data and feedback signals that determine what models learn.

Strong judgment around data is particularly important. You should be able to determine whether a dataset has the signal quality, label fidelity, coverage, diversity, and clinical relevance required to support useful training tasks, rewards, and evaluations.

Prior healthcare experience is not required.


This is a full-time in-person role based in San Francisco, CA. 


What you will do:


  • Build healthcare-specific RL environments, including tasks, action spaces/tool interfaces, reward functions, verifiers, and evaluation harnesses.
  • Run post-training experiments on language models and agents using techniques such as SFT, RLVR, RLHF/RLAIF, and reward modeling.
  • Turn clinical and biomedical datasets into training environments with measurable, verifiable outcomes.
  • Design rewards and verifiers that capture correctness across clinical reasoning and longitudinal decision-making tasks.
  • Train and evaluate multi-step agents operating across patient histories, clinical tools, and structured/unstructured medical data.
  • Build scalable pipelines for rollouts, training, evaluation, experiment tracking, and dataset iteration.
  • Analyze model failures and use them to improve environments, rewards, datasets, and subsequent training runs.


You might thrive in this role if:

  • You are excited by the idea of applying frontier RL methods to healthcare, medicine, and biological data.
  • You have experience with reinforcement learning, language model post-training, agent environments, reward modeling, evaluation, or related ML systems.
  • You have strong judgment around data and can assess whether a dataset has sufficient signal quality, label fidelity, coverage, longitudinal depth, and clinical relevance to support meaningful training tasks, environments, rewards, and evaluations.
  • You can move quickly from research concept to working prototype, then iterate based on empirical results.
  • You are comfortable designing controlled experiments, building baselines, and drawing trustworthy conclusions from noisy real-world data.
  • You are comfortable working in large ML codebases and can debug training runs, data pipelines, eval harnesses, and model behavior.
  • You care about building systems that are technically rigorous, clinically grounded, and useful beyond demos.
  • You are a self-starter who can own ambiguous problems, define the right technical path, and drive projects to completion.
  • You thrive in a fast-moving startup environment where research, engineering, product, and customer needs all intersect.
  • You have 2+ years of experience working on reinforcement learning, post-training, agents, or related ML systems.




Compensation and benefits:


  • Base salary: $200,000–$350,000 per year
  • Equity: 0.1%–1.0%
  • 401(k): Company match
  • Health benefits: Medical, dental, and vision insurance
  • Meals: Complimentary lunch provided daily at the office
  • AI tools: Unlimited budget for AI tools and software
  • Relocation & joining bonus: Available based on role and circumstances


Work authorization: Visa sponsorship is available for suitable candidates


Equal Opportunity

BioStack is an equal opportunity employer. We are committed to building a diverse and inclusive team and do not discriminate on the basis of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or any other characteristic protected by applicable law. All qualified applicants will receive consideration for employment.


A note for you:

You may be early in your career—just graduating or coming in with a few internships. That is completely fine. We are a young team too.

The real question is how you are wired.

You work toward something for months, finally achieve it, and almost immediately start thinking about what comes next. You want harder problems, more responsibility, and a steeper learning curve. If that sounds like you, you will fit in here. There is no ceiling at BioStack.

We are building our own version of a small group of unconventional, relentlessly driven people who perform exceptionally well when the stakes are high.

The work is technically difficult, operationally messy, and deeply consequential. We need people who want to become world-class, not merely competent.

You should want to become one of the best engineers of your generation. We will give you ambitious problems, real ownership, direct feedback, and the pressure and support to discover abilities you may not know you have.

It will be intense, demanding, and—for the right person—one of the most rewarding periods of their career.


Similar Jobs

Explore other opportunities that match your interests

Founding Forward Deployed Engineer

Programming
•
4h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

inventure

San Francisco Bay Area

Forward Deployed Engineer

Programming
•
4h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

inventure

San Francisco Bay Area
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

leadervest

San Francisco Bay Area

Subscribe our newsletter

New Things Will Always Update Regularly