P

Mid-Senior Level Research Engineer (AI for Scientific Discovery)

periodic labs • United State
Visa Sponsorship
Apply
AI Summary

Join periodic labs to advance frontier AI models for scientific discovery by curating data, generating synthetic datasets, and optimizing large-scale training experiments. This role focuses on improving scientific reasoning in models, collaborating with cross-functional teams, and scaling infrastructure across thousands of GPUs. Requires expertise in LLM training, evals, and distributed systems with a passion for AI-driven scientific breakthroughs.

Key Highlights
Train frontier AI models to enhance scientific reasoning and discovery in materials, energy, and beyond
Curate novel scientific datasets, generate synthetic data, and build performance-correlated evaluations
Design and execute large-scale training experiments with distributed compute scaling expertise
Key Responsibilities
Identify, process, and curate novel scientific datasets for large-scale model training
Generate high-quality synthetic data to address gaps in scientific knowledge and reasoning
Develop and refine evaluations that measure downstream scientific task performance
Apply techniques like self-distillation and on-policy distillation to improve model capabilities
Design and execute large-scale training experiments with supercompute engineers
Build tools to analyze how data choices impact model intelligence
Technical Skills Required
Large Language Model (LLM) Training Distributed Training Systems Data Curation and Synthetic Data Generation
Benefits & Perks
$250,000–$350,000 annual compensation + equity
Visa sponsorship available
Hybrid work model (Menlo Park, CA + San Francisco, CA)
Nice to Have
Experience optimizing distributed training throughput and reliability
Background in AI for scientific domains (e.g., protein, materials)
Experience creating synthetic data for non-verifiable tasks

Job Description


We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and a drive to push the boundaries of what's scientifically possible.

About The Role

We're training frontier models to develop deep scientific knowledge and reasoning for scientific discovery. As a Midtraining Research Engineer, you'll take base models and improve their scientific reasoning: curating and generating data, building evals, and running large-scale training experiments. Your work will also lay the groundwork for our pre-training efforts down the line.

What You'll Do

  • Identify, process, and curate novel sources of scientific data for large-scale model training.
  • Generate high-quality synthetic data to fill gaps in scientific knowledge and reasoning.
  • Build evaluations that correlate with downstream scientific task performance, working closely with RL researchers, physicists, and chemists.
  • Develop and apply techniques such as self-distillation and on-policy distillation to improve model capability.
  • Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs.
  • Build tools for yourself and the team to investigate how data choices shape model intelligence.

You Will Thrive in This Role If You Have

  • Experience training LLMs on curated mixes of trillions of tokens.
  • Experience on a dedicated evals team supporting a large production training run.
  • Hands-on use of self-distillation, on-policy distillation, or similar methods in a real training pipeline.
  • Experience with scaling laws and compute-optimal hyperparameters.
  • Comfort working across data, evals, and training infrastructure.

Especially Strong Candidates May Also Have

  • Experience optimizing throughput and reliability for large-scale distributed training runs.
  • A background in AI for science or training on specialized domain data (e.g., protein, materials, or other scientific datasets).
  • Experience creating evals or synthetic data for non verifiable tasks and tracking performance over live runs.

Mechanics

  • Minimum education: Bachelor's degree or similar experience
  • Location: Menlo Park, CA (Soon: San Francisco, too)
  • Compensation: $250,000–$350,000 + equity
  • Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process.

Similar Jobs

Explore other opportunities that match your interests

Principal Engineer, Agentic AI Systems (Lead Architect)

Programming
•
17m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

inflection ai

United State

Advanced Process Controls Engineer (Manufacturing & Optimization)

Programming
•
29m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Corning Incorporated

United State

Senior Manager, AI Engineering - Intelligent Foundations and Experiences (IFX)

Programming
•
51m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Capital One

United State

Subscribe our newsletter

New Things Will Always Update Regularly