Build real-time inference infrastructure for next-gen multimodal foundation models, bridging research and production. Design scalable, low-latency systems for transformers, SSMs, and hybrid architectures while collaborating with researchers. Requires expertise in distributed systems, ML serving, and productionizing cutting-edge AI models.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
Help build the inference stack behind the next generation of multimodal foundation models.
Most inference roles are about making existing models run faster.
This one is about helping define how entirely new model architectures are served at scale.
You'll be building real-time multimodal AI capable of processing enormous streams of text, audio and video. The research is pushing beyond today's transformer limitations, and your work will make those models usable in production.
You'll sit between frontier research and product engineering, designing the infrastructure that allows cutting-edge models to run with low latency, high reliability and at scale. If you enjoy solving systems problems where every millisecond matters, you'll feel at home here.
Your focus
Looking to advance your Development & Programming career with relocation support? Explore Development & Programming Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
- Build low-latency inference and serving infrastructure for foundation models across Transformers, SSMs and hybrid architectures.
- Design scalable, reliable distributed systems that support production AI workloads.
- Develop monitoring and observability across the inference stack.
- Work closely with researchers to productionise new model architectures.
- Help shape technical direction with significant ownership from day one.
You'll bring
- Strong software engineering fundamentals and experience building large-scale distributed systems.
- Experience with ML inference pipelines or serving generative models in production.
- The ability to work through ambiguous technical challenges and deliver zero-to-one systems.
- Experience implementing modern machine learning research into production environments.
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
Experience with vLLM, SGLang, Continuous Batching, CUDA or Triton would be particularly valuable but isn't essential.
Salary: Competitive + equity.
Location: San Francisco (onsite).
Alongside compensation, you'll receive fully covered medical, dental and vision insurance, 401(k), relocation and immigration support, plus daily meals in the office.
If you're interested in building the infrastructure that enables the next generation of AI models to run in real time, we'd love to tell you more.
All applicants will receive a response.
Similar Jobs
Explore other opportunities that match your interests