Build and optimize an autonomous coding agent that develops, validates, and improves LLM inference stack components—including GPU kernels, runtime, and serving infrastructure—with strict correctness and performance requirements. Design evaluation loops, orchestrate parallel development, and ensure hardware-validated correctness. Requires deep expertise in production AI agents, LLM inference systems, or GPU kernel development with strong Python and low-level systems coding skills.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
WorkorAI is recruiting on behalf of an early-stage AI infrastructure company building a coding agent that creates and improves an entire LLM inference stack: GPU kernels, runtime, and serving infrastructure.
This is a senior engineering role for someone whose experience combines production AI agents with either LLM inference systems or GPU kernel development.
About the role
The company is developing an agent that receives specifications and test results, writes low-level code, runs it against real hardware or simulation, analyzes failures, and keeps iterating until strict correctness and performance gates pass.
There is no room for plausible-looking output that does not work. Every component is validated against reference implementations, test suites, and hardware benchmarks.
On its first target, a proprietary accelerator with no existing inference ecosystem, the system reached working tensor-parallel matrix multiplication in approximately 10 hours and ran three frontier models end to end within 10 days. It is now serving production traffic.
What you will do
• Build the central agent controller that writes, executes, evaluates, and improves inference code.
Searching for Development & Programming roles that provide visa sponsorship? Connect with international employers through Development & Programming Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
• Design evaluation loops that detect incorrect output, compare implementations against references, and enforce performance gates.
• Own the test ladder used to validate every layer before other components are built on top of it.
• Orchestrate parallel work across kernels, runtime components, and serving infrastructure, including retry, dependency, escalation, and human-review logic.
• Review and validate low-level code produced by the agent against real hardware and simulation results.
What we are looking for
You have production experience building with LLMs or coding agents, including tool-use loops, constrained code generation, evaluation harnesses, and systems where tests catch model errors.
You also have meaningful depth in at least one of these areas:
• LLM inference systems: vLLM, KV cache, paged attention, continuous batching, quantization, speculative decoding, FlashAttention, or low-latency serving.
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
• GPU and accelerator kernels: CUDA, Triton, ROCm/HIP, Metal, attention, matrix multiplication, normalization, MoE, kernel fusion, or performance optimization.
Strong Python skills are required. You should also be comfortable reading or reviewing C++, Rust, CUDA C, or similarly low-level systems code.
Experience with LLVM, MLIR, TVM, compiler code generation, accelerator bring-up, NCCL, Megatron-LM, DeepSpeed, distributed inference, firmware, or non-GPU execution models is valuable but not required.
Why this role
• $200,000–$420,000 compensation range.
• San Francisco preferred; remote work can be discussed for a strong fit.
• Visa sponsorship may be available, including H-1B.
• $15M raised from investors focused on AI infrastructure and silicon.
Interested in opportunities specifically in United State? Discover our dedicated Visa Sponsorship Jobs in United State page featuring roles from top employers in this location.
• Approximately 14 engineers, with engineering and product operating as one team.
• Founded by a former Google Brain researcher.
• Production systems with measurable correctness and performance feedback rather than demo-only agent workflows.
How to apply
Apply through WorkorAI:
https://workorai.com/candidate/apply/cmsrca19o0001tpbqjf9vw7kl
The application includes a WorkorAI profile and a short role-focused AI interview. The complete process should take no more than 15 minutes.
WorkorAI is managing sourcing and initial technical evaluation for this search.
Similar Jobs
Explore other opportunities that match your interests