Founding GPU Infrastructure Engineer (HPC & AI Systems)
Build and scale a GPU neocloud infrastructure from bare metal to production, designing high-performance HPC and AI systems. Own GPU cluster commissioning, Kubernetes/Slurm orchestration, distributed training, and observability for enterprise-grade AI workloads. Founding role with meaningful equity, competitive compensation, and relocation support in San Francisco.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
FOUNDING GPU INFRASTRUCTURE ENGINEER [HPC & AI SYSTEMS] - MEANINGFUL EQUITY
REPEAT FOUNDERS WITH SERIOUSLY IMPRESSIVE EXITS.
Build a GPU neocloud from bare metal to inference
San Francisco | On-site (5 days with some flex) | Relocation support
Most infrastructure roles ask you to keep somebody else’s platform alive.
This one asks you to build the platform.
Sentiro Partners has been retained to find the Founding GPU Infrastructure Engineer for my client, a seed-stage GPU neocloud operating in stealth in San Francisco.
The company is building managed GPU infrastructure and high-performance inference for demanding AI workloads. It has secured backing from well-known VCs and is pursuing deployments at hundreds-of-GPUs scale.
THE FOUNDERS HAVE DONE THIS BEFORE
One founder built an AI infrastructure company operating at the HW layer, bringing low-latency AI computation to resource-constrained devices. Acquired.
The other founded & scaled an enterprise technology company to 400+ customers.
Their previous companies were backed by major VCs. Ivy league & top tier finance backgrounds. Seriously down to earth. Seriously hard workers. They are in the SoMa office right now, as you read this. Now they’re building again.
WHAT SUCCESS LOOKS LIKE
Success means taking new GPU capacity from delivered racks to reliable, customer-ready infrastructure and building the systems and team required to repeat that process at increasing scale.
You will be the company’s founding infrastructure engineer and one of its first technical hires.
When the racks arrive, you will help turn them into a production platform:
Looking to advance your Development & Programming career with relocation support? Explore Development & Programming Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
- Commission and burn in NVIDIA GPU clusters.
- Build and operate Kubernetes and Slurm infrastructure.
- Develop the orchestration layer above the underlying clusters.
- Enable distributed training, MPI and batch workloads.
- Architect distributed object storage and high-performance NVMe systems.
- Build around InfiniBand, RDMA and other high-bandwidth networking.
- Own telemetry, observability, reliability and cluster lifecycle tooling.
- Help develop low-latency, high-throughput inference infrastructure.
- Create the interfaces through which customers access and operate the platform.
- Establish the patterns that make each deployment faster and more reliable than the last.
THIS IS GPU INFRASTRUCTURE BUILT WITH HPC DISCIPLINE
You should understand what happens between a rack arriving at a facility and a customer successfully running a distributed GPU workload.
You don’t need to be a data-centre real-estate or cooling specialist. You do need enough hardware and systems depth to reason about GPU topology, storage, networking, workload behaviour and production reliability as one connected system.
WE SHOULD TALK IF YOU HAVE
- c. 3-7 years in GPU infrastructure, HPC, AI systems or distributed systems (we're not limited by exp. but you need to be able to go deep).
- Meaningful production experience with both Kubernetes and Slurm (either tbh)
- Built infrastructure rather than only deployed platforms created by other teams.
- Strong knowledge of distributed object storage, NVMe and high-performance networking.
- Experience with distributed training, MPI, InfiniBand, RDMA or similar technologies.
- Worked across hardware, systems software and production operations.
- The ability to move quickly and make sound decisions with incomplete information.
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
- The appetite to own outcomes rather than defend a narrow area of responsibility.
Particularly relevant backgrounds include AI labs, GPU neoclouds, quantitative trading firms, hyperscale infrastructure teams and serious HPC environments.
Experience with GPU power capping, power-aware scheduling, energy telemetry or grid-flexible compute would be unusually valuable, but it isn’t essential.
WHY NOW?
Because the upside is real.
You will work directly with the founders who have already raised capital, scaled teams and delivered successful exits. You will influence the architecture before the boundaries harden, help deliver the company’s first major GPU clusters and have the opportunity to build the infrastructure team around you.
This is a founding role with a competitive San Francisco salary, performance bonus, benefits and meaningful equity.
The team is building in person in San Francisco. Relocation support is available for exceptional candidates ready to move.
If your ideal role begins where the racks arrive—and ends with customers running serious AI workload... contact me directly.
Adrian Clarke
Founder, MD, Executive Search Partner Sentiro Partners
Frontier AI Search & Advisory
(https://sentiropartners.com)
ABOUT SENTIRO PARTNERS
- Sentiro Partners is a global executive search firm specialising in frontier AI, deep technology and high-performance technical leadership.
- We partner with ambitious founders, AI labs and technology companies to identify the engineers, researchers and leaders building the next generation of intelligent infrastructure.
- Headquartered in Dublin, Sentiro Partners conducts searches across North America, Europe and Asia-Pacific.
Similar Jobs
Explore other opportunities that match your interests