B

Senior Infrastructure Engineer (GPU Compute & AI Orchestration)

Boundless United State
Remote
Apply
AI Summary

Lead the design, optimization, and operation of a globally distributed GPU compute fleet for AI inference workloads, maximizing performance, reliability, and cost efficiency. Orchestrate heterogeneous hardware across cloud and on-prem environments while driving deep GPU/bare-metal optimizations. Requires 5+ years in large-scale infrastructure with Kubernetes, Linux, and infrastructure-as-code expertise.

Key Highlights
Build and operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) for AI inference workloads
Maximize GPU utilization, reliability, and cost efficiency through orchestration, scheduling, and optimization
Deep expertise in bare-metal and GPU performance tuning (PCIe P2P, CUDA, NUMA, network topology) required
Key Responsibilities
Orchestrate GPU workloads across multi-region, heterogeneous fleets using Kubernetes, SkyPilot, and cloud/on-prem providers
Optimize GPU performance through bare-metal tuning (PCIe P2P, ReBAR, NUMA, CUDA, memory configuration) and network topology
Maximize GPU utilization and cost efficiency via intelligent workload placement (spot instances, on-prem/cloud hybrid scheduling)
Build secure fleet access (Tailscale, Teleport), observability, and zero-downtime deployment systems
Drive down cost per GPU-hour while maintaining reliability and high throughput
Technical Skills Required
Kubernetes Linux Systems Administration Infrastructure-as-Code (Terraform, Ansible, Pulumi)
Benefits & Perks
Competitive salary + equity allocation
Health, dental, and vision insurance (U.S. employees; region-adjusted globally)
Flexible PTO, professional development budget, and remote-first work environment
Nice to Have
Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
Familiarity with GPU fleet orchestration tools (SkyPilot, Ray, Slurm)
Proficiency in Rust or low-level systems programming

Job Description


Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization.

You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.

What You'll Do

GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably.

Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the "always-on inference substrate" saturated and economical.

Bare-Metal & GPU Optimization: Go deep on GPU performance — PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology — to push throughput per node.

Reliability, Access & Observability: Build secure fleet access (Tailscale, Teleport), robust observability and alerting, and zero-downtime rollouts across a distributed node fleet.

Cost Optimization: Drive down $/GPU-hr through spot instance management, intelligent workload placement between on-prem and cloud, and resource scheduling — without sacrificing reliability.

Requirements

  • 5+ years of infrastructure/DevOps experience operating large-scale production systems
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale
  • Strong Linux systems administration skills
  • Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
  • Track record of managing mission-critical, high-throughput systems
  • Strong infrastructure-as-code background in heterogeneous environments
  • Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
  • Comfort navigating ambiguity with a strong bias for action

Nice to Have

  • Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
  • Experience operating ML training or other large-scale distributed compute infrastructure
  • Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
  • Familiarity with fleet access and networking tooling (Tailscale, Teleport)
  • Knowledge of network optimization and topology design
  • Experience with multi-region, globally distributed systems
  • Proficiency in Rust or low-level systems programming
  • Experience with on-premises data center operations

Additional Requirements

  • Candidates must include a public GitHub profile in their application
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered

Benefits

At Boundless, we take care of our people, because building the future of AI compute starts with an empowered team. Here's what you can expect when you join us:

  • Competitive salary + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment

We are a global team, and applicants from around the world are welcome to apply.

Similar Jobs

Explore other opportunities that match your interests

Azure DevOps Engineer

Devops
5m ago
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

IT Associates

United State
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Entry level

aadhya technologies inc

United State

Senior DevOps Engineer

Devops
1h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

AAA Life Insurance Company

United State

Subscribe our newsletter

New Things Will Always Update Regularly