Own and operate the multi-cloud Kubernetes infrastructure supporting AI model workloads, including GPU orchestration and real-time networking. Lead the expansion of infrastructure across AWS and GPU cloud providers using GitOps and infrastructure-as-code. Requires deep production experience with Kubernetes, GPU scheduling, and observability stacks.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
Department: Engineering
Location: San Francisco
You'll own the infrastructure platform that our AI models run on. This isn't a CI/CD-focused DevOps role. You'll work across GPU orchestration, multi-cloud Kubernetes, real-time networking, and observability. You'll be the person who knows why a model pod took 4 minutes to schedule, why cross-region latency spiked, or why a media relay is dropping packets.
We run production today across multiple Kubernetes clusters, regions, and GPU types, and we're actively expanding to additional cloud providers. You'll lead that expansion and keep everything running.
What You'll Do
- Provision and manage multi-region Kubernetes clusters across AWS and GPU cloud providers using infrastructure-as-code.
- Own the GitOps deployment lifecycle (Helm charts, Kustomize overlays, image automation, and continuous delivery.)
- Manage GPU node infrastructure: scheduling, model weight caching, image prefetching for fast cold starts, and GPU observability.
- Operate and improve our networking layer: ingress and gateway management, load balancing, media relay infrastructure, and cross-region connectivity.
- Build and maintain our observability stack: metrics, logs, traces, and profiling across all services and GPU workloads.
- Maintain infrastructure security: IAM, secret management, certificate automation, and encryption at rest.
- Own CI/CD pipelines for monorepo builds spanning Go services, Python model containers, and Helm chart releases.
- Partner with ML engineers on model serving: container optimization, health checks and startup tuning, media pipeline performance, and multi-GPU configuration.
Looking to advance your Devops career with relocation support? Explore Devops Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
- You've operated Kubernetes in production at scale, not just deployed to it, but debugged node-level scheduling issues, tuned autoscalers, and managed cluster upgrades.
- Strong infrastructure-as-code experience (Terraform, Pulumi, or similar) across multiple environments and regions.
- You've worked with GPU workloads on Kubernetes: device plugins, node taints/tolerations, GPU-aware scheduling. You understand why bin-packing matters
for expensive hardware. - Experience with GitOps tooling (FluxCD, ArgoCD, or similar) and Helm chart authoring.
- Comfortable with Redis or similar in-memory data stores (replication, persistence, pub/sub or streaming patterns)
- Familiarity with modern observability stacks (Prometheus, Grafana, OpenTelemetry, or equivalent) and knowing when to reach for metrics vs. logs vs. traces.
- Solid networking fundamentals: load balancers, TLS, DNS, NAT. Real-time or low-latency networking experience is a strong plus.
- You've worked in a startup where you owned infrastructure end-to-end, not just one slice of it.
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
- Experience with GPU cloud providers beyond AWS (Crusoe, CoreWeave, Lambda Labs, Nebius)
- Real-time media or streaming infrastructure
- Go or Python proficiency
- Familiarity with ML model serving (container image optimization, weight loading, GPU driver and runtime management)
- FinOps and GPU cost optimization
Interested in relocating to United State? Check out our comprehensive Relocation Jobs in United State page with detailed relocation packages and benefits.
- Pure CI/CD pipeline engineers who haven't operated Kubernetes clusters directly
- Candidates whose infrastructure experience is limited to managed PaaS (Heroku, Vercel, Railway)
- People who need a fully defined scope, this role requires figuring out what to build next, not just executing tickets
- Competitive San Francisco salary and meaningful equity
- We sponsor visas and support relocation to the US
- Generous health, dental, and vision coverage
Similar Jobs
Explore other opportunities that match your interests