S

Senior Platform Engineer (AWS, Kubernetes, Backend Systems)

Second Talent • San Francisco Bay Area
Visa Sponsorship Relocation
Apply
AI Summary

Own the reliability, scalability, and developer experience of a high-growth AI infrastructure platform. Design and optimize cloud-native systems, CI/CD pipelines, and observability while balancing performance, cost-efficiency, and production reliability. Requires deep expertise in AWS, Kubernetes, Terraform, and backend engineering principles.

Key Highlights
Own production uptime, latency, cloud costs, and incident response for a high-availability AI infrastructure platform.
Combine production infrastructure expertise with backend engineering to design scalable, cost-efficient systems.
Work across AWS, Kubernetes/EKS, Terraform, CI/CD, observability, and backend services to improve developer productivity.
Key Responsibilities
Own production uptime, latency, infrastructure provisioning, cloud costs, and incident response.
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, and networking services.
Design and improve backend and platform systems for scale, including autoscaling, queueing, retries, and rollback strategies.
Develop and maintain observability systems, dashboards, alerts, logging, tracing, SLOs, and on-call processes.
Develop CI/CD pipelines, release automation, and deployment workflows for reliable software delivery.
Write production-quality code for automation, backend services, and internal engineering tools.
Identify and implement improvements for system performance, reliability, developer productivity, and cloud efficiency.
Technical Skills Required
Amazon Web Services Kubernetes Terraform
Benefits & Perks
Competitive compensation based on experience and location
Relocation and visa sponsorship for strong candidates moving to the US
Comprehensive medical, dental, and vision coverage
Nice to Have
Experience with AI/ML, data-heavy, or enterprise platforms
Handling bursty workloads, long-running jobs, or distributed systems
Optimizing cloud infrastructure costs through architecture and observability
Building internal platforms and developer tools
Experience in early-stage or high-growth technology companies

Job Description


About Our Client

Our client is a rapidly growing AI infrastructure company building the systems and data infrastructure that power the next generation of AI agents. Its platform supports frontier AI labs, Fortune 500 companies, and high-growth technology companies.

Backed by significant venture funding and experiencing strong revenue growth, our client is scaling its engineering team to meet growing demand. The engineering team includes AI researchers, startup founders, and multiple international Olympiad medalists.


About the Role

Our client is looking for a Platform Engineer to own the reliability, scalability, performance, and developer experience of its core infrastructure and backend systems.

This is not a pure infrastructure role. The ideal candidate combines strong production infrastructure experience with solid backend engineering skills and can reason about service architecture, APIs, databases, queues, performance, deployment safety, and production reliability.


You'll work across AWS, Kubernetes, Terraform, CI/CD, observability, and backend services to make systems more reliable, scalable, cost-efficient, and easier for engineers to build on.


Key Responsibilities

  • Own production uptime, latency, infrastructure provisioning, cloud costs, and incident response.
  • Build and maintain AWS infrastructure using technologies such as Terraform, Kubernetes/EKS, Helm, Docker, EC2, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale, including autoscaling, queueing, retries, backpressure, capacity planning, and rollback strategies.
  • Build dashboards, alerts, logging, tracing, SLOs, runbooks, and on-call processes.
  • Develop reliable CI/CD, release automation, environment management, and deployment workflows.
  • Write clean, maintainable code for automation, backend services, and internal engineering tools.
  • Identify opportunities to improve system performance, reliability, developer productivity, and cloud efficiency.


Requirements

  • Experience owning production cloud infrastructure for a high-availability, user-facing platform.
  • Strong experience with AWS and containerized systems.
  • Hands-on experience with Terraform, Kubernetes/EKS, Docker, EC2, CI/CD, networking, IAM, and cloud security.
  • Experience operating observability, alerting, incident response, and deployment systems.
  • Strong backend engineering fundamentals, including service architecture, APIs, databases, asynchronous systems, queues, and distributed systems.
  • Ability to write production-quality code and apply sound engineering judgment across infrastructure and backend systems.
  • Strong ownership mindset and ability to operate effectively in a fast-moving environment.


Nice to Have

  • Experience with AI/ML, data-heavy, workflow, marketplace, developer-tool, or enterprise platforms.
  • Experience handling bursty workloads, long-running jobs, distributed workers, sandboxed execution, or high-concurrency services.
  • Experience optimizing cloud infrastructure costs through architecture, autoscaling, workload placement, caching, cleanup, or observability.
  • Experience building internal platforms and developer tools.
  • Experience working in an early-stage or high-growth technology company.


Our client prioritizes technical aptitude, ownership, and learning potential over years of experience.


What Our Client Offers

  • Competitive compensation based on experience and location.
  • Relocation and visa sponsorship for strong candidates moving to the US.
  • Comprehensive medical, dental, and vision coverage.
  • Office-provided lunch and dinner.
  • Paid time off and company-wide holiday break.
  • 401(k) and commuter benefits.
  • Fitness membership benefits.
  • Generous access to leading AI development tools, including ChatGPT, Claude Code, and Cursor.
  • Opportunity to work alongside exceptional AI researchers, engineers, and founders on cutting-edge infrastructure.

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

parallel works

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Wypoon Technologies

Netherlands
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Helsing

United Kingdom

Subscribe our newsletter

New Things Will Always Update Regularly