S

Senior Site Reliability Engineer (SRE) - Cloud & Platform Infrastructure

superbot (bot) United Arab Emirates
Visa Sponsorship
Apply
AI Summary

Lead the design and maintenance of scalable, observable, and resilient cloud infrastructure at superbot, shaping reliability practices and on-call culture for a growing engineering team. Own Kubernetes, AWS, and observability stacks while partnering with product teams to drive infrastructure-as-code and CI/CD improvements. Requires 4+ years of hands-on SRE/DevOps experience with production-scale ownership.

Key Highlights
Own reliability and observability for critical services, defining SLIs/SLOs/SLAs and reducing MTTR through incident management
Design and maintain AWS-based cloud infrastructure using Terraform, ensuring reproducibility and auditability
Scale and optimize Kubernetes clusters with capacity planning, autoscaling, and lifecycle management
Key Responsibilities
Define and track SLIs/SLOs/SLAs, lead blameless post-mortems, and reduce MTTR through systematic incident management
Design, build, and maintain AWS cloud infrastructure using Terraform, ensuring version-controlled, reproducible environments
Scale and optimize Kubernetes clusters with capacity planning, resource tuning, autoscaling, and lifecycle management
Strengthen observability by instrumenting services with Prometheus/Grafana, building actionable alerts, and reducing alert noise
Accelerate CI/CD pipelines (GitHub Actions/GitLab CI) to support frequent, safe deployments and embed reliability checks
Partner with engineering teams as an internal reliability advisor, advocating SRE best practices and running game days
Technical Skills Required
Amazon Web Services Kubernetes Terraform
Benefits & Perks
Competitive compensation with performance-based bonuses
Work visa processing assistance
Generous holiday and New Year bonuses

Job Description


About The Role

You’ll sit at the intersection of software engineering and infrastructure, partnering with product and platform teams to keep our systems observable, scalable, and resilient. From shaping our on-call culture to driving infrastructure-as-code adoption, you’ll have real influence over how BOT builds and operates software at scale — for 500+ people and growing.

Key Responsibilities

  • Own reliability across critical services — define and track SLIs/SLOs/SLAs, lead blameless post-mortems, and drive down MTTR through systematic incident management.
  • Design, build, and maintain cloud infrastructure on AWS (primary) using Terraform, ensuring environments are reproducible, version-controlled, and auditable.
  • Scale and optimize our Kubernetes-based container platform — capacity planning, resource tuning, autoscaling, and cluster lifecycle management.
  • Strengthen observability end-to-end: instrument services with Prometheus and Grafana, build actionable alerting, and reduce alert noise through continuous tuning.
  • Accelerate CI/CD pipelines (GitHub Actions / GitLab CI) to support frequent, safe deployments — shift reliability left by embedding checks into the delivery workflow.
  • Partner with engineering teams as an internal reliability advisor — run game days, advocate for SRE best practices, and help developers build services that operate well from the start.

Job Requirements

  • 4+ years in an SRE, DevOps, or platform engineering role with production ownership at scale.
  • Hands-on Kubernetes experience — deployment, scaling, networking, and troubleshooting in a production environment.
  • Infrastructure-as-code fluency with Terraform (or Pulumi) across a major cloud provider (AWS preferred).
  • Observability stack experience — Prometheus, Grafana, and/or equivalent tools with a track record of building meaningful dashboards and alerts.
  • Proficiency in at least one scripting or programming language (Python, Go, or Bash) for automation and tooling.
  • AWS certification (Solutions Architect, DevOps Engineer, or SysOps Administrator).

What We Offer

  • Competitive Compensation: Enjoy a salary package tailored to your skills and experience, along with performance-based bonuses.
  • Comprehensive Benefits: We support your well-being with accommodation, meal allowances, and assistance with work visa processing.
  • Work-Life Balance: Unwind with generous holiday and New Year bonuses.
  • Top-Tier Equipment: Stay productive with the latest tools, including a MacBook and iPhone.
  • Thriving Culture: Immerse yourself in a dynamic, inclusive work environment that fosters growth.

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

deeplight ai

United Arab Emirates
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

jungle gaming

United Arab Emirates

Engineering Manager

Devops
2w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

talabat

United Arab Emirates

Subscribe our newsletter

New Things Will Always Update Regularly