Ensure systems stay reliable, observable, and resilient at scale across multi-cluster environments. Build and maintain CI/CD pipelines, Infrastructure as Code, and observability stacks. Define reliability standards, manage SLOs/SLIs, and lead post-incident reviews for critical systems.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Job Description
At BairesDev®, we've been leading the way in technology projects for over 15 years. We deliver cutting-edge solutions to giants like Google and the most innovative startups in Silicon Valley.
Our diverse 4,000+ team, composed of the world's Top 1% of tech talent, works remotely on roles that drive significant impact worldwide.
When you apply for this position, you're taking the first step in a process that goes beyond the ordinary. We aim to align your passions and skills with our vacancies, setting you on a path to exceptional career development and success.
Senior Site Reliability Engineer (SRE) at BairesDev
In this role, you'll ensure systems stay reliable, observable, and resilient at scale, combining deep infrastructure expertise with a data-driven approach to reliability. Working across multi-cluster environments and modern observability stacks, you'll define what reliability actually means for critical systems and build the practices that keep them there. This is your opportunity to work where engineering rigor meets operational excellence, directly shaping the uptime and performance that users and businesses depend on.
What You'll Do
- Manage and scale Kubernetes environments across multiple clusters.
- Build and maintain CI/CD pipelines and Infrastructure as Code.
- Implement and maintain observability across systems for proactive issue detection.
- Define and track reliability standards, driving continuous improvement through incident learnings.
Interested in remote work opportunities in Devops? Discover Devops Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
- 5+ years of experience in Site Reliability Engineering or infrastructure engineering.
- Strong expertise in Kubernetes, including operators, autoscaling, and multi-cluster management.
- Experience with CI/CD pipeline engineering.
- Proficiency in Infrastructure as Code using Terraform and Helm.
- Hands-on experience with observability stacks such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Experience defining SLOs/SLIs, managing error budgets, and leading post-incident reviews.
- Background in cloud security tooling and IAM hardening.
- Advanced proficiency in English.
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
- 100% remote work (from anywhere).
- Excellent compensation in USD or your local currency if preferred
- Hardware and software setup for you to work from home.
- Flexible hours: create your own schedule.
- Paid parental leaves, vacations, and national holidays.
- Innovative and multicultural work environment: collaborate and learn from the global Top 1% of talent.
- Supportive environment with mentorship, promotions, skill development, and diverse growth opportunities.
Apply now!
#BD-PRIO-2026
Similar Jobs
Explore other opportunities that match your interests