S

Senior Cloud Operations Engineer (NOC) – AWS & Kubernetes (24/7 Support)

Spectrum IT Recruitment United Kingdom
Remote
Apply
AI Summary

Build and maintain highly resilient cloud platforms in AWS while supporting 24/7 production operations. Focus on incident response, automation, observability, and collaboration with engineering teams to enhance reliability. Requires expertise in Linux, AWS, Kubernetes, and scripting for large-scale environments.

Key Highlights
24/7 on-call support with a 28-day rotating shift pattern
Engineering-led role emphasizing automation, resilience, and continuous improvement
Collaborative environment with cross-functional teams (Software, Platform, Cloud, Security)
Key Responsibilities
Monitor and maintain highly available production platforms running in AWS
Respond to and manage production incidents across a 24/7 service
Develop automation to reduce manual operational tasks and improve platform resilience
Build and improve monitoring, alerting, and observability across cloud environments
Investigate complex technical issues and restore services quickly and effectively
Support containerized workloads using Kubernetes and Docker
Contribute to post-incident reviews and drive continuous service improvements
Technical Skills Required
Amazon Web Services Kubernetes Linux systems administration
Benefits & Perks
Competitive salary + bonus
Excellent benefits
Fully remote (UK-based)
Nice to Have
Infrastructure as Code (Terraform)
Site Reliability Engineering (SRE) principles (SLIs, SLOs)
Experience in regulated environments
Python, Bash, or Go scripting
Monitoring and observability platforms (Grafana, Prometheus, Datadog, Splunk, CloudWatch)
Networking fundamentals (DNS, TCP/IP, load balancing)

Job Description


NOC Engineer + AWS | Kubernetes

  • Fully Remote (UK)
  • 24/7 Shift Pattern (28-day rota including days & nights)
  • £ Competitive + Bonus + Excellent Benefits

Build resilient cloud platforms that support critical national services.

This is far more than a traditional NOC role.

You'll be joining an engineering-led organisation where reliability, automation and continuous improvement sit at the heart of the platform. Rather than simply responding to incidents, you'll work to prevent them by improving systems, automating operational processes and helping shape the future of highly resilient cloud services.

If you're passionate about building reliable cloud platforms and enjoy solving complex technical problems in large-scale production environments, we'd love to hear from you.

What you'll be doing

  • Monitoring and maintaining highly available production platforms running in AWS
  • Responding to and managing production incidents across a 24/7 service
  • Investigating complex technical issues and restoring services quickly and effectively
  • Developing automation to reduce manual operational tasks and improve platform resilience
  • Building and improving monitoring, alerting and observability across cloud environments
  • Working alongside Software, Platform, Cloud and Security Engineers to improve reliability and operational excellence
  • Contributing to post-incident reviews and driving continuous service improvements
  • Supporting containerised workloads using Kubernetes and Docker

What we're looking for

You'll ideally have experience in a Production Engineering, Cloud Operations or NOC environment with exposure to:

  • Linux systems administration
  • AWS cloud infrastructure
  • Kubernetes and Docker
  • Production support and incident management
  • Python, Bash or Go scripting
  • Monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk or CloudWatch
  • Networking fundamentals including DNS, TCP/IP and load balancing
  • A passion for automation, continuous improvement and operational excellence

Experience with Infrastructure as Code (Terraform), SRE principles (SLIs, SLOs), or regulated environments would be beneficial but isn't essential.

Why join?

Apply today or contact Dave Carlisle at Spectrum IT Recruitment for a confidential discussion.

Spectrum IT Recruitment (South) Limited is acting as an Employment Agency in relation to this vacancy.


Similar Jobs

Explore other opportunities that match your interests

Senior Remote Service Engineer (EV Charging Infrastructure)

Devops
1d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

abb e-mobility

United Kingdom

EMEA Events Manager

Devops
3d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Grafana Labs

United Kingdom
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

Eden Scott

United Kingdom

Subscribe our newsletter

New Things Will Always Update Regularly