Lead enterprise-grade chaos engineering to validate fault tolerance and zero-downtime reliability for mission-critical digital systems. Design, automate, and execute fault injection experiments across distributed environments, bridging SRE and architecture teams. Collaborate to harden systems and embed resilience into CI/CD pipelines while ensuring observability and blast-radius control.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
We are partnering with an enterprise client to strengthen platform resilience, validate fault tolerance, and ensure zero-downtime reliability for mission-critical digital systems. We are seeking a skilled Chaos Engineer (Resilience Automation Specialist) to represent our organisation and take technical ownership of fault injection experiments, automated resilience testing, and blast-radius verification.
In this role, you will bridge Site Reliability Engineering (SRE) and systems architecture within our client’s environment. You will intentionally simulate real-world failure scenarios—such as network latency, node dropouts, database failovers, and region outages—to uncover hidden failure modes and prove that distributed systems self-heal under duress.
As the employer of record, we provide full visa and migration support for qualified engineering talent deployed to our clients:
- 482 On-Hire Sponsorship Transfers: Fully supported for qualified candidates currently in Australia on an existing 482 visa looking to transfer sponsorship to work with our clients.
- New 482 Visa Sponsorship: Available for qualified candidates meeting the commercial experience and technical requirements.
- Temporary & Working Visa Holders: Open to all working visa holders seeking a direct pathway to employer sponsorship.
Searching for Devops roles that provide visa sponsorship? Connect with international employers through Devops Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
- Chaos Experimentation: Design, execute, and automate controlled chaos experiments across distributed microservices and containerized environments.
- Fault Injection Tooling: Deploy and manage enterprise chaos engineering platforms (e.g., Chaos Mesh, Gremlin, LitmusChaos, AWS Fault Injection Simulator).
- GameDays & Failure Drills: Organize and facilitate systematic GameDay scenarios to test system recovery, automated failover mechanisms, and operational runbooks.
- Resilience in CI/CD: Embed automated resilience assertions and steady-state validation directly into continuous deployment pipelines.
- Observability & Blast Radius Control: Monitor real-time telemetry (metrics, traces, logs) during experiments using tools like Prometheus, Grafana, Datadog, or OpenTelemetry, ensuring safe rollback triggers.
- Architecture Hardening: Collaborate with software engineering and cloud infrastructure teams to translate experiment findings into actionable architectural improvements (circuit breakers, retry policies, auto-healing).
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
- SRE & Resilience Mastery: Commercial experience in Site Reliability Engineering (SRE) or platform reliability with a strong focus on distributed systems resilience.
- Chaos Engineering Tooling: Hands-on experience configuring fault-injection tools such as Chaos Mesh, Gremlin, LitmusChaos, or AWS FIS.
- Container & Cloud Ecosystems: Deep understanding of Kubernetes architecture, container networking, service meshes, and cloud infrastructure (AWS/Azure).
- Scripting & Automation: Strong programming or scripting skills in Python, Go, or Bash to build custom failure-injection scenarios and verification logic.
- Location Requirements: Currently residing in Australia with valid work rights or eligibility for 482 visa sponsorship/transfer.
Interested in opportunities specifically in Australia? Discover our dedicated Visa Sponsorship Jobs in Australia page featuring roles from top employers in this location.
- Practical experience implementing distributed tracing (OpenTelemetry, Jaeger).
- Certified Kubernetes Administrator (CKA) or relevant Cloud certifications.
- Experience in high-availability, low-latency domains (Banking, Fintech, Telecommunications, E-Commerce).
Similar Jobs
Explore other opportunities that match your interests