A

Senior Chaos Engineer (Resilience Automation Specialist) - Enterprise Client Partnership

ai talent Australia
Visa Sponsorship
Apply
AI Summary

Lead enterprise-grade chaos engineering to validate fault tolerance and zero-downtime reliability for mission-critical digital systems. Design, automate, and execute fault injection experiments across distributed environments, bridging SRE and architecture teams. Collaborate to harden systems and embed resilience into CI/CD pipelines while ensuring observability and blast-radius control.

Key Highlights
Design and automate fault injection experiments for distributed systems resilience testing
Deploy and manage enterprise chaos engineering platforms (e.g., Chaos Mesh, Gremlin)
Collaborate with SRE and engineering teams to translate findings into architectural improvements
Key Responsibilities
Design, execute, and automate controlled chaos experiments across distributed microservices and containerized environments
Deploy and manage enterprise chaos engineering platforms (e.g., Chaos Mesh, Gremlin, LitmusChaos, AWS Fault Injection Simulator)
Organize and facilitate GameDay scenarios to test system recovery, automated failover mechanisms, and operational runbooks
Embed automated resilience assertions and steady-state validation into continuous deployment pipelines
Monitor real-time telemetry (metrics, traces, logs) during experiments to ensure safe rollback triggers
Collaborate with software engineering and cloud infrastructure teams to translate experiment findings into architectural improvements
Technical Skills Required
Site Reliability Engineering (SRE) Chaos Engineering Kubernetes
Benefits & Perks
Full visa and migration support (482 On-Hire Sponsorship Transfers, New 482 Visa Sponsorship, Temporary/Working Visa Holders)
Remote work eligibility
Employer of record status
Nice to Have
Practical experience implementing distributed tracing (OpenTelemetry, Jaeger)
Certified Kubernetes Administrator (CKA) or relevant Cloud certifications
Experience in high-availability, low-latency domains (Banking, Fintech, Telecommunications, E-Commerce)

Job Description


We are partnering with an enterprise client to strengthen platform resilience, validate fault tolerance, and ensure zero-downtime reliability for mission-critical digital systems. We are seeking a skilled Chaos Engineer (Resilience Automation Specialist) to represent our organisation and take technical ownership of fault injection experiments, automated resilience testing, and blast-radius verification.

In this role, you will bridge Site Reliability Engineering (SRE) and systems architecture within our client’s environment. You will intentionally simulate real-world failure scenarios—such as network latency, node dropouts, database failovers, and region outages—to uncover hidden failure modes and prove that distributed systems self-heal under duress.


🌏 Visa & Sponsorship Options

As the employer of record, we provide full visa and migration support for qualified engineering talent deployed to our clients:

  • 482 On-Hire Sponsorship Transfers: Fully supported for qualified candidates currently in Australia on an existing 482 visa looking to transfer sponsorship to work with our clients.
  • New 482 Visa Sponsorship: Available for qualified candidates meeting the commercial experience and technical requirements.
  • Temporary & Working Visa Holders: Open to all working visa holders seeking a direct pathway to employer sponsorship.
Core Responsibilities
  • Chaos Experimentation: Design, execute, and automate controlled chaos experiments across distributed microservices and containerized environments.
  • Fault Injection Tooling: Deploy and manage enterprise chaos engineering platforms (e.g., Chaos Mesh, Gremlin, LitmusChaos, AWS Fault Injection Simulator).
  • GameDays & Failure Drills: Organize and facilitate systematic GameDay scenarios to test system recovery, automated failover mechanisms, and operational runbooks.
  • Resilience in CI/CD: Embed automated resilience assertions and steady-state validation directly into continuous deployment pipelines.
  • Observability & Blast Radius Control: Monitor real-time telemetry (metrics, traces, logs) during experiments using tools like Prometheus, Grafana, Datadog, or OpenTelemetry, ensuring safe rollback triggers.
  • Architecture Hardening: Collaborate with software engineering and cloud infrastructure teams to translate experiment findings into actionable architectural improvements (circuit breakers, retry policies, auto-healing).
Selection Criteria
  • SRE & Resilience Mastery: Commercial experience in Site Reliability Engineering (SRE) or platform reliability with a strong focus on distributed systems resilience.
  • Chaos Engineering Tooling: Hands-on experience configuring fault-injection tools such as Chaos Mesh, Gremlin, LitmusChaos, or AWS FIS.
  • Container & Cloud Ecosystems: Deep understanding of Kubernetes architecture, container networking, service meshes, and cloud infrastructure (AWS/Azure).
  • Scripting & Automation: Strong programming or scripting skills in Python, Go, or Bash to build custom failure-injection scenarios and verification logic.
  • Location Requirements: Currently residing in Australia with valid work rights or eligibility for 482 visa sponsorship/transfer.
Preferred Qualifications (Nice to Have)
  • Practical experience implementing distributed tracing (OpenTelemetry, Jaeger).
  • Certified Kubernetes Administrator (CKA) or relevant Cloud certifications.
  • Experience in high-availability, low-latency domains (Banking, Fintech, Telecommunications, E-Commerce).



Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

ai talent

Australia

Platform Engineer

Devops
2w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

csiro

Australia
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

ai talent

Australia

Subscribe our newsletter

New Things Will Always Update Regularly