As a Reliability Engineer at Anthropic, you will ensure the dependability of large language model serving systems, focusing on high availability, low latency, and safety. Your responsibilities include defining SLOs, designing monitoring systems, leading incident response, and enhancing infrastructure resilience across cloud providers. You need a strong background in distributed systems, ML infrastructure, and chaos engineering with excellent cross-team collaboration skills.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Job Description
About Anthropic
Anthropic is dedicated to developing reliable, interpretable, and steerable artificial intelligence systems. Our mission is to ensure that AI technologies are safe and beneficial for users and society at large. We are a rapidly expanding team comprising researchers, engineers, policy experts, and business leaders who collaborate closely to build AI systems that prioritize safety, transparency, and societal good. Our commitment to innovation and responsible AI development positions us at the forefront of the industry, working tirelessly to create AI that aligns with human values and needs.
About The Role
As a Reliability Engineer at Anthropic, you will play a crucial role in maintaining and enhancing the dependability of our AI systems, particularly focusing on Claude, our flagship large language model. The AI Reliability Engineering (AIRE) team collaborates across various departments to improve the robustness and resilience of our critical serving pathways—from SDKs and network layers to API infrastructure and hardware accelerators. Your work will involve designing and implementing systems that ensure high availability and low latency, managing incident responses, and supporting infrastructure that underpins our AI safety commitments. This position offers a unique opportunity to influence the reliability of cutting-edge AI systems at a company committed to safety and societal benefit, providing a dynamic and cross-disciplinary environment that values holistic system thinking.
Qualifications
- Bachelor’s degree or equivalent in a relevant field such as Computer Science, Engineering, or related disciplines
- Strong background in distributed systems, infrastructure, or reliability engineering
- Experience in operating large-scale model serving or training infrastructure, preferably with over 1000 GPUs
- Knowledge of ML hardware accelerators including GPUs, TPUs, or Trainium
- Understanding of ML-specific networking optimizations like RDMA and InfiniBand
- Familiarity with AI-specific observability tools and frameworks
- Experience with chaos engineering and resilience testing methodologies
- Contributions to open-source infrastructure or ML tooling are advantageous
- Excellent communication and collaboration skills with the ability to build strong cross-team relationships
- Demonstrated ownership and user-centric approach to system reliability
Searching for Development & Programming roles that provide visa sponsorship? Connect with international employers through Development & Programming Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
- Develop and define Service Level Objectives (SLOs) for large language model serving systems, balancing availability, latency, and development velocity
- Design, implement, and maintain monitoring and observability systems across the token processing pipeline
- Assist in designing and deploying high-availability serving infrastructure across multiple regions and cloud providers
- Lead incident response efforts for critical AI services, ensuring rapid resolution, conducting thorough incident reviews, and implementing systematic improvements
- Support the reliability and safety of safeguard model serving, ensuring alignment with safety commitments and operational excellence
- Collaborate with cross-functional teams to identify system vulnerabilities and implement resilience enhancements
- Contribute to the development of best practices for system reliability, scalability, and safety
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
Interested in opportunities specifically in United Kingdom? Discover our dedicated Visa Sponsorship Jobs in United Kingdom page featuring roles from top employers in this location.
- Competitive annual salary ranging from £325,000 to £390,000 GBP
- Comprehensive health and wellness benefits
- Opportunities for professional growth and development in a pioneering AI environment
- Flexible hybrid work policy with a minimum of 25% in-office presence
- Visa sponsorship available for eligible candidates
- Collaborative and innovative work culture focused on societal impact and safety
Anthropic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, ethnicity, gender, sexual orientation, age, disability, or any other protected characteristic. We believe that diverse perspectives and backgrounds foster innovation and excellence, and we actively encourage candidates from all backgrounds to apply.
Similar Jobs
Explore other opportunities that match your interests