Site Reliability Engineer II [Performance - Chaos]
As a Site Reliability chaos Engineer, your role is to provide reliability engineering services through chaos engineering, and performance engineering techniques. Using monitoring, fault-injection, and performance tools, you will deliver detailed feedback to product owners and development teams. You will collaborate with cross-functional teams to design, build, automate, and maintain scalable, resilient infrastructure. Your responsibilities will include ensuring high availability, monitoring system performance, executing chaos experiments, and aiding support staff with resolving incidents. This role requires a strong background in scripting, cloud platforms, and a passion for optimizing operational efficiency and system resilience. You will use Site Reliability Engineering chaos practices to deliver a seamless user experience.
Responsibilities:
- Design and develop custom, automated fault-injection playbooks to test system recovery mechanisms under pressure.
- Execute end-to-end chaos experiments (such as network latency, container crashes, and region failovers) across non-production environments.
- Analyze infrastructure bottlenecks and application dependencies exposed during simulated system failures.
- Draft comprehensive post-mortem reports and present findings to software development teams to guide code-hardening efforts.
- Construct custom dashboards in Grafana and Prometheus to visualize real-time system degradation during chaos testing.
- Implement automated chaos tests directly into active CI/CD deployment pipelines.
- Perform extensive performance and load testing using JMeter to baseline system stability before injecting faults.
- Configure advanced Dynatrace alerting policies to ensure simulated outages are immediately and accurately detected by monitoring tools.
- Simulate full Disaster Recovery (DR) and business continuity scenarios to verify the reliability of backup and failover procedures.
- Partner with software engineers to refactor application code for better fault tolerance and self-healing capabilities.
- Collaborate with the broader operations team to define system boundaries and minimize risk when executing testing windows.
- Build clean, modular automation scripts in Python, Shell, or JavaScript to provision test environments and schedule recurring chaos runs.
“I love working here because Quest has been my second family and second home. I've experienced a wholesome work environment, and good management.”
- Quest Employee
- Software Engineer I - CMS Hyderabad, India 07/24/2026
- Lead Software Engineer - Site Reliability Engineering Hyderabad, India 07/24/2026
- Site Reliability Engineer II [Performance - Chaos] Hyderabad, India 07/24/2026
No jobs have been saved.
No jobs have been saved.
Quest Diagnostics is an equal employment opportunity employer. Our policy is to recruit, hire and promote qualified individuals without regard to race, color, religion, sex, age, national origin, disability, veteran status, sexual orientation, gender identity, or any any other legally protected status . Quest Diagnostics observes minimum age requirements established by federal, state and/or local laws, and will ask an applicant for verification when deemed necessary.
Quest Diagnostics is committed to working with and providing reasonable accommodations to individuals with disabilities. If you need a reasonable accommodation because of a disability for any part of the employment process, please complete the accommodation request form.