Skip to content

Find and eliminate unused AWS resources with Cloud Zombie Hunter → Try it

CDOps Tech logo - Cloud and DevOps consulting services.
  • Services
    • Cloud Engineering & Architecture (The Foundation)
    • Platform Engineering & IDP (The Forge)
    • Cloud Security & Compliance (The Shield)
    • Fractional SRE & Interim DevOps (The “Air Cover” Wedge)
    • All Services
  • Pricing
    • Cloud Foundation
    • Managed Cloud Services
    • Site Reliability Engineering
    • MLOps
    • Cloud Cost Optimization Audit
    • Cloud & DevOps Roadmap
    • CI/CD Pipelines
    • Containerization
  • Resources
    • Blog
    • Case Studies
    • Newsroom
  • About Us
    • Company
    • Careers
  • Contact
CDOps Tech Logo
CONSULT AN EXPERT
Guide, Insights

Chaos engineering: is your infrastructure as resilient as you think?

Simarpreet S Chandhok

•

September 28, 2026

Every team says they run blameless postmortems. Few actually do. The moment someone feels like they're defending themselves, they start editing the story, and you lose the exact details that would have prevented the next incident. Here's what separates a genuinely blameless process from one that just borrows the name.
Share This Post :
Facebook
Twitter
LinkedIn

Most teams find out how resilient their infrastructure actually is during an incident, not before one. By then it’s expensive to learn.

Chaos engineering exists to close that gap. It’s the discipline of deliberately injecting failure into a system while it’s still healthy, so you find the weak points on your own schedule instead of the internet’s. Netflix built the first well-known version of this, Chaos Monkey, back in 2011, to kill random production instances and confirm the system kept serving traffic anyway. The idea has since become a standard practice at companies running distributed infrastructure at scale.

The uncomfortable part: most infrastructure that looks resilient on paper has never actually been tested. Redundant nodes, multi-AZ deployments, autoscaling groups, and failover databases are all resilience claims. Chaos engineering is how you verify them.

Why "we have redundancy" isn't the same as "we're resilient"

A few public postmortems make the gap between design and reality concrete.

On July 2, 2019, Cloudflare went down globally for 27 minutes because a single new WAF rule contained a regular expression prone to catastrophic backtracking. The CPU exhaustion caused by this single rule spiked to 100% on every CPU core handling HTTP and HTTPS traffic across Cloudflare’s worldwide network. The rule passed code review and functional tests. Nobody had tested what happened to the fleet under pathological input at global rollout scale, because nothing in the deployment pipeline was designed to surface that failure mode before it hit every edge server at once.

On December 7, 2021, AWS’s us-east-1 region went down for hours, taking Netflix, Slack, Tinder, and a long list of other services with it. The root cause traced back to problems within the region’s internal control plane, the system responsible for managing and orchestrating network traffic, and the resulting congestion cascaded into EC2, DynamoDB, and console access failures. Teams that had architected for AZ-level redundancy discovered that a control-plane failure doesn’t respect availability zone boundaries the way EC2 instance failures do.

Neither company was careless. Both had invested heavily in reliability engineering. What they hadn’t done, in the specific failure paths that caused these incidents, was inject the failure ahead of time and watch what actually happened. That’s the gap chaos engineering is built to close: the distance between “we designed for this” and “we’ve confirmed this holds under real failure conditions.”

What chaos engineering actually tests

Chaos engineering is often reduced to “randomly kill servers and see what breaks.” That’s Chaos Monkey’s original, narrow use case, and it undersells what the practice covers today. A mature chaos engineering program tests failure across several layers:

  • Infrastructure failure: terminating instances, killing containers, simulating AZ or region outages
  • Network failure: injecting latency, packet loss, DNS failures, or partitioning services from each other
  • Dependency failure: taking down a downstream API, database, or third-party service your system depends on
  • Resource exhaustion: forcing CPU, memory, or disk pressure to see how the system degrades
  • State and data failure: corrupting or delaying data, testing what happens when a cache goes stale or a queue backs up

The common thread across all of these is that you’re not testing whether a component can fail. You already know it can. You’re testing whether the system behaves the way you assumed it would when that component fails, and whether the humans on call can actually diagnose and respond to it.

The method: hypothesis first, blast radius second

A chaos experiment isn’t “break something and see what happens.” It follows a structure, borrowed from the scientific method, that turns it into a controlled test rather than a stunt:

  1. Define steady state. Pick a measurable signal for “the system is healthy” — request success rate, latency percentile, throughput. This is your baseline.
  2. Form a hypothesis. State what you expect to happen: “If the primary database read replica fails, request latency should stay under 200ms because the query router fails over within 5 seconds.”
  3. Inject the failure, with a defined blast radius. Start small. Target one instance, one availability zone, or a single percentage of traffic before running the same experiment against a wider surface. This is the step teams skip when they’re in a hurry, and it’s the one that turns a useful experiment into a self-inflicted outage.
  4. Observe and compare. Did steady state hold? If not, where did the assumption break — the failover logic, the alerting, the on-call runbook, the retry behavior of upstream services?
  5. Fix, then repeat. The output of a chaos experiment is a remediation item, not just a report. Fix the gap, then re-run the experiment to confirm the fix actually holds.

Teams running this on Kubernetes typically use tools like LitmusChaos or Chaos Mesh to inject pod failures, network delays, and resource pressure directly into the cluster. Teams on AWS have AWS Fault Injection Simulator as a managed option for the same purpose. Gremlin remains a common commercial platform for running these experiments across a mixed infrastructure estate. The tool matters less than the discipline of running the experiment with a hypothesis and a bounded blast radius.

Common misconceptions

Most engineering teams don’t have a full-time reliability function, and chaos engineering tends to fall off the roadmap for exactly that reason: it takes bandwidth that’s already spoken for by feature work. That’s usually where a fractional or interim SRE engagement earns its keep — running a structured chaos program, building the blast-radius controls and rollback plans around it, and handing the team a resilience baseline they can maintain without full-time reliability headcount. CDOps Tech’s Air Cover offering is built for exactly that gap.

For teams running Kubernetes specifically, failure injection is most useful when it’s built into how the cluster is managed day to day, not bolted on as a one-off exercise. That’s part of what a managed Kubernetes engagement should include: pod disruption budgets, network policies, and resource limits that are actually load-tested against failure, not just configured and left alone.

Frequently Asked Questions

Is chaos engineering the same as penetration testing?
No. Penetration testing looks for security vulnerabilities an attacker could exploit. Chaos engineering looks for resilience gaps that would surface under operational failure, whether that’s a dead instance, a slow dependency, or a full AZ outage. The methods overlap (both are controlled, deliberate attacks on your own system) but the goals are different.

How often should we run chaos experiments?
Continuously is the end state for mature programs, often as part of the CI/CD pipeline so every deploy is tested against known failure modes automatically. Teams starting out typically run scheduled GameDay sessions monthly or quarterly, then increase frequency as tooling and confidence mature.

What’s the difference between chaos engineering and disaster recovery testing?
DR testing usually validates a specific, planned scenario — full region failover, backup restoration — on a known schedule. Chaos engineering is broader and more continuous: it tests smaller, more frequent failure modes as an ongoing practice, and DR testing can be thought of as one large-blast-radius chaos experiment within that broader practice.

Do we need to test in production to get value from this?
No. Most of the value comes from finding gaps in staging, where the cost of a bad hypothesis is a failed test, not a customer-facing outage. Production testing adds confidence once you’ve already fixed what staging surfaced, and once you have the observability and rollback tooling to contain a real surprise.

What’s the first experiment a team new to this should run?
Pick one dependency your system treats as “always available” and kill it in staging. A downstream API, a cache layer, a queue. Watch whether your system degrades the way you assumed it would, or whether it falls over in a way nobody expected. That single experiment usually surfaces enough gaps to build a backlog from.

The Bottom Line

Resilience that hasn’t been tested is a hypothesis, not a fact. The Cloudflare and AWS incidents above didn’t happen because those companies were negligent about reliability. They happened because a specific failure path went untested until it happened live, in front of every customer at once. Chaos engineering is how you find that path first.

If you want a second set of eyes on where your own infrastructure’s untested assumptions are likely sitting, that’s a conversation worth having before an incident forces it. Book a discovery call, or see how this plays out for teams like yours in our case studies.

Run postmortems that actually change something

Still seeing the same incidents resurface? CDOps Tech embeds fractional SRE and interim DevOps expertise into your team to close the reliability gaps your postmortems keep surfacing.

GET STARTED
Share This Post :
Facebook
Twitter
LinkedIn

Navigation

Got Questions About Your Cloud Strategy?

Don’t hesitate to reach out. Our cloud and DevOps experts are here to help you navigate everything from migration to optimization.
CONTACT US NOW

Recommended Reading

Blameless Postmortem

The Blameless Postmortem: How to Actually Run One

The Blameless Postmortem: How to Actually Run One
September 18, 2026
When to Hire Your First DevOps Engineer vs. When to Outsource
September 11, 2026
DevOps vs. Platform Engineering: What’s the Actual Difference?
September 1, 2026
cdops tech contact

Thinking about outsourcing your tech operations?

Get in touch and discover how working with CDOps Tech gives your business an edge with top-tier engineers and cloud experts – ready to support DevOps, Cloud, Security, AI, SRE, and more from leading global talent hubs. Fill out the form to get started.

Countries Served
0
Support Coverage
20 /7
Core Service Areas
0 +
Technologies & Tools
0 +
CDOps Tech Logo

Transforming businesses through cutting-edge cloud infrastructure and seamless DevOps automation

Useful Links
  • About Us
  • Pricing
  • Contact
  • Case Studies
  • Blogs
  • Privacy Policy
Solutions
  • Fractional SRE & Interim DevOps (The “Air Cover” Wedge)
  • Cloud Engineering & Architecture (The Foundation)
  • Platform Engineering & IDP (The Velocity)
  • Cloud Security & Compliance (The Shield)
Contact Information

Feel free to contact & reach us !!

  • contact@cdops.tech
  • +65 60288048​

CDOps Tech Singapore

  • #14-04 SBF Center, 160 Robinson Road, Singapore (068914)

CDOps Tech India

  • 117/L/188 Naveen Nagar, Kakadeo, Kanpur, Uttar Pradesh, India
Linkedin Instagram Facebook

Copyright © 2026 CDOps Tech.  All rights reserved.