# Load Testing, Capacity Planning and Chaos Engineering — Scalability, Availability & Reliability

Source: https://www.skillbyai.com/en/scalability/o-testing

> Find limits before users do and verify that failure handling works.

## Testing the system, not just the code

**Load testing** applies realistic traffic to find capacity and bottlenecks: a **load test** at expected peak, a **stress test** beyond it to see how the system fails, a **soak test** for hours to catch leaks, and a **spike test** for sudden bursts. Use production-like data and environments, model real user journeys, and watch latency percentiles and saturation, not just averages. **Capacity planning** turns results and growth forecasts into a plan: current peak, expected growth, headroom target and the next bottleneck to fix. **Chaos engineering** runs controlled experiments that inject failures (kill instances, add latency, block a dependency, fail a zone) to verify that redundancy and resilience patterns actually work. Start with a hypothesis, a small blast radius and an abort switch, in staging first; teams then run **game days** where they practise failures and incident response together.

## A k6 load test with thresholds

The test fails if p95 latency or the error rate exceeds the targets.

```js
import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  stages: [
    { duration: '5m', target: 200 },   // ramp up
    { duration: '20m', target: 200 },  // hold at expected peak
    { duration: '5m', target: 400 },   // push past peak
    { duration: '5m', target: 0 },
  ],
  thresholds: {
    http_req_duration: ['p(95)<300'],
    http_req_failed: ['rate<0.001'],
  },
};

export default function () {
  const res = http.get('https://staging.shop.example.com/api/products?page=1');
  check(res, { 'status is 200': (r) => r.status === 200 });
  sleep(1);
}
```

## Find the bottleneck, then the next one

Fixing one bottleneck reveals the next: first the app servers, then the database connection pool, then a lock in one table. Capacity work is iterative; record each limit you find and the change that moved it.

**Quiz:** What is the purpose of a chaos engineering experiment?

- [ ] To randomly break production with no plan
- [ ] To replace monitoring
- [x] To verify, with a hypothesis and limited blast radius, that the system tolerates a specific failure
- [ ] To measure developer productivity

*Answer:* To verify, with a hypothesis and limited blast radius, that the system tolerates a specific failure. Chaos experiments test resilience assumptions deliberately and safely.
