Guide · Reliability
What Breaks When Traffic Spikes in SaaS?
Traffic spikes break things. Database connections, Lambda, memory, rate limits. Here's what to fix before you go viral.
Capacity cliffs are Performance & SLO failures—load test and budget the tail before you go viral. Primary control: Performance & SLO
Launch day traffic finds every soft limit
Viral newsletter hits and bot storms break apps that "worked fine" at Tuesday load—DB pools, Lambda concurrency, memory, or your own rate limits choking good users. One team spiked 50× after a feature; the pool emptied in minutes with no pooling discipline, no concurrency plan, and no prior load test. They repaired under fire and lost users who bounced during the outage.
Prepare the bottleneck before the spike
Pool DB connections and set ceilings (RDS Proxy / PgBouncer). Reserve Lambda concurrency for critical paths; know account limits. Watch memory and scale before OOM. Prefer per-user rate limits over a global knife that cuts VIPs. Run k6/Artillery against realistic peaks once—better than learning live.
Performance & continuity under load
APRF Performance & SLO and Reliability treat capacity and recovery as gates. Uptime alerts still matter: spikes without pages are silent failures until Twitter notices.
Next: Performance & SLO
Open the related pillar specification for mandatory checks, artifacts, and pass conditions. Self-attest is optional.
Related
Frequently asked questions
- What breaks when SaaS traffic spikes?
- Database connection pools, Lambda concurrency, memory, and rate limits. Connection pooling, reserved concurrency, and load testing help prevent failures.
- How do I prepare for traffic spikes?
- Set up connection pooling, configure Lambda concurrency, monitor memory. Load test before launch. Set up alerts for high resource usage.
- What is database connection pooling?
- Connection pooling reuses database connections instead of creating new ones per request. It prevents connection exhaustion under load. Use RDS Proxy or PgBouncer.