Keep origin traffic bounded when the cache becomes unavailable
A cache outage redirects the full request rate to the origin. The supposed fallback turns a small infrastructure fault into a wider outage.
- Focused work estimate
- 4h + prerequisites
- Priority in the scenario
- High
- Engineering practice
- Fault isolation · Backpressure
Estimated field mix
- Site reliability60%
- Performance engineering40%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional equipment-rental service caches availability summaries. A campaign sends repeated reads, while stock updates and shared expiry times create bursts against the origin. Build a local origin stub and cache-backed read API using synthetic depots and products.
Setup prerequisites
- Create a deterministic local availability origin and Redis-backed reader with synthetic tenant, depot and product data.
- Use a seeded hot-key distribution and controlled time; record cache capacity, TTLs, runtime and machine limits.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- PCACHE-101 · Specify which availability responses may be reused
- PCACHE-102 · Measure saved origin work instead of celebrating the hit ratio
- PCACHE-103 · Create a hot-key workload that exposes synchronized expiry
- PCACHE-104 · Coalesce simultaneous availability misses for one key
- PCACHE-106 · Serve stale availability only within an explicit degraded-read policy
Acceptance criteria
- Set cache-client timeouts and bound origin fallback concurrency and queue length.
- Return an explicit overload or unavailable response when the fallback budget is exhausted.
- Recover cache use without a synchronized refill storm or continuing to queue expired callers.
Implementation constraints
- Exercise cache disconnect, slow response and recovery separately; do not rely only on a clean process shutdown.
Verification to include
- Run the hot-key load while making cache operations hang and assert bounded connection and origin work counts.
- Restore the cache during overload and verify queue drainage, timeout accounting and correct tenant-specific responses.
Deliverables
- Degraded-cache policy, failure injection and recovery traces
Rollout and recovery
Canary the bounded fallback configuration locally; retain a switch that fails informational reads explicitly if fallback threatens the origin budget.
Value of the work
For the engineer: Learn to evaluate caching through avoided work, bounded staleness, concurrency and recovery rather than hit rate alone.
For the team: Produce a reviewable cache policy and failure exercise for a read-heavy service with changing business data.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.