Rehearse backlog recovery with a constrained worker budget
Adding workers may increase database contention and make recovery slower.
- Focused work estimate
- 5h + prerequisites
- Priority in the scenario
- High
- Engineering practice
- Performance experiments · Incident response
Estimated field mix
- Performance engineering60%
- Site reliability40%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional document service has a rising processing backlog. Operators cannot tell whether arrivals, worker failures, or a slow dependency is responsible.
Setup prerequisites
- Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- BINCIDENT-101 · Report oldest eligible job age alongside queue depth
- BINCIDENT-102 · Separate arrival rate from successful service rate
- BINCIDENT-104 · Pause selected producers while preserving accepted work
- BINCIDENT-103 · Identify poison jobs without starving unrelated work
- BINCIDENT-105 · Recover expired worker leases without concurrent duplicate effects
- BINCIDENT-106 · Tune retry delay for a throttled dependency
- BINCIDENT-107 · Estimate backlog drain time under explicit capacity assumptions
Acceptance criteria
- Compare two bounded concurrency settings under identical arrivals.
- Measure completion age, dependency pressure, and duplicate-effect checks.
- Choose a recovery setting with stated tradeoffs.
Implementation constraints
- Record machine limits; do not extrapolate local capacity to production.
Verification to include
- Drain the synthetic incident workload.
- Demonstrate a setting that worsens contention or violates limits.
Deliverables
- Recovery experiment report.
Rollout and recovery
Raise concurrency in measured steps; revert when dependency or integrity thresholds fail.
Value of the work
For the engineer: Practice incident timelines, queue diagnosis, and measured recovery.
For the team: Produce a repeatable response that preserves work integrity and exposes customer delay.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.