Evaluate flake repairs with controlled repeated runs
Twenty green reruns are being presented as proof that a defect is gone.
- Focused work estimate
- 5h + prerequisites
- Priority in the scenario
- High
- Engineering practice
- Experimental design · Reliability analysis
Estimated field mix
- Quality engineering80%
- Site reliability20%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional web team reruns red builds until they pass. Shared clocks, leaked state, and unawaited work make failures hard to trust.
Setup prerequisites
- Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- BTEST-101 · Record first-attempt results separately from rerun results
- BTEST-102 · Reproduce order-dependent failures with a saved shuffle seed
- BTEST-103 · Trace leaked database rows between test cases
- BTEST-104 · Replace wall-clock races with explicit clock control
- BTEST-105 · Remove readiness sleeps from browser tests
- BTEST-106 · Find async work that escapes test teardown
Acceptance criteria
- Declare seeds, repetitions, environment, and remaining uncertainty.
- Compare repaired and deliberately unfixed cases under matching conditions.
- Report first-attempt failures and confidence limits without universal reliability claims.
Implementation constraints
- Keep the run budget bounded and retain failed seeds.
Verification to include
- Recover the injected failure in the unfixed variant.
- Verify repaired runs and explicitly report any inconclusive result.
Deliverables
- Repair evaluation report.
Rollout and recovery
Remove quarantine only with the reproduction fixed and policy review complete.
Value of the work
For the engineer: Practice controlled diagnosis, isolation, and test reliability measurement.
For the team: Recover useful failure signals and reduce blind reruns without hiding product defects.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.