Replay a failed preview through an audited repair command
Operators repair failed jobs by deleting Redis keys. That loses failure history and can replay work for the wrong tenant when a copied ID is mistaken.
- Focused work estimate
- 2h 30m + prerequisites
- Priority in the scenario
- High
- Engineering practice
- Operational APIs · Audit trails · Least privilege
Estimated field mix
- Platform engineering40%
- Backend30%
- Security30%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional internal document portal creates previews through a trusted mock renderer. Large batches crowd out small teams, failed documents retry forever, and operators lack a safe replay command. This exercise never executes uploaded code or real document macros.
Setup prerequisites
- Queue semantics
- Database transactions
- Operational metrics
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- QUEUE-101 · Expose a document job's current phase and terminal reason
- QUEUE-102 · Reject oversized document batches before accepting work
- QUEUE-103 · Commit job acceptance and dispatch intent together
- QUEUE-104 · Stop retrying unsupported document formats
- QUEUE-105 · Keep one team's bulk import from occupying every worker
- QUEUE-106 · Fence a late worker after its rendering lease expires
Acceptance criteria
- Allow an authorized operator to create a linked repair attempt for a terminal failure.
- Require tenant scope, expected job revision, and a reason.
- Preserve original failure history and make duplicate repair requests idempotent.
Implementation constraints
- Do not reopen the failed record in place or expose a general queue-management console.
Verification to include
- Repair one terminal fixture failure and follow the linked attempts.
- Reject a foreign tenant, stale revision, and repair of a successful job.
Deliverables
- Scoped repair command and audit projection
Rollout and recovery
Grant repair permission to a synthetic operator role; revoke command access while retaining read-only audit history.
Value of the work
For the engineer: Practice asynchronous lifecycle control, fair scheduling, bounded retries, and artifact recovery with a deterministic provider.
For the team: See how an engineer accounts for accepted work and reduces operational toil without granting operators broad data access.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.