noCV
QUEUE-109 · Make recovery routine

Replay a failed preview through an audited repair command

Practice briefTaskIntermediate

Operators repair failed jobs by deleting Redis keys. That loses failure history and can replay work for the wrong tenant when a copied ID is mistaken.

Focused work estimate
2h 30m + prerequisites
Priority in the scenario
High
Engineering practice
Operational APIs · Audit trails · Least privilege

Estimated field mix

  • Platform engineering40%
  • Backend30%
  • Security30%

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Your next step

Review it, then add it to your workspace.

The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.

Project context

A fictional internal document portal creates previews through a trusted mock renderer. Large batches crowd out small teams, failed documents retry forever, and operators lack a safe replay command. This exercise never executes uploaded code or real document macros.

Setup prerequisites

  • Queue semantics
  • Database transactions
  • Operational metrics

Preceding work

Complete these dependencies, or supply their agreed outputs before taking this ticket.

Acceptance criteria

  • Allow an authorized operator to create a linked repair attempt for a terminal failure.
  • Require tenant scope, expected job revision, and a reason.
  • Preserve original failure history and make duplicate repair requests idempotent.

Implementation constraints

  • Do not reopen the failed record in place or expose a general queue-management console.

Verification to include

  • Repair one terminal fixture failure and follow the linked attempts.
  • Reject a foreign tenant, stale revision, and repair of a successful job.

Deliverables

  • Scoped repair command and audit projection

Rollout and recovery

Grant repair permission to a synthetic operator role; revoke command access while retaining read-only audit history.

Value of the work

For the engineer: Practice asynchronous lifecycle control, fair scheduling, bounded retries, and artifact recovery with a deterministic provider.

For the team: See how an engineer accounts for accepted work and reduces operational toil without granting operators broad data access.

Evidence boundaries

Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.