# noCV engineering task library

Content version 5

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.

## PRECOVER — Recover interrupted returns without refunding twice

A fictional equipment retailer lets customers choose a refund or replacement after inspection. Its payment and stock providers can time out after accepting an operation, so retrying the whole return is unsafe. Create a local TypeScript application, PostgreSQL state/outbox tables, and synthetic provider doubles with controllable outcomes; no starter code or fixtures are supplied. Use invented orders and integer minor-unit amounts only, with no real payments or external provider calls.

**Field:** Distributed systems. **Suggested stack:** TypeScript, Node.js, PostgreSQL, Vitest.

**Engineer value:** Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

**Company value:** Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

**Delivery agreement:** Ten tickets over command intake, provider recovery, and operational rehearsal. Build controllable synthetic dependencies for an individual issue or deliver the complete local returns workflow with failure drills.

### Setup prerequisites

- SQL transactions

- Idempotent commands

- Async failure handling

- State modeling

### Accept one durable intent

Make return decisions, commands, and pending work consistent before calling providers.

#### PRECOVER-101 — Reject return decisions that skip inspection or reverse a completed refund

**Bug · High priority · Foundational**

noCV practice brief v5 · PRECOVER-101 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Accept one durable intent. Depends on: No preceding ticket.

Difficulty: Foundational. Estimated focused work: 90 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Backend 50% · System design 30% · Security 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: State (compare).

State — Compare: Compare guarded state objects with an explicit transition table for preventing illegal return actions without requiring a class for every status.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

The return endpoint accepts a caller-supplied status string. A replacement can be requested before inspection, and a stale browser tab can move a refunded return back to awaiting-inspection.

Acceptance criteria

- Define explicit commands and permitted transitions for awaiting inspection, approved, rejected, processing, completed, and needs-review states.

- Require an expected revision and authorized merchant scope for each decision; terminal outcomes cannot be overwritten by a stale command.

- Return stable invalid-transition and revision-conflict errors while preserving an append-only transition history.

Implementation constraints

- Compare a transition table with state objects; choose the representation that makes allowed transitions and guards easiest to inspect.

Verification

- Exercise the approved-refund and rejected-inspection paths and compare their recorded transition order.

- Attempt a pre-inspection replacement, a post-refund reversal, and a cross-merchant decision; verify no state or history mutation.

Deliverables

- Return transition contract and invalid-transition regression cases

Rollout and recovery: Route local return decisions through the explicit transition API before adding provider effects; retain transition history when reverting the UI.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PRECOVER-102 — Bind a repeated refund request to its original merchant and intent

**Bug · High priority · Intermediate**

noCV practice brief v5 · PRECOVER-102 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Accept one durable intent. Depends on: PRECOVER-101.

Difficulty: Intermediate. Estimated focused work: 150 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Distributed systems 50% · Backend 30% · Database engineering 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: Command (apply).

Command — Apply: Capture refund intent as an immutable durable command so retries refer to the same operation and cannot quietly change its merchant, amount, or target.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

The browser retries an approved 4,500-minor-unit refund after a lost response. The current deduplication key is global, and changing the amount while reusing the key overwrites the queued request.

Acceptance criteria

- Persist an immutable refund command with merchant, return, amount, currency, and a canonical intent hash scoped to its request key.

- Identical retries return the existing command identity and current status; changed intent under the same scoped key returns a conflict.

- Concurrent identical requests create one accepted command, and command history never rewrites the original amount or target return.

Implementation constraints

- Represent the operation as a durable command rather than relying on an in-memory callback or HTTP response cache; amounts use integer minor units.

Verification

- Submit the same intent concurrently and retry after a simulated lost response; compare one command identity and one accepted history entry.

- Reuse the key with a new amount and from another merchant; assert conflict for changed scoped intent and independent authorized identities across merchants.

Deliverables

- Durable refund command contract and scoped idempotency reproduction

Rollout and recovery: Enable durable command intake before dispatching synthetic refunds; disable new intake without deleting accepted command identities.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PRECOVER-103 — Commit approved return work and its dispatch record together

**Bug · High priority · Advanced**

noCV practice brief v5 · PRECOVER-103 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Accept one durable intent. Depends on: PRECOVER-102.

Difficulty: Advanced. Estimated focused work: 210 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Distributed systems 60% · Database engineering 40%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: Transactional Outbox (apply).

Transactional Outbox — Apply: Bind state and dispatch intent to one commit while accepting repeat delivery and preserving deterministic identities for downstream deduplication.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

The application marks a return processing, then publishes a refund job. If it crashes between those steps, the return is stuck forever; publishing first can instead process a refund for a rolled-back decision.

Acceptance criteria

- Commit the accepted command, return transition, and outbox record in one database transaction with a deterministic dispatch identity.

- A bounded dispatcher claims due records safely across two instances and retries using the same operation identity.

- Treat dispatch as at-least-once: acknowledgment loss can redeliver, but downstream handling deduplicates the effect and pending records remain observable.

Implementation constraints

- Do not hold the database transaction open during a provider call; use explicit lease expiry and attempt limits for dispatch recovery.

Verification

- Crash before and after commit and compare return, command, and outbox records for atomic presence or absence.

- Lose a dispatch acknowledgment and run two dispatchers through lease expiry; assert one logical refund intent despite repeat delivery.

Deliverables

- Transactional outbox write path, dispatcher, and crash-boundary checks

Rollout and recovery: Start with dispatch paused and reconcile synthetic pending counts; enable one dispatcher before rehearsing a second instance and restart.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

### Handle partial and uncertain outcomes

Coordinate effects and compensation while containing independent provider failures.

#### PRECOVER-104 — Persist replacement progress across stock reservation and shipment creation

**Story · High priority · Advanced**

noCV practice brief v5 · PRECOVER-104 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Handle partial and uncertain outcomes. Depends on: PRECOVER-103.

Difficulty: Advanced. Estimated focused work: 240 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Distributed systems 60% · Integrations 20% · Backend 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: Saga (apply) · State (apply).

Saga — Apply: Coordinate durable replacement steps and fallible compensation so a partial provider success can be resumed or reconciled without repeating earlier effects.

State — Apply: Represent uncertain provider outcomes separately from confirmed failures so workflow transitions cannot invent a successful rollback or completed refund.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

A replacement reserves the last unit, then shipment creation fails. Retrying from the beginning makes another reservation, while immediately issuing a refund would ignore stock still held for the customer.

Acceptance criteria

- Persist separate reservation and shipment steps with stable provider operation keys and explicit pending, confirmed, failed, and uncertain outcomes.

- Resume from recorded progress after restart; a confirmed reservation is reused instead of repeated when shipment creation is retried.

- On a confirmed shipment failure, schedule release of the reservation before offering the declared refund fallback; uncertain effects remain unresolved until checked.

Implementation constraints

- Use a durable orchestrated workflow in the local application; compensation is another fallible operation, not a database rollback of an external effect.

Verification

- Reserve stock, crash before shipment confirmation, and resume; assert one reservation lineage and the correct next step.

- Return a confirmed shipment rejection and then a reservation-release timeout; verify the workflow remains visible and does not silently refund or reserve again.

Deliverables

- Replacement step model and reservation/shipment recovery cases

Rollout and recovery: Enable the replacement path for one synthetic order type; pause new workflows while allowing existing recorded steps to reconcile on rollback.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PRECOVER-105 — Open the refund circuit without declaring timed-out refunds failed

**Bug · High priority · Advanced**

noCV practice brief v5 · PRECOVER-105 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Handle partial and uncertain outcomes. Depends on: PRECOVER-103, PRECOVER-104.

Difficulty: Advanced. Estimated focused work: 210 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Distributed systems 50% · Site reliability 30% · Integrations 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: Circuit Breaker (apply).

Circuit Breaker — Apply: Bound attempts during provider failure while keeping circuit health separate from each refund's durable and potentially uncertain business outcome.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

The synthetic refund provider begins timing out. The proposed circuit breaker maps every timeout to failed and releases queued retries when it closes, even though some timed-out refunds were accepted by the provider.

Acceptance criteria

- Define closed, open, and half-open behavior with a controllable clock, a bounded failure window, and a limited number of half-open probes.

- Opening the circuit defers new provider attempts without converting existing uncertain refund outcomes into confirmed failures.

- Reconcile ambiguous operation identities through the provider's status lookup before resubmission; a reopened circuit preserves pending work and retry timing.

Implementation constraints

- Use provider idempotency and status lookup contracts in addition to the breaker; a circuit breaker alone cannot prevent duplicate money movement.

Verification

- Advance a fake clock through the failure threshold, open window, and half-open success/failure paths and assert bounded probe calls.

- Simulate provider acceptance followed by timeout, then close the circuit; verify reconciliation finds the original refund and no second refund is created.

Deliverables

- Refund circuit policy, uncertain-outcome reconciliation, and clock-driven cases

Rollout and recovery: Observe the synthetic failure window before enabling deferral; disabling the breaker must not bypass uncertain-outcome reconciliation.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PRECOVER-106 — Keep stalled stock calls from consuming every refund execution slot

**Bug · High priority · Intermediate**

noCV practice brief v5 · PRECOVER-106 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Handle partial and uncertain outcomes. Depends on: PRECOVER-104, PRECOVER-105.

Difficulty: Intermediate. Estimated focused work: 180 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Site reliability 40% · Distributed systems 40% · Performance engineering 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: Bulkhead (apply).

Bulkhead — Apply: Separate provider execution capacity so a stalled stock dependency cannot consume refund slots, with bounded queues and explicit permit cleanup.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

Stock reservations and refunds share a 12-slot provider pool. Twelve hanging stock requests prevent an otherwise healthy refund provider from receiving work, and the pending promise list grows without a limit.

Acceptance criteria

- Give stock and refund operations independent concurrency limits and bounded waiting queues with explicit admission outcomes.

- Timeout, cancellation, and provider rejection each release a slot exactly once while preserving durable work for the declared retry policy.

- When stock capacity is saturated, refund work continues within its configured limit and neither queue exceeds its cap.

Implementation constraints

- Document what is isolated by the pools and what remains shared, including database capacity; do not claim full fault isolation from separate counters alone.

Verification

- Hold all stock operations behind a barrier and complete a refund batch without exceeding either provider's concurrency limit.

- Cancel queued work, time out running calls, and overfill both queues; assert no leaked permits or unbounded pending promises.

Deliverables

- Provider capacity limits, admission contract, and saturation reproduction

Rollout and recovery: Start with conservative synthetic pool limits and expose queue depth/rejection counts; pause intake before lowering limits below active work.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PRECOVER-107 — Stop compensation retries when a replacement has already shipped

**Bug · High priority · Expert**

noCV practice brief v5 · PRECOVER-107 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Handle partial and uncertain outcomes. Depends on: PRECOVER-104, PRECOVER-105, PRECOVER-106.

Difficulty: Expert. Estimated focused work: 300 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Distributed systems 60% · System design 20% · Backend 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: Saga (refactor).

Saga — Refactor: Make compensation conditional on current durable and provider-confirmed facts, including irreversible progress and a visible unresolved state.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

A shipment-create response was lost, so the return entered compensation. A later status lookup confirms the replacement shipped; the queued compensation still plans to release stock and refund the customer.

Acceptance criteria

- Revalidate confirmed external progress and the workflow revision before each compensation step; a stale plan cannot execute against newer shipment state.

- Represent irreversible shipment completion and conflicting provider facts explicitly, moving unresolved cases to needs-review with a bounded reason trail.

- An authorized resolution records its cited synthetic provider references and chosen next action without deleting earlier uncertainty or issuing an unapproved second benefit.

Implementation constraints

- Compensation must follow the declared return policy and current confirmed facts; never model an external shipment as a reversible local transaction.

Verification

- Queue compensation, then reveal a confirmed shipment before its next step; assert the stale compensation is stopped and no refund is issued.

- Race two resolution attempts and return contradictory provider status responses; verify one revision wins and unresolved contradictions remain visible.

Deliverables

- Compensation guards, conflict-resolution contract, and late-confirmation race cases

Rollout and recovery: Enable automatic compensation only for confirmed reversible states; keep conflict resolution explicit and retain a pause control for all pending compensation.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

### Rehearse bounded recovery

Make operator replay, restart behavior, and overload limits inspectable.

#### PRECOVER-108 — Preview the exact return command before an operator requests replay

**Story · Medium priority · Foundational**

noCV practice brief v5 · PRECOVER-108 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Rehearse bounded recovery. Depends on: PRECOVER-102, PRECOVER-107.

Difficulty: Foundational. Estimated focused work: 120 minutes; setup and prerequisite tickets are additional.

Estimated field mix: API design 40% · Backend 40% · Security 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: Command (apply).

Command — Apply: Expose durable command intent for safe preview and explicit replay while keeping the original operation immutable and rechecking eligibility at execution.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

Support can currently click Retry beside a return without seeing whether it will check a refund status, retry a reservation, or attempt compensation. The button also lets a stale page replay a command that already completed.

Acceptance criteria

- Show the immutable command identity, target return, intended next operation, current revision, and the reason replay is eligible or blocked.

- A preview is read-only; replay requires an explicit authorized command with expected revision and a fresh eligibility check.

- Reusing the replay request key returns the existing replay result, while completed or unresolved-ineligible operations cannot be forced through this path.

Implementation constraints

- Show only synthetic operational identifiers and bounded explanations; replay cannot edit the original refund amount or target operation.

Verification

- Preview one eligible status-check command and one completed command; confirm the preview performs no provider calls or writes.

- Complete a command after preview, then request replay with its stale revision; also try an unauthorized merchant and assert rejection.

Deliverables

- Replay preview and command endpoint contracts with stale/denied cases

Rollout and recovery: Expose preview before enabling the replay action; disable replay writes independently while keeping history and explanations available.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PRECOVER-109 — Rehearse returns recovery at every commit and provider acknowledgment boundary

**Task · High priority · Expert**

noCV practice brief v5 · PRECOVER-109 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Rehearse bounded recovery. Depends on: PRECOVER-103, PRECOVER-107, PRECOVER-108.

Difficulty: Expert. Estimated focused work: 330 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Distributed systems 40% · Quality engineering 40% · Site reliability 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: Transactional Outbox (apply) · Saga (apply).

Transactional Outbox — Apply: Verify the outbox's durable intent and repeat-delivery behavior at actual restart boundaries instead of assuming publication and acknowledgment are atomic.

Saga — Apply: Exercise workflow resumption and compensation after partial external effects, accepting explicit unresolved states when provider facts cannot be established.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

The happy-path demo succeeds, but no one has restarted the application between a provider accepting an operation and the database recording its result. A green HTTP test does not establish whether accepted refunds survive recovery correctly.

Acceptance criteria

- Create a deterministic crash matrix covering command commit, outbox claim, provider acceptance, acknowledgment persistence, and compensation progress.

- After each restart, reconcile durable commands, provider operations, and return state; no confirmed effect disappears and no logical refund is applied twice.

- Distinguish eventual completion under a recovered provider from explicit unresolved state when the provider remains unavailable or contradictory.

Implementation constraints

- Use process restarts and durable local database state, not only thrown exceptions inside one transaction; reset synthetic provider state deliberately between cases.

Verification

- Run every crash point twice with the same seeded workload and compare final command/effect counts and unresolved reason codes.

- Keep status lookup unavailable after an acceptance timeout and verify the drill stops at a visible unresolved state rather than manufacturing completion.

Deliverables

- Restart harness, crash-boundary matrix, and reconciled synthetic effect report

Rollout and recovery: Run the full drill before enabling a new workflow version; pause new intake and drain or reconcile recorded work before reverting code.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PRECOVER-110 — Tune recovery capacity without creating a retry surge when a provider returns

**Task · Medium priority · Advanced**

noCV practice brief v5 · PRECOVER-110 · Recover interrupted returns without refunding twice

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Rehearse bounded recovery. Depends on: PRECOVER-105, PRECOVER-106, PRECOVER-109.

Difficulty: Advanced. Estimated focused work: 240 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Performance engineering 40% · Site reliability 40% · Distributed systems 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics: Bulkhead (refactor) · Circuit Breaker (refactor).

Bulkhead — Refactor: Extend dependency capacity isolation with fair new/recovery admission and an explicit shared database budget so bounded pools do not hide another bottleneck.

Circuit Breaker — Refactor: Coordinate half-open probes and reopening with durable retry scheduling so a healthy transition does not release the entire outage backlog at once.

Pattern topics identify design choices to practice. Read the ticket's acceptance criteria and justify the simplest suitable approach. Tags are not capability or ownership evidence; an untagged ticket has no curated pattern topic assigned.

A ten-minute synthetic outage leaves 2,000 deferred returns. When the circuit closes, every due retry enters the refund pool together, crowding out new work and exhausting the shared database connection budget.

Acceptance criteria

- Define bounded admission and fair scheduling between new and recovery work, with retry jitter generated from a reproducible test seed.

- Coordinate half-open probes, provider pool limits, and the shared database budget so reopening cannot enqueue or execute the entire backlog at once.

- Report backlog age, queue occupancy, rejection/defer counts, and completion time for the declared local load; justify selected limits and remaining shared bottlenecks.

Implementation constraints

- No universal throughput target is assumed; state machine resources, workload, and acceptable service objectives before comparing capacity settings.

Verification

- Recover a seeded 2,000-return backlog while submitting new work and verify bounded concurrency, queue sizes, and progress for both classes.

- Make the provider fail again during recovery and cancel waiting work; assert permits are released, retries remain durable, and no tight retry loop forms.

Deliverables

- Recovery scheduling policy, reproducible saturation run, and capacity tradeoff report

Rollout and recovery: Enable recovery scheduling with conservative limits and a pause control; reduce admission before shrinking active pools and preserve pending command identities.

Project prerequisites: SQL transactions Idempotent commands Async failure handling State modeling

Engineer value: Practice durable workflow state, ambiguous provider outcomes, compensation, concurrency isolation, and recovery that survives process restarts.

Company value: Inspect whether a proposed workflow preserves refund and inventory invariants, contains provider failures, and gives operators a bounded recovery path.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.
