noCV
PPROVIDER-108 · Contain provider failure

Pause a failing provider and admit one recovery probe

Practice briefStoryAdvanced

Provider B returns transport failures for every synthetic request. Callers continue spending the full timeout on each attempt even though the service needs a short recovery window.

Focused work estimate
3h 30m + prerequisites
Priority in the scenario
High
Engineering practice
Circuit breakers · State transitions · Deterministic timing

Estimated field mix

  • Site reliability50%
  • Integrations30%
  • Distributed systems20%

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Pattern topics

  • Circuit BreakerApply

    Stop repeated attempts to a failing provider with explicit failure classification and a single controlled recovery probe.

Your next step

Review it, then add it to your workspace.

The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.

Project context

A fictional maintenance scheduler sends appointment notices through two providers. Their request shapes, error codes, and acknowledgement semantics differ. Build two local scripted provider doubles and a TypeScript application module; use synthetic recipients, block external network access, and never send real notifications. No provider accounts, starter repository, or production delivery qualification is supplied.

Setup prerequisites

  • HTTP contracts
  • Asynchronous cancellation
  • Dependency injection

Preceding work

Complete these dependencies, or supply their agreed outputs before taking this ticket.

Acceptance criteria

  • Open the circuit after three consecutive declared transport failures, reject further attempts locally for thirty simulated seconds, and then admit one recovery probe.
  • A successful probe closes the circuit; a failed probe reopens it. Validation errors and provider business rejections do not count as transport failures.
  • Track circuits independently per provider, return a distinct circuit-open outcome, and never report rejected local attempts as provider calls.

Implementation constraints

  • Use an injected monotonic clock and explicit transition logic; no real-time sleeps, external service, or automatic cross-provider failover is required.

Verification to include

  • Advance the fake clock through closed, open, and half-open states and assert allowed call counts.
  • Race several requests at the recovery boundary and confirm exactly one probe; verify A remains usable while B is open.

Deliverables

  • Circuit Breaker state transitions and deterministic recovery tests

Rollout and recovery

Enable the breaker around B's local stub first, observe transitions, and preserve unknown-operation records when resetting circuit state.

Value of the work

For the engineer: Practice placing provider boundaries, comparing wrappers and coordinators, and testing ordering, cancellation, and resource limits under controlled faults.

For the team: Review whether a provider change can be made without rewriting scheduling rules, leaking recipient data, or turning one provider failure into wider exhaustion.

Evidence boundaries

Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.