Pause a failing provider and admit one recovery probe
Provider B returns transport failures for every synthetic request. Callers continue spending the full timeout on each attempt even though the service needs a short recovery window.
- Focused work estimate
- 3h 30m + prerequisites
- Priority in the scenario
- High
- Engineering practice
- Circuit breakers · State transitions · Deterministic timing
Estimated field mix
- Site reliability50%
- Integrations30%
- Distributed systems20%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Pattern topics
- Circuit BreakerApply
Stop repeated attempts to a failing provider with explicit failure classification and a single controlled recovery probe.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional maintenance scheduler sends appointment notices through two providers. Their request shapes, error codes, and acknowledgement semantics differ. Build two local scripted provider doubles and a TypeScript application module; use synthetic recipients, block external network access, and never send real notifications. No provider accounts, starter repository, or production delivery qualification is supplied.
Setup prerequisites
- HTTP contracts
- Asynchronous cancellation
- Dependency injection
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- PPROVIDER-101 · Normalize the second provider's acknowledgement without inventing delivery
- PPROVIDER-102 · Keep provider SDK types out of appointment rules
- PPROVIDER-105 · Record one bounded telemetry event around each provider attempt
- PPROVIDER-103 · Validate notice inputs before choosing a provider
- PPROVIDER-104 · Decouple notice format from the selected transport
- PPROVIDER-106 · Replace the scheduler's six-call notification sequence with one bounded operation
- PPROVIDER-107 · Remove the transparent send proxy that retries an ambiguous timeout
Acceptance criteria
- Open the circuit after three consecutive declared transport failures, reject further attempts locally for thirty simulated seconds, and then admit one recovery probe.
- A successful probe closes the circuit; a failed probe reopens it. Validation errors and provider business rejections do not count as transport failures.
- Track circuits independently per provider, return a distinct circuit-open outcome, and never report rejected local attempts as provider calls.
Implementation constraints
- Use an injected monotonic clock and explicit transition logic; no real-time sleeps, external service, or automatic cross-provider failover is required.
Verification to include
- Advance the fake clock through closed, open, and half-open states and assert allowed call counts.
- Race several requests at the recovery boundary and confirm exactly one probe; verify A remains usable while B is open.
Deliverables
- Circuit Breaker state transitions and deterministic recovery tests
Rollout and recovery
Enable the breaker around B's local stub first, observe transitions, and preserve unknown-operation records when resetting circuit state.
Value of the work
For the engineer: Practice placing provider boundaries, comparing wrappers and coordinators, and testing ordering, cancellation, and resource limits under controlled faults.
For the team: Review whether a provider change can be made without rewriting scheduling rules, leaking recipient data, or turning one provider failure into wider exhaustion.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.