Restart crashed workers without retry storms
A corrupt fixture crashes every replacement worker, and immediate restart loops consume the supervisor CPU.
- Focused work estimate
- 4h + prerequisites
- Priority in the scenario
- High
- Engineering practice
- Supervision · Backoff · Failure isolation
Estimated field mix
- Systems programming70%
- Site reliability30%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Pattern topics
- Circuit BreakerApply
Open a bounded unavailable state after repeated worker crashes so replacement attempts stop until the declared recovery probe.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional document converter launches local sandbox substitutes as child processes. Its prototype parses newline-delimited output, leaks descriptors, and retries requests after ambiguous worker exits. Create a Rust supervisor and deterministic worker fixtures; no document content or production sandbox is supplied.
Setup prerequisites
- File descriptors
- Framing
- Process signals
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- SIPC-101 · Frame worker messages without treating newlines as boundaries
- SIPC-103 · Correlate out-of-order worker replies to the right caller
- SIPC-102 · Close inherited descriptors before executing the worker
- SIPC-104 · Separate protocol output from worker diagnostics
- SIPC-105 · Apply backpressure when a worker stops reading
- SIPC-106 · Classify worker exit before deciding whether to retry
Acceptance criteria
- Restart budget uses bounded attempts and backoff
- Stable runtime resets the failure window
- Exhaustion opens a visible unavailable state
Implementation constraints
- Use injected time and deterministic jitter; do not sleep in tests.
Verification to include
- Crash twice, recover, and verify the budget clears after stability.
- Crash every replacement and confirm restart attempts stop at the bound.
Deliverables
- Restart policy and crash-loop test
Rollout and recovery
Fail new requests fast while a worker class is unavailable.
Value of the work
For the engineer: Practice IPC contracts, resource inheritance, process supervision and ambiguous completion.
For the team: Review a local worker boundary that contains crashes and preserves request outcomes before connecting an isolated execution provider.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.