Rehearse a secret-provider outage and document recovery limits
The provider is unavailable during a planned rotation. On call needs to know which reads can continue, which writes must wait, and how to recover without pasting keys into environment files.
- Focused work estimate
- 3h + prerequisites
- Priority in the scenario
- Medium
- Engineering practice
- Security incident drills · Provider resilience · Recovery documentation
Estimated field mix
- Security50%
- Site reliability50%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional supplier integration signs incoming webhooks and uses an outbound API credential. Operators currently replace environment values by hand. Build with a deterministic secret-store adapter and fabricated keys only; no live provider account or production credential is part of the exercise.
Setup prerequisites
- Cryptographic hash APIs
- HTTP webhook handling
- Access control
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- VAULT-101 · Inventory credential consumers without exporting their values
- VAULT-102 · Move secret lookup behind a version-aware provider
- VAULT-103 · Redact credentials from failed supplier requests
- VAULT-104 · Model credential activation and retirement as explicit transitions
- VAULT-105 · Verify signed webhooks against raw bytes before parsing
- VAULT-106 · Accept old and new webhook signatures only during a bounded overlap
- VAULT-107 · Keep duplicate signed callbacks from creating duplicate supplier events
- VAULT-108 · Rotate outbound credentials without retrying an ambiguous write twice
- VAULT-109 · Make emergency revocation reach cached consumers
Acceptance criteria
- Exercise provider outage before activation, during overlap, and after old-version revocation.
- Document allowed cached behavior and fail-closed conditions for each stage.
- Restore processing through the provider while preserving operation IDs and audit history.
Implementation constraints
- Use fabricated credentials and deterministic failures; no manual secret-value fallback is allowed.
Verification to include
- Recover a staged fixture rotation after provider availability returns.
- Keep the revoked version unusable throughout outage and recovery, and inspect logs for secret absence.
Deliverables
- Credential incident drill report and stage-specific recovery runbook
Rollout and recovery
Version the runbook with provider and policy versions; repeat the drill when cache or overlap behavior changes.
Value of the work
For the engineer: Practice credential lifecycle design, overlap windows, authenticated webhooks, and failure recovery without handling real secrets.
For the team: Review whether an engineer can make rotation auditable and fail closed while preserving availability and replay safety.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.