Exercise configuration service loss during a rolling deployment
The configuration endpoint becomes unavailable while old and new workers are starting with different cached snapshots.
- Focused work estimate
- 5h + prerequisites
- Priority in the scenario
- Medium
- Engineering practice
- Fault injection · Operational readiness
Estimated field mix
- Platform engineering50%
- Site reliability30%
- Quality engineering20%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional internal reporting platform changes feature and timeout settings through environment edits. Partial rollouts leave API and worker processes interpreting different values.
Setup prerequisites
- Create local API and worker configuration consumers.
- Use fabricated settings without secrets.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- ACONFIG-101 · Inventory runtime settings with owners and restart requirements
- ACONFIG-102 · Reject invalid timeout combinations as one configuration unit
- ACONFIG-103 · Define an immutable configuration snapshot envelope
- ACONFIG-104 · Atomically adopt a validated configuration snapshot in each process
- ACONFIG-105 · Publish configuration only against the revision reviewed by the operator
- ACONFIG-106 · Keep incompatible consumers on their last valid configuration
- ACONFIG-107 · Stage a timeout revision for a named consumer cohort
- ACONFIG-108 · Report configuration staleness separately from application health
- ACONFIG-109 · Rollback configuration by selecting a prior immutable snapshot
Acceptance criteria
- Describe startup and running-process behavior separately.
- Prove critical settings cannot start from an unvalidated cache.
- Record revision selection and recovery for both consumer versions.
Implementation constraints
- Use local fault injection; no production configuration changes.
Verification to include
- Disconnect the endpoint after valid adoption and retain the declared bounded behavior.
- Start with a corrupt cache during outage and reject startup safely.
Deliverables
- Outage drill and compatibility trace
Rollout and recovery
Run before rollout; halt the deployment if a consumer cannot demonstrate its selected revision.
Value of the work
For the engineer: Practice configuration contracts, compatibility and failure recovery.
For the team: Review controlled configuration delivery with inspectable blast radius.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.