Recover coordination after a database connection partition
An instance loses database connectivity but can still contact the fake work provider, raising the risk of acting on expired authority.
- Focused work estimate
- 5h + prerequisites
- Priority in the scenario
- Medium
- Engineering practice
- Network partitions · Recovery
Estimated field mix
- Distributed systems60%
- Networking20%
- Site reliability20%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional document service runs periodic retention planning on multiple application instances. Paused instances resume after lease expiry and can overlap newer schedulers.
Setup prerequisites
- Create a synthetic maintenance queue and two independent scheduler clients.
- Use fake time where possible and no destructive real retention actions.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- ALEASE-101 · Specify scheduler lease state with a separate authority epoch
- ALEASE-102 · Use one authoritative clock boundary for lease decisions
- ALEASE-103 · Classify maintenance actions that require fencing before side effects
- ALEASE-104 · Acquire a scheduler lease atomically under competing instances
- ALEASE-105 · Renew a scheduler lease only for the current holder and epoch
- ALEASE-106 · Fence maintenance writes after a paused scheduler resumes
- ALEASE-107 · Make scheduled maintenance occurrences idempotent across leadership changes
- ALEASE-108 · Expose lease health without implying the holder is making progress
Acceptance criteria
- Stop new protected actions when authority cannot be verified.
- Reconcile lease and run state after reconnection.
- Resume only under a current valid epoch.
Implementation constraints
- Use controlled local adapter disconnection, not real network disruption.
Verification to include
- Disconnect one owner, acquire with another, and continue valid work.
- Reconnect the old owner and reject its cached authority before any action.
Deliverables
- Partition drill and authority recovery trace
Rollout and recovery
Run before multi-instance scheduling; fail closed on unresolved ownership.
Value of the work
For the engineer: Practice lease assumptions, fencing and stale-owner rejection.
For the team: Review safe coordination that remains explainable under pauses and recovery.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.