Keep one slow room client from exhausting gateway memory
A paused browser cannot consume events, but its send buffer grows throughout a synthetic incident drill. Add bounded per-connection delivery and a recoverable overload path.
- Focused work estimate
- 3h 30m + prerequisites
- Priority in the scenario
- High
- Engineering practice
- Backpressure · Resource limits · Load testing
Estimated field mix
- Performance engineering50%
- Real-time systems50%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
The fictional Harbor operations team coordinates synthetic incidents in a shared room. Engineers post status notes and claim response tasks while connections come and go. Scope is one application gateway, a durable room event store, and fixture clients; alert paging and external chat delivery are excluded.
Setup prerequisites
- Room membership and command contracts
- Synthetic incident events and a controllable transport harness
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- LIVE-101 · Show the difference between connected and caught up
- LIVE-102 · Reject malformed room events before they reach the reducer
- LIVE-103 · Authorize room subscription before replaying its history
- LIVE-104 · Stop duplicate and out-of-order notes confusing the timeline
- LIVE-105 · Recover when a reconnect cursor is older than retained history
- LIVE-106 · Reconcile a task claim when its acknowledgement is lost
- LIVE-107 · Expire disconnected presence without rewriting incident history
- LIVE-108 · Cut off room delivery when membership is revoked mid-replay
Acceptance criteria
- Each connection has explicit queued-byte and queued-event limits.
- Overloaded clients receive a safe resync/close outcome when possible; durable events are never silently skipped while claiming currency.
- Healthy clients continue receiving events and per-connection resources are reclaimed after closure.
Implementation constraints
- Keep a bounded load fixture; distinguish disposable presence updates from durable timeline events.
Verification to include
- Pause one consumer while a second reads normally; measure bounded memory and uninterrupted healthy delivery.
- Exceed limits, reconnect the slow client, and recover by cursor/snapshot with no silently missing timeline events.
Deliverables
- Backpressure policy and bounded slow-consumer load report
Rollout and recovery
Canary conservative queue limits with disconnect metrics; reduce admission or disable live updates if memory bounds fail.
Value of the work
For the engineer: Practice event identity, ordering, replay, authorization, and bounded backpressure in one understandable system.
For the team: Inspect how an engineer keeps collaborative state trustworthy during failures without relying on optimistic success messages.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.