# noCV engineering task library

Content version 5

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.

## BINCIDENT — Queue backlog incident response

A fictional document service has a rising processing backlog. Operators cannot tell whether arrivals, worker failures, or a slow dependency is responsible.

**Field:** Site reliability. **Suggested stack:** TypeScript, BullMQ, Redis.

**Engineer value:** Practice incident timelines, queue diagnosis, and measured recovery.

**Company value:** Produce a repeatable response that preserves work integrity and exposes customer delay.

**Delivery agreement:** Deliver local incident instrumentation, recovery controls, and a tabletop report.

### Setup prerequisites

- Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

### Observe backlog

Separate queue age, throughput, and failure symptoms.

#### BINCIDENT-101 — Report oldest eligible job age alongside queue depth

**Task · Medium priority · Foundational**

noCV practice brief v5 · BINCIDENT-101 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Observe backlog. Depends on: No preceding ticket.

Difficulty: Foundational. Estimated focused work: 60 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Site reliability 80% · Platform engineering 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

A shallow queue can still contain a single customer job stuck for hours.

Acceptance criteria

- Measure oldest eligible age separately from count.

- Exclude scheduled future work from overdue age.

- Show missing timestamp records explicitly.

Implementation constraints

- Use bounded job-type labels only.

Verification

- Measure a known delayed fixture.

- Keep future-scheduled jobs out of overdue calculations.

Deliverables

- Backlog age metric.

Rollout and recovery: Add read-only metrics before changing worker behavior.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BINCIDENT-102 — Separate arrival rate from successful service rate

**Task · Medium priority · Intermediate**

noCV practice brief v5 · BINCIDENT-102 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Observe backlog. Depends on: BINCIDENT-101.

Difficulty: Intermediate. Estimated focused work: 120 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Site reliability 60% · Performance engineering 40%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

The dashboard counts retries as new incoming customer work.

Acceptance criteria

- Track logical arrivals, attempts, and completions separately.

- Use consistent measurement windows.

- Expose worker concurrency with throughput.

Implementation constraints

- Do not infer capacity from a single peak sample.

Verification

- Reconcile a fixture with repeated attempts.

- Detect growing backlog despite high attempt throughput.

Deliverables

- Queue flow dashboard.

Rollout and recovery: Compare against synthetic queue inventory before relying on alerts.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

### Contain safely

Control arrivals and retries while preserving job identity.

#### BINCIDENT-103 — Identify poison jobs without starving unrelated work

**Bug · High priority · Advanced**

noCV practice brief v5 · BINCIDENT-103 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Contain safely. Depends on: BINCIDENT-102.

Difficulty: Advanced. Estimated focused work: 180 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Platform engineering 60% · Site reliability 40%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

One malformed document repeatedly consumes the first worker slot.

Acceptance criteria

- Bound attempts by failure category.

- Move exhausted jobs to an inspectable terminal holding state.

- Allow unrelated valid jobs to continue.

Implementation constraints

- Keep payload contents out of generic incident logs.

Verification

- Quarantine a deterministic poison job.

- Process a valid neighboring job and preserve its identity.

Deliverables

- Poison-job containment.

Rollout and recovery: Enable per job type; retain held work for reviewed correction.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BINCIDENT-104 — Pause selected producers while preserving accepted work

**Task · High priority · Advanced**

noCV practice brief v5 · BINCIDENT-104 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Contain safely. Depends on: BINCIDENT-102.

Difficulty: Advanced. Estimated focused work: 180 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Site reliability 50% · Platform engineering 30% · Security 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Operators stop every producer even though only one import path overloads the queue.

Acceptance criteria

- Scope pause controls to declared producer classes.

- Reject new work with explicit retry guidance.

- Preserve previously accepted job records.

Implementation constraints

- Privileged pause changes require an audit entry.

Verification

- Pause the synthetic import producer.

- Verify another producer works and unauthorized pause is denied.

Deliverables

- Scoped admission control.

Rollout and recovery: Start with one producer switch; restore intake gradually after backlog age recovers.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BINCIDENT-105 — Recover expired worker leases without concurrent duplicate effects

**Bug · High priority · Advanced**

noCV practice brief v5 · BINCIDENT-105 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Contain safely. Depends on: BINCIDENT-103.

Difficulty: Advanced. Estimated focused work: 210 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Distributed systems 50% · Platform engineering 30% · Site reliability 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

A worker crashes after producing output but before acknowledging completion.

Acceptance criteria

- Use stable job operation identity.

- Reclaim only expired ownership.

- Adopt existing completed output before repeating effects.

Implementation constraints

- Mock effects must be idempotent and locally scoped.

Verification

- Crash after effect acceptance and reclaim.

- Race two reclaimers and observe one terminal outcome.

Deliverables

- Lease recovery behavior.

Rollout and recovery: Reclaim in bounded batches; stop on conflicting output identities.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BINCIDENT-106 — Tune retry delay for a throttled dependency

**Task · High priority · Intermediate**

noCV practice brief v5 · BINCIDENT-106 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Contain safely. Depends on: BINCIDENT-103, BINCIDENT-104.

Difficulty: Intermediate. Estimated focused work: 150 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Site reliability 60% · Integrations 40%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Immediate retries amplify the dependency's temporary throttling.

Acceptance criteria

- Respect bounded retry-after guidance.

- Apply jitter with reproducible test control.

- Cap total attempts and elapsed time.

Implementation constraints

- Never retry permanent validation or authorization failures.

Verification

- Recover after a temporary throttle.

- Exhaust a persistent outage without a retry storm.

Deliverables

- Dependency retry policy.

Rollout and recovery: Trial on one job class; pause dispatch if throttling remains sustained.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

### Recover and learn

Drain work and validate recovery decisions.

#### BINCIDENT-107 — Estimate backlog drain time under explicit capacity assumptions

**Task · Medium priority · Advanced**

noCV practice brief v5 · BINCIDENT-107 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Recover and learn. Depends on: BINCIDENT-102, BINCIDENT-106.

Difficulty: Advanced. Estimated focused work: 180 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Performance engineering 60% · Site reliability 40%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Support promises a completion time using queue depth divided by the busiest worker's speed.

Acceptance criteria

- Use measured successful service and arrival rates.

- Report assumptions and uncertainty bounds.

- Return no finite estimate when arrivals meet or exceed service.

Implementation constraints

- Use synthetic measurements and avoid general capacity claims.

Verification

- Estimate a controlled draining workload.

- Show an unstable workload as non-draining.

Deliverables

- Drain-time estimator.

Rollout and recovery: Publish estimates with their observation window; retract stale estimates.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BINCIDENT-108 — Rehearse backlog recovery with a constrained worker budget

**Task · High priority · Expert**

noCV practice brief v5 · BINCIDENT-108 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Recover and learn. Depends on: BINCIDENT-104, BINCIDENT-105, BINCIDENT-107.

Difficulty: Expert. Estimated focused work: 300 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Performance engineering 60% · Site reliability 40%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Adding workers may increase database contention and make recovery slower.

Acceptance criteria

- Compare two bounded concurrency settings under identical arrivals.

- Measure completion age, dependency pressure, and duplicate-effect checks.

- Choose a recovery setting with stated tradeoffs.

Implementation constraints

- Record machine limits; do not extrapolate local capacity to production.

Verification

- Drain the synthetic incident workload.

- Demonstrate a setting that worsens contention or violates limits.

Deliverables

- Recovery experiment report.

Rollout and recovery: Raise concurrency in measured steps; revert when dependency or integrity thresholds fail.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BINCIDENT-109 — Create an incident timeline from control and queue events

**Story · Medium priority · Intermediate**

noCV practice brief v5 · BINCIDENT-109 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Recover and learn. Depends on: BINCIDENT-108.

Difficulty: Intermediate. Estimated focused work: 120 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Site reliability 80% · Data engineering 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

The post-incident discussion relies on memory of when pauses and retries changed.

Acceptance criteria

- Record control actions with actor and UTC time.

- Link changes to observed backlog metrics.

- Separate observations from causal hypotheses.

Implementation constraints

- Exclude customer payloads and personal operator commentary.

Verification

- Reconstruct the synthetic incident.

- Show a missing observation as a gap rather than inventing timing.

Deliverables

- Incident timeline exporter.

Rollout and recovery: Preserve original event records; append corrections to the timeline.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BINCIDENT-110 — Write a queue handoff with explicit resume conditions

**Chore · Low priority · Foundational**

noCV practice brief v5 · BINCIDENT-110 · Queue backlog incident response

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Recover and learn. Depends on: BINCIDENT-109.

Difficulty: Foundational. Estimated focused work: 60 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Site reliability 100%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

The next on-call engineer inherits paused intake without knowing when to reopen it.

Acceptance criteria

- List active controls and held job counts.

- Specify measurable intake-resume conditions.

- Document safe rollback of each temporary control.

Implementation constraints

- Commands target the local incident fixture only.

Verification

- Resume intake using the handoff.

- Keep intake paused when dependency health remains unknown.

Deliverables

- On-call handoff runbook.

Rollout and recovery: Require a handoff whenever containment outlives the current operator.

Project prerequisites: Create a bounded local queue simulation with synthetic jobs and controllable worker/dependency failures.

Engineer value: Practice incident timelines, queue diagnosis, and measured recovery.

Company value: Produce a repeatable response that preserves work integrity and exposes customer delay.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.
