# noCV engineering task library

Content version 5

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.

## BTEST — Flaky test containment and repair

A fictional web team reruns red builds until they pass. Shared clocks, leaked state, and unawaited work make failures hard to trust.

**Field:** Quality engineering. **Suggested stack:** TypeScript, Vitest, Playwright.

**Engineer value:** Practice controlled diagnosis, isolation, and test reliability measurement.

**Company value:** Recover useful failure signals and reduce blind reruns without hiding product defects.

**Delivery agreement:** Deliver repaired tests, deterministic reproductions, and a bounded quarantine policy.

### Setup prerequisites

- Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

### Find failure causes

Capture reproducible conditions and classify failures.

#### BTEST-101 — Record first-attempt results separately from rerun results

**Task · Medium priority · Foundational**

noCV practice brief v5 · BTEST-101 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Find failure causes. Depends on: No preceding ticket.

Difficulty: Foundational. Estimated focused work: 60 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 100%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

A green rerun erases the original failure from the summary.

Acceptance criteria

- Persist attempt number and original outcome.

- Link retries to the same test identity.

- Show initial failure counts separately from final status.

Implementation constraints

- Exclude screenshots or payloads containing real user data.

Verification

- Record fail-then-pass history.

- Ensure duplicate reporter delivery does not double-count attempts.

Deliverables

- Attempt-aware result report.

Rollout and recovery: Introduce reporting before changing retry policy; retain original raw fixture results.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BTEST-102 — Reproduce order-dependent failures with a saved shuffle seed

**Task · High priority · Intermediate**

noCV practice brief v5 · BTEST-102 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Find failure causes. Depends on: BTEST-101.

Difficulty: Intermediate. Estimated focused work: 120 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 100%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

One test fails only after a preference-setting test runs first.

Acceptance criteria

- Shuffle tests using a recorded seed.

- Replay the identical ordering from that seed.

- Emit a minimal command including environment assumptions.

Implementation constraints

- Keep the experiment inside the synthetic suite.

Verification

- Replay a known failing order.

- Verify a different seed and missing seed are reported accurately.

Deliverables

- Seeded order runner.

Rollout and recovery: Use in diagnostic jobs; preserve the normal suite order until the defect is fixed.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

### Remove nondeterminism

Repair isolation, time, and asynchronous behavior.

#### BTEST-103 — Trace leaked database rows between test cases

**Bug · High priority · Advanced**

noCV practice brief v5 · BTEST-103 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Remove nondeterminism. Depends on: BTEST-102.

Difficulty: Advanced. Estimated focused work: 180 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 60% · Database engineering 40%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

An integration test reads records inserted by an earlier test and passes for the wrong reason.

Acceptance criteria

- Assign isolated fixture ownership to each test.

- Clean only rows owned by that test.

- Fail when a test observes another fixture namespace.

Implementation constraints

- Do not use an unrestricted table truncation against shared databases.

Verification

- Run the tests independently and in reversed order.

- Inject foreign fixture rows and verify isolation.

Deliverables

- Fixture isolation repair.

Rollout and recovery: Adopt per-suite first; preserve failed fixture IDs for diagnosis without row contents.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BTEST-104 — Replace wall-clock races with explicit clock control

**Bug · High priority · Intermediate**

noCV practice brief v5 · BTEST-104 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Remove nondeterminism. Depends on: BTEST-102.

Difficulty: Intermediate. Estimated focused work: 120 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 100%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

A trial-expiry test fails around UTC midnight and daylight-saving changes.

Acceptance criteria

- Inject a clock into expiry calculations.

- Cover before, at, and after expiry boundaries.

- Restore real timers after each test.

Implementation constraints

- Business timestamps remain UTC; local zones are test inputs.

Verification

- Replay midnight and offset-boundary cases.

- Verify leaked fake timers are detected by teardown.

Deliverables

- Clock-based expiry tests.

Rollout and recovery: Land with the date behavior unchanged except for the documented boundary correction.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BTEST-105 — Remove readiness sleeps from browser tests

**Bug · Medium priority · Intermediate**

noCV practice brief v5 · BTEST-105 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Remove nondeterminism. Depends on: BTEST-101.

Difficulty: Intermediate. Estimated focused work: 150 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 100%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

A fixed delay is either too short on CI or unnecessarily long locally.

Acceptance criteria

- Wait for a specific observable ready state.

- Bound the wait and report the missing condition.

- Handle a failed request without waiting indefinitely.

Implementation constraints

- Do not increase global timeouts to mask failures.

Verification

- Run with fast and delayed local responses.

- Return an error and verify a useful bounded failure.

Deliverables

- Condition-based browser waits.

Rollout and recovery: Replace sleeps incrementally; keep traces for failures during adoption.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BTEST-106 — Find async work that escapes test teardown

**Bug · High priority · Advanced**

noCV practice brief v5 · BTEST-106 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Remove nondeterminism. Depends on: BTEST-103, BTEST-104.

Difficulty: Advanced. Estimated focused work: 210 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 80% · Developer tooling 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Background polling continues after a test completes and changes the next test's state.

Acceptance criteria

- Track owned timers, subscriptions, and requests.

- Cancel owned work during teardown.

- Fail the test on unhandled late rejections.

Implementation constraints

- Do not suppress process-level rejection reporting.

Verification

- Complete a polling test with clean teardown.

- Inject an ignored cancellation and detect the leaked work.

Deliverables

- Async lifecycle repair.

Rollout and recovery: Enable leak detection for the affected suite before expanding coverage.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

### Keep signals useful

Make retries and quarantine visible and temporary.

#### BTEST-107 — Separate product defects from environmental test failures

**Task · High priority · Advanced**

noCV practice brief v5 · BTEST-107 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Keep signals useful. Depends on: BTEST-101, BTEST-106.

Difficulty: Advanced. Estimated focused work: 180 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 70% · Site reliability 30%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

An unavailable local dependency and an incorrect application response currently look identical.

Acceptance criteria

- Define explicit failure categories with supporting observations.

- Keep unknown causes marked unknown.

- Prevent infrastructure classification from silently turning failures green.

Implementation constraints

- Classification rules must be auditable and versioned.

Verification

- Classify a refused connection and an assertion mismatch.

- Verify ambiguous evidence remains unclassified.

Deliverables

- Failure triage rules.

Rollout and recovery: Use categories for routing; retain the underlying failed status.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BTEST-108 — Add expiring quarantine entries with accountable owners

**Task · High priority · Intermediate**

noCV practice brief v5 · BTEST-108 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Keep signals useful. Depends on: BTEST-107.

Difficulty: Intermediate. Estimated focused work: 150 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 80% · Platform engineering 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Quarantined tests have accumulated without owners or a return date.

Acceptance criteria

- Require owner, reason, linked reproduction, and expiry.

- Continue executing quarantined tests in a visible lane.

- Fail policy checks on expired or unmatched entries.

Implementation constraints

- Quarantine cannot remove coverage silently.

Verification

- Accept a complete short-lived entry.

- Reject expired, wildcard-only, and ownerless entries.

Deliverables

- Quarantine policy validator.

Rollout and recovery: Start with a reviewed inventory; restore tests automatically only after policy conditions are met.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BTEST-109 — Evaluate flake repairs with controlled repeated runs

**Task · High priority · Expert**

noCV practice brief v5 · BTEST-109 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Keep signals useful. Depends on: BTEST-103, BTEST-104, BTEST-105, BTEST-106.

Difficulty: Expert. Estimated focused work: 300 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 80% · Site reliability 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Twenty green reruns are being presented as proof that a defect is gone.

Acceptance criteria

- Declare seeds, repetitions, environment, and remaining uncertainty.

- Compare repaired and deliberately unfixed cases under matching conditions.

- Report first-attempt failures and confidence limits without universal reliability claims.

Implementation constraints

- Keep the run budget bounded and retain failed seeds.

Verification

- Recover the injected failure in the unfixed variant.

- Verify repaired runs and explicitly report any inconclusive result.

Deliverables

- Repair evaluation report.

Rollout and recovery: Remove quarantine only with the reproduction fixed and policy review complete.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### BTEST-110 — Add a contributor guide for a trustworthy regression test

**Chore · Low priority · Foundational**

noCV practice brief v5 · BTEST-110 · Flaky test containment and repair

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Keep signals useful. Depends on: BTEST-108, BTEST-109.

Difficulty: Foundational. Estimated focused work: 60 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Quality engineering 100%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

New contributors copy retry-heavy tests and perpetuate the same failure patterns.

Acceptance criteria

- Show one deterministic fixture, clock, and cleanup example.

- State when a rerun is diagnostic rather than acceptance.

- Document how to submit a reproducible failure.

Implementation constraints

- Use actual repaired examples from this project.

Verification

- Follow the guide to add a passing boundary test.

- Check that an intentionally leaked resource fails review checks.

Deliverables

- Contributor testing guide.

Rollout and recovery: Link the guide from quarantine failures and the test command.

Project prerequisites: Create a small local application with three deliberately flaky tests using synthetic data and local endpoints.

Engineer value: Practice controlled diagnosis, isolation, and test reliability measurement.

Company value: Recover useful failure signals and reduce blind reruns without hiding product defects.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.
