# noCV engineering task library

Content version 5

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.

## PBATCH — Import a large supplier catalog without exhausting the worker

A fictional wholesaler imports supplier rows into a staging catalog. The current prototype reads the entire file into memory and restarts from zero after a failure. Build the prototype and generated CSV fixture locally before measuring improvements.

**Field:** Performance engineering. **Suggested stack:** TypeScript, Node.js streams, PostgreSQL, CSV.

**Engineer value:** Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

**Company value:** Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

**Delivery agreement:** Ten tickets in three phases. All data and failures are locally generated; throughput targets are relative to the declared baseline and do not imply a production service-level guarantee.

### Setup prerequisites

- Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size.

- Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

### Define input and resource behavior

Establish a representative file and honest end-to-end measurements.

#### PBATCH-101 — Generate a catalog file that includes difficult CSV boundaries

**Task · Medium priority · Foundational**

noCV practice brief v5 · PBATCH-101 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Define input and resource behavior. Depends on: No preceding ticket.

Difficulty: Foundational. Estimated focused work: 75 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Data engineering 50% · Quality engineering 30% · Performance engineering 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

The demonstration file has one short record per line. It cannot expose parsers that split quoted descriptions or allocate one enormous record.

Acceptance criteria

- Generate 250,000 deterministic rows including quoted delimiters, embedded newlines, multibyte text and declared invalid cases.

- Publish expected accepted/rejected counts and a digest of normalized accepted records.

- Declare a maximum record size and include an intentionally oversized record in a separate failure fixture.

Implementation constraints

- Use invented supplier identifiers and descriptions; fixture generation must not require a downloaded customer file.

Verification

- Regenerate with the same seed and compare file digest and expected counts.

- Parse with a reference implementation and verify quoted newlines do not change record boundaries.

Deliverables

- Fixture generator, manifest and expected normalized digest

Rollout and recovery: Version the fixture before profiling; keep the oversized fixture separate so the normal baseline has a defined completion result.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PBATCH-102 — Measure peak import memory across parsing, validation and writes

**Task · Medium priority · Foundational**

noCV practice brief v5 · PBATCH-102 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Define input and resource behavior. Depends on: PBATCH-101.

Difficulty: Foundational. Estimated focused work: 90 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Performance engineering 70% · Data engineering 30%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

The worker exits near the end of an import, but the existing memory sample is taken only after garbage collection and misses the peak.

Acceptance criteria

- Record process RSS, managed heap, buffered-record count and stage throughput at a fixed documented sampling interval.

- Report peak values and exit outcome, including resource-limit termination.

- Separate input-read completion from database-commit completion in elapsed time.

Implementation constraints

- Document sampling limitations and resource limits; do not report a killed run as a completed low-memory run.

Verification

- Run the whole-file baseline and retain its peak or termination outcome.

- Inject a slow writer and verify the timeline shows buffered work accumulating before completion.

Deliverables

- Resource timeline and baseline measurement command

Rollout and recovery: Keep measurement output outside imported data; retain failed-run artifacts for the streaming comparison.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PBATCH-103 — Attribute import time to parsing, validation and database waits

**Story · Medium priority · Intermediate**

noCV practice brief v5 · PBATCH-103 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Define input and resource behavior. Depends on: PBATCH-101, PBATCH-102.

Difficulty: Intermediate. Estimated focused work: 120 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Performance engineering 60% · Data engineering 40%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

A proposal to increase database concurrency assumes writes dominate, but expensive normalization may already saturate one CPU core.

Acceptance criteria

- Measure stage service time and time blocked on downstream capacity separately.

- Compare parse-only, parse-plus-validation and complete-import runs using the same input.

- Reconcile accepted and rejected records between stages and identify the supported bottleneck hypothesis.

Implementation constraints

- Use bounded aggregate timing rather than one log entry per record, which would alter the measured workload.

Verification

- Inject a known validation delay and show its contribution in the stage report.

- Inject writer delay instead and confirm it appears as downstream wait rather than parser CPU time.

Deliverables

- Stage comparison report and aggregate instrumentation

Rollout and recovery: Use the instrumentation in local benchmarks first; disable it independently if overhead makes comparisons unreliable.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

### Bound the import pipeline

Control buffering and write work without dropping or duplicating records.

#### PBATCH-104 — Stream records without buffering the rest of the catalog

**Story · High priority · Advanced**

noCV practice brief v5 · PBATCH-104 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Bound the import pipeline. Depends on: PBATCH-101, PBATCH-102, PBATCH-103.

Difficulty: Advanced. Estimated focused work: 240 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Performance engineering 50% · Data engineering 50%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Switching to a streaming file reader did not reduce memory because validation still accumulates every parsed row before writing starts.

Acceptance criteria

- Connect reading, parsing, validation and writing through bounded buffers with downstream backpressure.

- Reject an oversized record explicitly without allowing unbounded parser accumulation.

- Preserve normalized accepted-record digest and rejection reasons from the baseline fixture.

Implementation constraints

- Express buffer bounds in records and bytes where record sizes vary; a stream API alone does not establish bounded memory.

Verification

- Pause the writer and assert parser progress stops within the declared buffering bound.

- Split quoted multibyte records across small input chunks and compare complete results with the reference fixture.

Deliverables

- Streaming pipeline, buffer invariants and peak-memory comparison

Rollout and recovery: Keep the whole-file implementation only as a small-fixture reference; stop and retain the source file if the streaming result digest differs.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PBATCH-105 — Batch staging writes without changing duplicate-SKU behavior

**Story · High priority · Advanced**

noCV practice brief v5 · PBATCH-105 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Bound the import pipeline. Depends on: PBATCH-104.

Difficulty: Advanced. Estimated focused work: 210 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Database engineering 50% · Data engineering 30% · Performance engineering 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Single-row inserts dominate after streaming is introduced. A bulk-insert experiment is faster but resolves duplicate supplier SKUs differently.

Acceptance criteria

- Document duplicate-SKU ordering and preserve it across batch boundaries.

- Cap batch rows and serialized bytes, with bounded transaction duration.

- Return deterministic accepted/rejected outcomes when one batch includes invalid or conflicting records.

Implementation constraints

- Compare at least three bounded batch sizes on the fixed fixture; do not choose the largest solely from one fast run.

Verification

- Place duplicate SKUs on either side of a batch boundary and compare normalized results with the reference behavior.

- Inject a write failure midway through a batch and verify the declared atomicity and retry outcome.

Deliverables

- Batched writer, batch-size measurements and duplicate regressions

Rollout and recovery: Canary in staging with batch size configurable; reducing it must preserve semantics and allow work to continue from a valid checkpoint.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PBATCH-106 — Stop validation workers from creating an unbounded reorder queue

**Bug · High priority · Expert**

noCV practice brief v5 · PBATCH-106 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Bound the import pipeline. Depends on: PBATCH-103, PBATCH-104, PBATCH-105.

Difficulty: Expert. Estimated focused work: 330 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Performance engineering 50% · Data engineering 50%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Parallel validation improves throughput until one slow record delays output. Later completed records accumulate while the writer waits for order.

Acceptance criteria

- Define a bounded in-flight window and preserve required source ordering without retaining unlimited completed work.

- Propagate cancellation and fatal validation failure through every active stage.

- Select worker concurrency from repeated measurements that include serialization and coordination overhead.

Implementation constraints

- If the measured validation stage is not CPU-bound, document that result and retain a simpler bounded path instead of adding workers without benefit.

Verification

- Delay the first record while later records complete and assert both in-flight and retained-result bounds.

- Fail one worker during the import and verify controlled shutdown, deterministic checkpoint state and no silent record loss.

Deliverables

- Concurrency decision, bounded ordering implementation and fault tests

Rollout and recovery: Keep a single-worker configuration available; drain or cancel active work before changing concurrency during a rehearsal.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

### Recover and verify

Resume safely, report useful progress and gate throughput improvements.

#### PBATCH-107 — Report import progress from committed records instead of bytes read

**Bug · Medium priority · Intermediate**

noCV practice brief v5 · PBATCH-107 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Recover and verify. Depends on: PBATCH-104, PBATCH-105.

Difficulty: Intermediate. Estimated focused work: 120 minutes; setup and prerequisite tickets are additional.

Estimated field mix: API design 40% · Data engineering 40% · Frontend 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

The UI reaches 100% while thousands of records are still queued for the database. Operators assume they can close the import.

Acceptance criteria

- Expose bytes consumed, records accepted/rejected, committed records and terminal state as separate counters.

- Mark completion only after all accepted records are committed and final reconciliation succeeds.

- Avoid a remaining-time promise when throughput is unstable; report an unknown estimate explicitly.

Implementation constraints

- Keep progress updates throttled and monotonic within one attempt; resumed attempts must disclose their checkpoint origin.

Verification

- Pause the final write and verify the import remains nonterminal despite input exhaustion.

- Inject rejection and retry cases and reconcile counters with the final staging contents.

Deliverables

- Progress contract and end-of-input regression

Rollout and recovery: Introduce the progress fields before changing the UI completion rule; retain committed counters if the display is reverted.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PBATCH-108 — Resume an interrupted catalog import from a committed checkpoint

**Story · High priority · Expert**

noCV practice brief v5 · PBATCH-108 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Recover and verify. Depends on: PBATCH-104, PBATCH-105, PBATCH-107.

Difficulty: Expert. Estimated focused work: 360 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Data engineering 50% · Database engineering 30% · Distributed systems 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

An import loses its process after several committed batches. Restarting from a byte offset inside a quoted multiline record corrupts the next batch.

Acceptance criteria

- Bind checkpoints to source digest, parser configuration and import identity.

- Checkpoint only a valid record boundary whose writes are committed, with an explicit idempotency rule for uncertain completion.

- Reject a changed source or configuration and preserve the previous import for inspection.

Implementation constraints

- A raw newline is not a CSV record boundary. State whether resumption seeks to a validated boundary or replays and skips already committed record identities.

Verification

- Terminate after commit but before checkpoint acknowledgement, resume and compare the final digest with an uninterrupted import.

- Change one source byte and attempt resumption; no additional staging writes may occur.

Deliverables

- Checkpoint protocol, resume implementation and interruption matrix

Rollout and recovery: Enable resumption only after the failure matrix passes; retain the source and last valid checkpoint when recovery is refused.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PBATCH-109 — Prove temporary import resources disappear after cancellation

**Chore · Medium priority · Advanced**

noCV practice brief v5 · PBATCH-109 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Recover and verify. Depends on: PBATCH-104, PBATCH-106, PBATCH-108.

Difficulty: Advanced. Estimated focused work: 180 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Site reliability 40% · Data engineering 30% · Performance engineering 30%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Cancelled imports leave temporary files and open connections. A subsequent run inherits resource pressure and appears slower for an unrelated reason.

Acceptance criteria

- Define ownership and cleanup for temporary files, file descriptors, workers and database connections.

- Cancel safely at read, validation and commit stages while preserving the last valid recovery checkpoint.

- Make repeated cancellation safe and keep the source fixture intact.

Implementation constraints

- Check resource inventories before and after; avoid asserting cleanup from a completion message alone.

Verification

- Cancel at each controlled stage and verify owned resources return to the declared baseline.

- Run cancellation twice, then resume or start a fresh import and confirm correct results without restarting the environment.

Deliverables

- Cleanup implementation, resource inventory and cancellation regressions

Rollout and recovery: Run cleanup rehearsal before accepting performance results; quarantine uncertain staging state for explicit recovery rather than deleting it silently.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.

#### PBATCH-110 — Set a throughput gate that also enforces import memory and correctness

**Chore · High priority · Expert**

noCV practice brief v5 · PBATCH-110 · Import a large supplier catalog without exhausting the worker

Fictional engineering practice briefs. Starter repositories, fixtures, automated grading, and verified ownership are not included.

Phase: Recover and verify. Depends on: PBATCH-105, PBATCH-106, PBATCH-108, PBATCH-109.

Difficulty: Expert. Estimated focused work: 300 minutes; setup and prerequisite tickets are additional.

Estimated field mix: Performance engineering 50% · Data engineering 30% · Quality engineering 20%.

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

The fastest batch configuration exceeds the worker memory budget and fails on interrupted runs. Throughput alone is selecting the wrong candidate.

Acceptance criteria

- Run three paired baseline/candidate comparisons with the same source, database state and resource limits.

- Require completion within the 256 MiB process budget, exact normalized digest and at least 20% higher median committed-record throughput than a completing baseline.

- If the whole-file baseline cannot complete, report that fact and compare throughput against a documented completing bounded baseline; never calculate improvement from a killed run.

Implementation constraints

- Include reset and warmup procedures, peaks, repetitions and recovery observations in the report; the budgets are local exercise conditions.

Verification

- Run normal and slow-writer fixtures and publish all memory peaks and completion outcomes.

- Repeat the commit/checkpoint interruption case with the chosen batch and concurrency settings, then verify final digest and resource cleanup.

Deliverables

- Performance gate, raw measurements and chosen operating configuration

Rollout and recovery: Promote only the configuration satisfying every gate in the synthetic environment; restore the prior bounded configuration if memory or recovery regresses.

Project prerequisites: Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size. Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Engineer value: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

Company value: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

AI tools are welcome during implementation. Record assumptions, review the result, and verify its behavior.

Planning status does not create Outcome Evidence or Ownership Evidence.
