noCV
PBATCH-101 · Define input and resource behavior

Generate a catalog file that includes difficult CSV boundaries

Practice briefTaskFoundational

The demonstration file has one short record per line. It cannot expose parsers that split quoted descriptions or allocate one enormous record.

Focused work estimate
1h 15m + prerequisites
Priority in the scenario
Medium
Engineering practice
CSV contracts · Synthetic data

Estimated field mix

  • Data engineering50%
  • Quality engineering30%
  • Performance engineering20%

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Your next step

Review it, then add it to your workspace.

The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.

Project context

A fictional wholesaler imports supplier rows into a staging catalog. The current prototype reads the entire file into memory and restarts from zero after a failure. Build the prototype and generated CSV fixture locally before measuring improvements.

Setup prerequisites

  • Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size.
  • Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Preceding work

No earlier ticket is required. Complete the project setup above.

Acceptance criteria

  • Generate 250,000 deterministic rows including quoted delimiters, embedded newlines, multibyte text and declared invalid cases.
  • Publish expected accepted/rejected counts and a digest of normalized accepted records.
  • Declare a maximum record size and include an intentionally oversized record in a separate failure fixture.

Implementation constraints

  • Use invented supplier identifiers and descriptions; fixture generation must not require a downloaded customer file.

Verification to include

  • Regenerate with the same seed and compare file digest and expected counts.
  • Parse with a reference implementation and verify quoted newlines do not change record boundaries.

Deliverables

  • Fixture generator, manifest and expected normalized digest

Rollout and recovery

Version the fixture before profiling; keep the oversized fixture separate so the normal baseline has a defined completion result.

Value of the work

For the engineer: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

For the team: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

Evidence boundaries

Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.