Generate a catalog file that includes difficult CSV boundaries
The demonstration file has one short record per line. It cannot expose parsers that split quoted descriptions or allocate one enormous record.
- Focused work estimate
- 1h 15m + prerequisites
- Priority in the scenario
- Medium
- Engineering practice
- CSV contracts · Synthetic data
Estimated field mix
- Data engineering50%
- Quality engineering30%
- Performance engineering20%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional wholesaler imports supplier rows into a staging catalog. The current prototype reads the entire file into memory and restarts from zero after a failure. Build the prototype and generated CSV fixture locally before measuring improvements.
Setup prerequisites
- Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size.
- Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.
Preceding work
No earlier ticket is required. Complete the project setup above.
Acceptance criteria
- Generate 250,000 deterministic rows including quoted delimiters, embedded newlines, multibyte text and declared invalid cases.
- Publish expected accepted/rejected counts and a digest of normalized accepted records.
- Declare a maximum record size and include an intentionally oversized record in a separate failure fixture.
Implementation constraints
- Use invented supplier identifiers and descriptions; fixture generation must not require a downloaded customer file.
Verification to include
- Regenerate with the same seed and compare file digest and expected counts.
- Parse with a reference implementation and verify quoted newlines do not change record boundaries.
Deliverables
- Fixture generator, manifest and expected normalized digest
Rollout and recovery
Version the fixture before profiling; keep the oversized fixture separate so the normal baseline has a defined completion result.
Value of the work
For the engineer: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.
For the team: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.