noCV
PBATCH-105 · Bound the import pipeline

Batch staging writes without changing duplicate-SKU behavior

Practice briefStoryAdvanced

Single-row inserts dominate after streaming is introduced. A bulk-insert experiment is faster but resolves duplicate supplier SKUs differently.

Focused work estimate
3h 30m + prerequisites
Priority in the scenario
High
Engineering practice
Batching · Transaction semantics

Estimated field mix

  • Database engineering50%
  • Data engineering30%
  • Performance engineering20%

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Your next step

Review it, then add it to your workspace.

The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.

Project context

A fictional wholesaler imports supplier rows into a staging catalog. The current prototype reads the entire file into memory and restarts from zero after a failure. Build the prototype and generated CSV fixture locally before measuring improvements.

Setup prerequisites

  • Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size.
  • Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.

Preceding work

Complete these dependencies, or supply their agreed outputs before taking this ticket.

Acceptance criteria

  • Document duplicate-SKU ordering and preserve it across batch boundaries.
  • Cap batch rows and serialized bytes, with bounded transaction duration.
  • Return deterministic accepted/rejected outcomes when one batch includes invalid or conflicting records.

Implementation constraints

  • Compare at least three bounded batch sizes on the fixed fixture; do not choose the largest solely from one fast run.

Verification to include

  • Place duplicate SKUs on either side of a batch boundary and compare normalized results with the reference behavior.
  • Inject a write failure midway through a batch and verify the declared atomicity and retry outcome.

Deliverables

  • Batched writer, batch-size measurements and duplicate regressions

Rollout and recovery

Canary in staging with batch size configurable; reducing it must preserve semantics and allow work to continue from a valid checkpoint.

Value of the work

For the engineer: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.

For the team: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.

Evidence boundaries

Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.