Measure peak import memory across parsing, validation and writes
The worker exits near the end of an import, but the existing memory sample is taken only after garbage collection and misses the peak.
- Focused work estimate
- 1h 30m + prerequisites
- Priority in the scenario
- Medium
- Engineering practice
- Memory measurement · Pipeline profiling
Estimated field mix
- Performance engineering70%
- Data engineering30%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional wholesaler imports supplier rows into a staging catalog. The current prototype reads the entire file into memory and restarts from zero after a failure. Build the prototype and generated CSV fixture locally before measuring improvements.
Setup prerequisites
- Generate a deterministic 250,000-row synthetic CSV with quoted newlines, invalid records and a stated maximum record size.
- Create a disposable staging database and an import process constrained to 256 MiB; record runtime and available CPU.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
Acceptance criteria
- Record process RSS, managed heap, buffered-record count and stage throughput at a fixed documented sampling interval.
- Report peak values and exit outcome, including resource-limit termination.
- Separate input-read completion from database-commit completion in elapsed time.
Implementation constraints
- Document sampling limitations and resource limits; do not report a killed run as a completed low-memory run.
Verification to include
- Run the whole-file baseline and retain its peak or termination outcome.
- Inject a slow writer and verify the timeline shows buffered work accumulating before completion.
Deliverables
- Resource timeline and baseline measurement command
Rollout and recovery
Keep measurement output outside imported data; retain failed-run artifacts for the streaming comparison.
Value of the work
For the engineer: Practice streaming, backpressure, allocation analysis and resumable work while retaining exact import semantics.
For the team: Develop a repeatable import performance and recovery exercise that exposes memory, throughput and data-quality tradeoffs.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.