noCV
BEXTRACT-109 · Support correction

Design a field-level extraction evaluation with abstention costs

Practice briefTaskExpert

Document-level success hides whether the extractor frequently invents totals or simply leaves them for review.

Focused work estimate
6h + prerequisites
Priority in the scenario
High
Engineering practice
Evaluation design · AI quality

Estimated field mix

  • Applied AI60%
  • Quality engineering40%

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Your next step

Review it, then add it to your workspace.

The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.

Project context

A fictional procurement team receives inconsistent supplier documents. It wants draft records that staff can review without treating model output as authoritative accounting data.

Setup prerequisites

  • Author synthetic text invoices with ambiguous dates, currencies, and line items; use a deterministic provider double.

Preceding work

Complete these dependencies, or supply their agreed outputs before taking this ticket.

Acceptance criteria

  • Report supported correct, incorrect, missing, and abstained fields separately.
  • Include layout, date, currency, and injection variations.
  • Compare manual-review volume against unsupported-field risk under declared assumptions.

Implementation constraints

  • Deterministic doubles validate plumbing; live model accuracy remains unmeasured.

Verification to include

  • Detect a fabricated total in the evaluation.
  • Show an appropriate abstention separately from an incorrect value.

Deliverables

  • Evaluation protocol and synthetic results.

Rollout and recovery

Expand document coverage only after reviewing failure categories and unresolved assumptions.

Value of the work

For the engineer: Practice schema constraints, numerical checks, review workflows, and safe model boundaries.

For the team: Produce an inspectable extraction prototype that can reduce manual transcription while preserving review control.

Evidence boundaries

Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.