noCV
SEARCH-110 · Evaluate and operate the service

Publish an evaluation report that separates retrieval and answer failures

Practice briefTaskIntermediate

A single accuracy percentage conceals whether failures came from missing sources or invalid generated answers. Create a small versioned evaluation set and report distinct failure categories.

Focused work estimate
2h 30m + prerequisites
Priority in the scenario
Medium
Engineering practice
AI evaluation · Test design · Failure taxonomy

Estimated field mix

  • Applied AI60%
  • Quality engineering40%

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Your next step

Review it, then add it to your workspace.

The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.

Project context

Employees search fictional travel and equipment policies. The current prototype blends draft and approved text and sometimes answers using a policy the employee cannot open. Use a small synthetic corpus and a deterministic answer-provider double.

Setup prerequisites

  • Author synthetic policy documents with revisions, effective dates, and access groups.
  • Use a local deterministic provider double; paid model access is optional and not required.

Preceding work

Complete these dependencies, or supply their agreed outputs before taking this ticket.

Acceptance criteria

  • Cases cover answerable, unsupported, conflicting, restricted, stale, and injection-bearing questions.
  • Reports separate eligible-source retrieval, citation validity, abstention behavior, and provider failures.
  • Each result records corpus, implementation, provider-double, and case-set versions with expected and actual observations.

Implementation constraints

  • Do not claim deterministic-double results measure a real model quality level.
  • Keep evaluation examples synthetic and publishable.

Verification to include

  • Run the full set twice and compare deterministic results.
  • Introduce an access-filter defect and confirm it appears as a security failure rather than an aggregate quality dip.

Deliverables

  • Versioned evaluation cases and categorized report

Rollout and recovery

Use category-specific gates for changes; require a separate measured report before replacing the deterministic adapter.

Value of the work

For the engineer: Practice access-aware retrieval, source provenance, structured model boundaries, and reproducible evaluation.

For the team: Inspect how an engineer prevents unsupported or unauthorized answers and handles changing policy content.

Evidence boundaries

Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.