Publish an evaluation report that separates retrieval and answer failures
A single accuracy percentage conceals whether failures came from missing sources or invalid generated answers. Create a small versioned evaluation set and report distinct failure categories.
- Focused work estimate
- 2h 30m + prerequisites
- Priority in the scenario
- Medium
- Engineering practice
- AI evaluation · Test design · Failure taxonomy
Estimated field mix
- Applied AI60%
- Quality engineering40%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
Employees search fictional travel and equipment policies. The current prototype blends draft and approved text and sometimes answers using a policy the employee cannot open. Use a small synthetic corpus and a deterministic answer-provider double.
Setup prerequisites
- Author synthetic policy documents with revisions, effective dates, and access groups.
- Use a local deterministic provider double; paid model access is optional and not required.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- SEARCH-101 · Import policy documents with stable revision identities
- SEARCH-102 · Keep chunk citations anchored to the original policy text
- SEARCH-103 · Apply group access before ranking or sending context to a provider
- SEARCH-104 · Exclude drafts and future policies from current-policy answers
- SEARCH-105 · Reject answers with invented or mismatched citations
- SEARCH-106 · Return insufficient information when the corpus cannot answer
- SEARCH-107 · Contain instructions embedded in a retrieved document
- SEARCH-108 · Bound provider latency, retries, and retained query data
- SEARCH-109 · Invalidate answers when policy authority or access changes
Acceptance criteria
- Cases cover answerable, unsupported, conflicting, restricted, stale, and injection-bearing questions.
- Reports separate eligible-source retrieval, citation validity, abstention behavior, and provider failures.
- Each result records corpus, implementation, provider-double, and case-set versions with expected and actual observations.
Implementation constraints
- Do not claim deterministic-double results measure a real model quality level.
- Keep evaluation examples synthetic and publishable.
Verification to include
- Run the full set twice and compare deterministic results.
- Introduce an access-filter defect and confirm it appears as a security failure rather than an aggregate quality dip.
Deliverables
- Versioned evaluation cases and categorized report
Rollout and recovery
Use category-specific gates for changes; require a separate measured report before replacing the deterministic adapter.
Value of the work
For the engineer: Practice access-aware retrieval, source provenance, structured model boundaries, and reproducible evaluation.
For the team: Inspect how an engineer prevents unsupported or unauthorized answers and handles changing policy content.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.