noCV
BRAG-108 · Evaluate limits

Evaluate grounded answering separately from retrieval quality

Practice briefTaskExpert

A single answer score hides whether failures came from missing documents or unsupported generation.

Focused work estimate
6h + prerequisites
Priority in the scenario
High
Engineering practice
Evaluation design · Retrieval

Estimated field mix

  • Applied AI60%
  • Quality engineering40%

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Your next step

Review it, then add it to your workspace.

The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.

Project context

A fictional internal support team needs answers from product guides. Some guides are outdated or restricted, and fluent unsupported answers would create operational mistakes.

Setup prerequisites

  • Author a synthetic document collection with two tenants, conflicting versions, and a deterministic model double; no model account required.

Preceding work

Complete these dependencies, or supply their agreed outputs before taking this ticket.

Acceptance criteria

  • Create separate retrieval and citation-validity checks.
  • Include answerable, absent, conflicting, and denied questions.
  • Report deterministic-double results separately from unmeasured live-model behavior.

Implementation constraints

  • Use public synthetic expectations; do not embed hidden evaluator answers in product APIs.

Verification to include

  • Detect a missing retrieval result.
  • Detect a fluent answer unsupported by retrieved text.

Deliverables

  • Evaluation report with failure categories.

Rollout and recovery

Require both boundaries before expanding the corpus; retain an explicit unmeasured label for live quality.

Value of the work

For the engineer: Practice retrieval boundaries, citation validation, and deterministic evaluation.

For the team: Create a reviewable assistant prototype with clear refusal, cost, and freshness behavior.

Evidence boundaries

Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.