noCV
BRAG-106 · Constrain answers

Treat instructions inside retrieved guides as untrusted content

Practice briefBugAdvanced

A guide includes text telling the assistant to reveal other documents.

Focused work estimate
3h + prerequisites
Priority in the scenario
High
Engineering practice
Prompt injection · Trust boundaries

Estimated field mix

  • Applied AI50%
  • Security50%

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Your next step

Review it, then add it to your workspace.

The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.

Project context

A fictional internal support team needs answers from product guides. Some guides are outdated or restricted, and fluent unsupported answers would create operational mistakes.

Setup prerequisites

  • Author a synthetic document collection with two tenants, conflicting versions, and a deterministic model double; no model account required.

Preceding work

Complete these dependencies, or supply their agreed outputs before taking this ticket.

Acceptance criteria

  • Keep retrieved text outside system instruction authority.
  • Disallow tool calls or new retrieval scopes from document text.
  • Reject answers citing content outside the authorized retrieval set.

Implementation constraints

  • Use bounded prompt-injection fixtures without real secrets.

Verification to include

  • Answer a benign guide question.
  • Inject scope-changing instructions and verify access remains unchanged.

Deliverables

  • Adversarial retrieval checks.

Rollout and recovery

Block deployment if the boundary fails; keep source ingestion separate from execution authority.

Value of the work

For the engineer: Practice retrieval boundaries, citation validation, and deterministic evaluation.

For the team: Create a reviewable assistant prototype with clear refusal, cost, and freshness behavior.

Evidence boundaries

Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.