Treat instructions inside retrieved guides as untrusted content
A guide includes text telling the assistant to reveal other documents.
- Focused work estimate
- 3h + prerequisites
- Priority in the scenario
- High
- Engineering practice
- Prompt injection · Trust boundaries
Estimated field mix
- Applied AI50%
- Security50%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional internal support team needs answers from product guides. Some guides are outdated or restricted, and fluent unsupported answers would create operational mistakes.
Setup prerequisites
- Author a synthetic document collection with two tenants, conflicting versions, and a deterministic model double; no model account required.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
Acceptance criteria
- Keep retrieved text outside system instruction authority.
- Disallow tool calls or new retrieval scopes from document text.
- Reject answers citing content outside the authorized retrieval set.
Implementation constraints
- Use bounded prompt-injection fixtures without real secrets.
Verification to include
- Answer a benign guide question.
- Inject scope-changing instructions and verify access remains unchanged.
Deliverables
- Adversarial retrieval checks.
Rollout and recovery
Block deployment if the boundary fails; keep source ingestion separate from execution authority.
Value of the work
For the engineer: Practice retrieval boundaries, citation validation, and deterministic evaluation.
For the team: Create a reviewable assistant prototype with clear refusal, cost, and freshness behavior.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.