Evaluate summary utility without rewarding confident guesses
The team needs a useful assessment that does not equate fluent prose with correct incident analysis.
- Focused work estimate
- 5h + prerequisites
- Priority in the scenario
- High
- Engineering practice
- AI evaluation · Incident response
Estimated field mix
- Applied AI50%
- Quality engineering30%
- Site reliability20%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional on-call team wants an assistant that joins alerts, deployment notes, and runbooks. Incomplete telemetry and hostile log content make unsupported conclusions dangerous.
Setup prerequisites
- Author synthetic alerts, deployment records, and runbooks plus deterministic tool and model doubles.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- BTRIAGE-101 · Define a common timeline record for operational events
- BTRIAGE-102 · Limit incident retrieval to the caller's service grants
- BTRIAGE-103 · Bound tool arguments and returned telemetry volume
- BTRIAGE-104 · Keep log content from issuing new tool instructions
- BTRIAGE-105 · Require source references for incident summary statements
- BTRIAGE-106 · Represent competing incident hypotheses explicitly
- BTRIAGE-107 · Cancel obsolete drafts when new incident context arrives
- BTRIAGE-108 · Prevent repeated refreshes from exhausting incident budgets
Acceptance criteria
- Measure citation validity, omitted critical observations, and unsupported assertions separately.
- Include partial, contradictory, and malicious context cases.
- Report deterministic boundary results separately from live-model quality.
Implementation constraints
- Do not claim reduced incident duration from this local prototype.
Verification to include
- Detect an invented root cause.
- Accept a concise unresolved summary with valid diagnostic next steps.
Deliverables
- Evaluation protocol and failure report.
Rollout and recovery
Expand incident coverage only after reviewing unsupported-claim failures.
Value of the work
For the engineer: Practice constrained tool use, operational uncertainty, and traceable AI outputs.
For the team: Produce concise incident drafts that preserve source traceability and operator control.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.