Compare two routing revisions on the same ambiguous-ticket set
A new prompt moves more messages to Billing, but the team cannot tell whether it improved routing or merely reduced abstention. Compare both revisions on a fixed synthetic case set.
- Focused work estimate
- 2h 30m + prerequisites
- Priority in the scenario
- Medium
- Engineering practice
- AI evaluation · Regression analysis · Release criteria
Estimated field mix
- Applied AI60%
- Quality engineering40%
Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.
Review it, then add it to your workspace.
The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.
Project context
A fictional software vendor receives billing, account-access, bug, and security reports. A prototype silently moves tickets based on vague model confidence. Replace it with bounded suggestions, deterministic safety rules, and an auditable review flow.
Setup prerequisites
- Create a synthetic support-ticket corpus with no real customer messages.
- Implement a deterministic local classifier double with success, malformed-output, and timeout modes.
Preceding work
Complete these dependencies, or supply their agreed outputs before taking this ticket.
- TRIAGE-101 · Define queue labels with examples and an unknown outcome
- TRIAGE-102 · Normalize inbound messages without discarding the original record
- TRIAGE-103 · Route declared security incidents to review before calling a classifier
- TRIAGE-104 · Accept only bounded suggestions from the classifier
- TRIAGE-107 · Keep provider outages from blocking the support inbox
- TRIAGE-105 · Show a suggested queue as an explicit agent action
- TRIAGE-106 · Discard a suggestion when the ticket changed while classification ran
- TRIAGE-108 · Record corrections without turning agent clicks into automatic training data
Acceptance criteria
- Cases include mixed issues, short messages, quoted text, injection attempts, and explicit security intake.
- Reports show per-queue outcomes, review rate, rule violations, invalid outputs, and changed decisions.
- A mandatory-review violation blocks adopting the new revision regardless of aggregate routing accuracy.
Implementation constraints
- Pin normalization, taxonomy, provider-double, and case-set versions.
- Keep abstentions visible rather than counting them as successful classifications.
Verification to include
- Run both revisions on identical inputs and reproduce the decision diff.
- Inject a security downgrade into one revision and confirm it fails the adoption gate.
Deliverables
- Routing comparison harness and adoption report
Rollout and recovery
Keep the current revision active while evaluating the new one; switch by versioned configuration after category gates pass.
Value of the work
For the engineer: Practice constrained classification, human correction workflows, model versioning, and evaluation under ambiguous inputs.
For the team: Review whether automation saves triage effort while preserving queue ownership and safe escalation.
Evidence boundaries
Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.
Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.