noCV
TRIAGE-109 · Operate and learn from corrections

Compare two routing revisions on the same ambiguous-ticket set

Practice briefTaskAdvanced

A new prompt moves more messages to Billing, but the team cannot tell whether it improved routing or merely reduced abstention. Compare both revisions on a fixed synthetic case set.

Focused work estimate
2h 30m + prerequisites
Priority in the scenario
Medium
Engineering practice
AI evaluation · Regression analysis · Release criteria

Estimated field mix

  • Applied AI60%
  • Quality engineering40%

Field percentages are editorial estimates of the ticket's engineering focus. They total 100%; they are not measured time, proficiency scores, or ownership evidence.

Your next step

Review it, then add it to your workspace.

The board opens an editable draft; nothing is saved until you confirm it. Sign-in and workspace permissions apply, and Demo boards remain ephemeral.

Project context

A fictional software vendor receives billing, account-access, bug, and security reports. A prototype silently moves tickets based on vague model confidence. Replace it with bounded suggestions, deterministic safety rules, and an auditable review flow.

Setup prerequisites

  • Create a synthetic support-ticket corpus with no real customer messages.
  • Implement a deterministic local classifier double with success, malformed-output, and timeout modes.

Preceding work

Complete these dependencies, or supply their agreed outputs before taking this ticket.

Acceptance criteria

  • Cases include mixed issues, short messages, quoted text, injection attempts, and explicit security intake.
  • Reports show per-queue outcomes, review rate, rule violations, invalid outputs, and changed decisions.
  • A mandatory-review violation blocks adopting the new revision regardless of aggregate routing accuracy.

Implementation constraints

  • Pin normalization, taxonomy, provider-double, and case-set versions.
  • Keep abstentions visible rather than counting them as successful classifications.

Verification to include

  • Run both revisions on identical inputs and reproduce the decision diff.
  • Inject a security downgrade into one revision and confirm it fails the adoption gate.

Deliverables

  • Routing comparison harness and adoption report

Rollout and recovery

Keep the current revision active while evaluating the new one; switch by versioned configuration after category gates pass.

Value of the work

For the engineer: Practice constrained classification, human correction workflows, model versioning, and evaluation under ambiguous inputs.

For the team: Review whether automation saves triage effort while preserving queue ownership and safe escalation.

Evidence boundaries

Outcome Evidence: Tests, patches, and runbooks are requested deliverables. They become Outcome Evidence only through a qualified Mission and immutable Evidence IDs.

Ownership Evidence: Independent adaptation must be observed under a declared verification policy and cite immutable Evidence IDs. Completing a planning ticket establishes no Ownership Evidence.