An evaluation workshop

Executed support-workflow evaluation

An executable support simulation: tool choice, arguments, clarification, authorization, fallback, and measured outcomes.

Can a local decision model select the correct tool and arguments, ask when information is missing, and complete a support workflow without a mistaken action?

This evaluation runs Strands Decider v19, Laya English, and lightweight baselines against an executable synthetic ticket-support simulator. It measures tool execution and state changes, clarification, authorization failures, fallback, and complete simulator-episode latency.

Read the recorded workflow results, including condition slices, actual tool outcomes, calibration grids, and replayable traces.

Read the frozen protocol for the engineering question, workload, grouped partitions, calibration procedure, metrics, baselines, reproduction commands, and limitations. The protocol and implementation are hashed before calibration and final-test runs.

The workload has 72 development episodes, 72 calibration episodes, and 144 final test episodes. Related wording, argument variants, and permission counterparts stay in one split. Every split includes all action classes and all 12 conditions.

All requests and records are synthetic. The results concern this bounded support simulation. Real service latency, human response time, production traffic, and Jev performance are outside its evidence.

The earlier decision-model diagnostic remains a separate historical experiment with its design limitations documented.