Connecting requests, policy decisions, execution, and observed state
I trace one synthetic agent action from its request to its observed effect, then examine what event hashes, snapshots, and replay can establish—and what needs additional verification.
Lessons from an executed tool-selection evaluation
I replaced a label-matching diagnostic with an executable support workflow. On 144 synthetic test episodes, the rules baseline outperformed the tested zero-shot model configurations.