ISSUE 001 / A PROPOSED EVALUATION METHOD
Before an AI agent touches your customer workflow write its receipt
You have seen a promising AI tool demo. Now imagine the same tool processing a customer request while you are away. What would you need to see the next morning to know that the work was actually completed?
A fluent answer is useful. A receipt is stronger: the input, the action, the result, the failure path, and the person or system that accepted the outcome.
Start with one observable task
Consider a hypothetical pilot that turns a support request into a draft ticket. Use synthetic examples. Save the request, environment, proposed action, observed result, and acceptance decision for every run. If a value is unknown, write “unknown.”
DeepSeek documents a boundary between requesting a tool call and executing it. Read the official tool-call documentation ↗ (checked 15 September 2026).
Define the acceptance rule first
Write twenty synthetic requests, including missing information, duplicates, and requests outside the agreed scope. Require each run to leave an inspectable draft or a clear handoff. A repeat submission should not silently create a second final ticket.
Twenty cases are a starting exercise, not statistical proof of production reliability. We have not run a model benchmark for this issue.
Count the work around the model
Keep the requests and acceptance rules steady when you compare candidates. Record metered usage, retries, and human review. Separate new cash spending from the allocation of subscriptions you already own. The proposed business metric is total cost per accepted result.
The useful output is a list of concrete failure modes and an estimate of the work needed to fix them. We are not claiming a winning model.