Evos
Resources

143 exceptions, one week: what a shadow deployment measured

Aug 5, 2026ResearchBy Evos

Most claims about AI in operations are untestable. A vendor describes what a system could do, shows a demo on clean data, and leaves the buyer to guess how it behaves on a Tuesday afternoon when three carriers have gone quiet and a customs hold has just landed. We would rather publish a number that can be checked.

So here is one. In February 2026 we ran a one-week shadow deployment against a live mid-market freight operator, on sea and air freight, on the workflow that consumes more ops time than any other: shipment exception handling. The operator processed 143 exceptions. It agreed with the human team on 82% of them. It handled 70% end to end with no person involved. It went from first system connection to first exception processed in under 24 hours.

What a shadow deployment is

A shadow deployment runs the operator on real work without letting it touch anything. We ingested one week of the operator's historical data in real time — every shipment record, status update, document and message — and let the system process each case as it would have live. Every decision was logged. None were executed. At the end of the week we compared the operator's decision on each exception against what the human ops team actually did during the same window.

That design matters, because it removes the two easiest ways to flatter a result. There is no cherry-picked dataset: it is one continuous week of whatever came in. And there is no moving scoreboard: the humans who did the work set the benchmark, case by case.

It was a cold start

This was the first time the operator had seen this company's operation. No tuning window, no months of supervised learning on their data, no hand-built rules for their carriers. It connected to five systems — a Cargowise TMS, an ERPNext instance, carrier portals, customs portals and email logs — and started working.

It is worth being clear that this remains the earliest reference point in our record, and later deployments have exceeded these figures materially. We are publishing the cold-start numbers because they are the honest floor, not the ceiling.

82% agreement with the human team

On 117 of 143 exceptions, the operator reached the same decision as the ops team. On 26 it did not. That 18% gap is the most useful number in the whole test, because it is where the work is. Disagreements clustered in the cases where the right answer depends on something not present in any system — a relationship with a specific carrier's night dispatcher, a customer who tolerates a late delivery but not a surprise, a rule that exists because of something that went wrong in 2021.

This is exactly the tacit layer that never made it into software, and it is why generic models underperform experienced operators on operational work. Closing that gap is not a matter of a larger model. It is a matter of capturing the reasoning from the people who hold it.

70% handled end to end

100 of the 143 exceptions were resolved end to end with no human involvement required — detected, investigated, decided and actioned. The remaining 30% were not failures. They were escalated with the full case assembled: what happened, what was checked, what the options are, and a recommended action. An escalation that arrives complete is worth a great deal more than an alert that arrives empty.

Where the work actually sat

The exception mix is instructive for anyone sizing this work in their own operation. Documentation was the largest share at roughly 35% — cross-referencing documents against shipment records, flagging missing or incorrect BOLs, compliance gaps and filing errors, then generating the correction requests. Carrier communications accounted for around 25%: detecting non-responsiveness, missed pickup and delivery windows, and status-update failures, then drafting outreach carrying the load context and the urgency.

Customer delay handling was another 25% — identifying which delays warranted notification and drafting the communication with the reason, the revised ETA and the recommended next step. Customs made up the final 15%, identifying clearance holds and document requirements and producing resolution recommendations.

Note what that distribution says: none of this is exotic. It is reading, cross-checking, deciding and writing. It is precisely the work that consumes an ops desk and never appears as a line item on any budget.

What it replaced, in hours

At roughly 12 minutes per exception, 143 exceptions a week is about 28.6 hours of manual work — on a two-person ops team, close to 15 hours per person per week. That figure is not incidental. Across legacy industries, 15-plus hours per employee per week disappears into manual operational tasks, and it is the single largest recoverable cost most operations carry.

Detection was never the hard part

Under manual portal checks, this team learned about problems on a four-to-eight-hour delay. The operator detected them in real time. But faster detection on its own changes very little, and this is where most AI in operations stops. The team already knew about the problems. What they lacked was the capacity to resolve them. Cutting detection latency to zero while leaving every resolution on the same two people does not return a single hour.

Why publish the 82% and not just the 70%

Because a number you cannot fail is not a measurement. 95% of GenAI pilots deliver zero P&L impact, and one reason that statistic survives is that the industry reports capability rather than performance. If we only published the 70%, a buyer would have no way to judge what happens in the other 30%, or how often the system is confidently wrong.

An operator should be measured the way a strong team member is measured: did the work get done, did it stay done, how much came back. On a cold start, on one week of real freight exceptions, against the people who do the job: 82% agreement, 70% autonomous, live in under 24 hours. Those are the numbers. We will keep publishing them as they move.

See what this looks like on your desk

A shadow deployment is the least disruptive way to find out what an autonomous operator would do in your operation, because nothing is actioned and nothing changes until you decide it should. Book an assessment and we will run the same test on your exceptions.