Evos
Resources

560 hours in 21 days: what live autonomous output looks like

Aug 12, 2026ResearchBy Evos

Freight desks do not run out of knowledge. They run out of hours.

Every load tendered, every exception worked, every invoice checked is somebody's time, and there is never enough of it. That is the constraint on a mid-market desk, and it is the one thing more software has never fixed.

So the question worth asking about an autonomous operator is not how accurate it is. It is how many of those hours it can take.

560 hours, 21 days, live

On a live freight brokerage deployment in Kansas City, Evos operators completed 560 hours of freight workload autonomously across 21 days. Not proposed, not flagged for review, not scored against a human afterwards. Completed, on a live desk, in production.

For scale: one full-time expert working the same window delivers roughly 123 hours, against a baseline of about 173 hours a month. Measured against that baseline, the output is equivalent to roughly 4.8 full-time people.

Why hours is the honest unit

An accuracy percentage answers a question about the model. Hours of completed work answers a question about the operation, which is the only one an ops leader is actually asking.

Hours is also the unit that survives contact with a budget. A hiring plan is denominated in people. So is the gap that made the plan necessary. Output measured in the same unit can be compared directly against the thing you were going to do instead.

What it means that it acted

In production a decision is not a record. It is a message that reaches a broker, a load that gets covered, an invoice that gets queried. Every one of those 560 hours is an action that landed in a real system and stood.

That is what calibration is for. An operator is scored against your own team on a decision type before it is allowed to act on it, which is the process we published in full, 143 exceptions at 82% agreement. By the time an operator is working unsupervised on a case, its judgment on that exact kind of case has already been checked.

The trajectory across our deployments points the same way. The February figures were a cold start and remain the earliest reference point in our record, not current performance. On current deployments, some customers have reached full autonomy on specific roles within four days of onboarding, and the level reached varies role by role inside the same customer.

What the number covers

Two things worth knowing about it.

  • It is one deployment, in one segment of freight, across 21 days.
  • It measures completed workload rather than decision accuracy. The two are tracked separately, and the per-role breakdown behind the 560 is being compiled now.

What it means for a desk your size

The useful question is not whether 560 hours is impressive. It is which hours on your desk are the same shape: repetitive, judgment-heavy, spread across systems that do not talk, and permanently short of the person who was supposed to do them.

An operator runs at 30 to 50% the cost of a person in the same role and is live on the systems you already run inside 24 hours. Against a role you have been trying to fill for months, that is the comparison that matters.

Measure it on your own desk

We start every deployment in shadow, against your own data, so the first number you see is your operation rather than ours. Book an assessment and we will map which hours on your desk an operator would carry, and what that is worth.

Sources: Evos live freight brokerage deployment, Kansas City, 21-day measurement (560 hours of freight workload completed autonomously; ~123-hour human baseline for one full-time expert over the same window against ~173 hours per month; ~4.8 FTE equivalent output). Shadow-mode accuracy figures from the Evos shadow deployment with a mid-market freight operator, February 2026 (143 exceptions, 82% decision accuracy, 70% handled end to end, cold start). Autonomy depth across current deployments as of August 2026. Operator cost as a share of an equivalent human role per Evos capability benchmarks.