In logistics and transportation, implementing AI is not easy. There is messy data, customer-specific processes, emails, documents, phone calls, exceptions, multiple systems, and frontline teams that already have full plates. So, when an AI project or pilot starts working, there is a natural tendency to prove it.
One metric I keep seeing is some version of:
“Our AI automated 500,000 tasks.”
Or maybe two million actions. Ten million transactions. Pick a big number.
It sounds impressive. But what does it actually mean?
A “task” isn't a standard unit
Think about the word load. If a transportation company told you it moved one million loads last year, you would probably have a few questions.
Was a load a pallet?
A full truckload?
A 20-foot ocean container?
An airfreight ULD?
All of those could legitimately be called a load. But they clearly don't represent the same amount of work, revenue, capacity, or complexity. An “AI task” is even less standardized.
One vendor might count reading an email as a task. Classifying that email could be another. Extracting the shipment number could be another. Looking up the shipment could be another. Drafting the response could be another. Another company might call that entire workflow one task.
Neither is necessarily wrong. But, it makes the statement “we automated one million tasks” pretty hard to interpret.
At best, task counts tell you how much activity is happening inside the technology. They don't necessarily tell you how much value is being created outside of it.
So what should an operations leader measure instead?
I think there are three places to start.
1. Capacity Created
If we are trying to understand what AI did to human work, there is one unit everyone can agree on:
A minute.
If a dispatcher, customer-service rep, pricing coordinator, or documentation specialist used to spend six minutes handling something and now spends one minute, something tangible changed. Five minutes of capacity was created.
Sometimes those minutes represent work that disappeared entirely. Other times, they are repurposed.
The employee doesn't work fewer hours. Instead, they have more time to deal with exceptions, solve difficult customer problems, price another load, follow up on a sales opportunity, or simply get through a busy day without the backlog piling up.
For most trucking companies and freight forwarders, that distinction matters.
The business case for AI doesn't have to be, “How many people can I eliminate?”
It might be:
Can this team handle 20% more volume?
Can we grow without adding headcount at the same rate?
Can we respond to customers faster?
Can experienced operators spend more time on difficult problems?
Can we absorb peak periods without burning everyone out?
There is good evidence for looking at AI this way.
A large real-world study published in the Quarterly Journal of Economics followed 5,172 customer-service agents after an AI assistant was introduced. Instead of counting how many AI actions occurred, the researchers measured what happened to the operation.
Agents with AI resolved about 15% more customer issues per hour on average. That's a number an operations leader can do something with.
2. Operator Trust
But capacity alone isn't enough.
You also have to ask the people doing the work:
Do you trust it?
That can sound like a soft metric compared with minutes and dollars.
I don't think it is.
You can force someone to log into a new system. You cannot force them to depend on it. And that difference becomes increasingly important as AI moves from simply helping an operator toward actually doing work on their behalf.
There are practical ways to see trust developing.
Are operators accepting the AI's work or constantly rewriting it?
Are they comfortable letting it handle routine situations?
When something is wrong, do they understand why?
Do they want to use it on more workflows?
Would they complain if you took it away?
And eventually:
Are they willing to let it act without reviewing every step?
There is a progression here:
Usage → Reliance → Autonomy
Seeing users log in tells you that you've deployed software.
Seeing operators increasingly rely on it tells you that you might actually have something scalable.
3. Customer Impact
Finally, ask a very simple question:
Did anything get better for the customer?
That doesn't always mean NPS.
For a motor carrier, it might mean:
Faster status responses
Fewer missed emails
Quicker quote turnaround
Fewer service failures
Faster exception resolution
For a freight forwarder, it might be:
Fewer documentation errors
Faster shipment updates
Fewer delays caused by missing information
Quicker responses to customers and overseas offices
The metric should fit the workflow.
Interestingly, the same Quarterly Journal of Economics study found that the impact wasn't limited to employee productivity. Customer sentiment also improved, and customers became less likely to ask to escalate a conversation to a manager.
That matters because a workflow can become faster internally without necessarily becoming better externally. You want both.
The three-question AI scorecard
So before scaling an AI pilot across more people, terminals, branches, or workflows, I would ask three questions.
Capacity Created
Did we give meaningful time back to the operation?
Operator Trust
Do the people responsible for the work trust the system enough to rely on it?
Customer Impact
Did the customer or the business experience a better outcome?
You don't need a complicated formula.
A simple red, yellow, green assessment may tell you plenty.
High Capacity + Low Trust + High Customer Impact
Valuable, but difficult to scale.
High Capacity + High Trust + Low Customer Impact
Efficient, but possibly improving the wrong thing.
Low Capacity + High Trust + High Customer Impact
People like it, but the economic case is weak.
High Capacity + High Trust + High Customer Impact
Scale it.
There will, of course, be other considerations. Accuracy, risk, cost, integrations, and implementation effort all matter.
But these three questions tell you something fundamental:
Is the AI actually making the operation better?
Measure what happens around the AI
That question is becoming more important as AI adoption in logistics accelerates.
Their conclusion was that adoption itself is becoming less of a differentiator. The real challenge is turning AI capabilities into measurable and repeatable operational and financial outcomes.
I think that's exactly right.
Today, saying “we use AI” will be about as remarkable as saying “we use software.” And, saying the AI performed two million tasks won't tell us much more.
Instead, look at what happened around the AI.
Did your people get meaningful time back?
Do they trust it enough to rely on it?
Did your customers experience something better?
If the answer to all three is yes, you probably have something worth scaling.
