LEARN · ARTICLE

Stop Counting AI Tasks: A Better Way to Measure AI in Logistics

person holding white mini bell alarmclock

By:

John K Chang

In logistics and transportation, implementing AI is not easy. There is messy data, customer-specific processes, emails, documents, phone calls, exceptions, multiple systems, and frontline teams that already have full plates. So, when an AI project or pilot starts working, there is a natural tendency to prove it.

One metric I keep seeing is some version of:

“Our AI automated 500,000 tasks.”

Or maybe two million actions. Ten million transactions. Pick a big number.

It sounds impressive. But what does it actually mean?

A “task” isn't a standard unit

Think about the word load. If a transportation company told you it moved one million loads last year, you would probably have a few questions.

Was a load a pallet?

A full truckload?

A 20-foot ocean container?

An airfreight ULD?

All of those could legitimately be called a load. But they clearly don't represent the same amount of work, revenue, capacity, or complexity. An “AI task” is even less standardized.

One vendor might count reading an email as a task. Classifying that email could be another. Extracting the shipment number could be another. Looking up the shipment could be another. Drafting the response could be another. Another company might call that entire workflow one task.

Neither is necessarily wrong. But, it makes the statement “we automated one million tasks” pretty hard to interpret.

At best, task counts tell you how much activity is happening inside the technology. They don't necessarily tell you how much value is being created outside of it.

So what should an operations leader measure instead?

I think there are three places to start.

1. Capacity Created

If we are trying to understand what AI did to human work, there is one unit everyone can agree on:

A minute.

If a dispatcher, customer-service rep, pricing coordinator, or documentation specialist used to spend six minutes handling something and now spends one minute, something tangible changed. Five minutes of capacity was created.

Sometimes those minutes represent work that disappeared entirely. Other times, they are repurposed.

The employee doesn't work fewer hours. Instead, they have more time to deal with exceptions, solve difficult customer problems, price another load, follow up on a sales opportunity, or simply get through a busy day without the backlog piling up.

For most trucking companies and freight forwarders, that distinction matters.

The business case for AI doesn't have to be, “How many people can I eliminate?”

It might be:

  • Can this team handle 20% more volume?

  • Can we grow without adding headcount at the same rate?

  • Can we respond to customers faster?

  • Can experienced operators spend more time on difficult problems?

  • Can we absorb peak periods without burning everyone out?

There is good evidence for looking at AI this way.

A large real-world study published in the Quarterly Journal of Economics followed 5,172 customer-service agents after an AI assistant was introduced. Instead of counting how many AI actions occurred, the researchers measured what happened to the operation.

Agents with AI resolved about 15% more customer issues per hour on average. That's a number an operations leader can do something with.

2. Operator Trust

But capacity alone isn't enough.

You also have to ask the people doing the work:

Do you trust it?

That can sound like a soft metric compared with minutes and dollars.

I don't think it is.

You can force someone to log into a new system. You cannot force them to depend on it. And that difference becomes increasingly important as AI moves from simply helping an operator toward actually doing work on their behalf.

There are practical ways to see trust developing.

  • Are operators accepting the AI's work or constantly rewriting it?

  • Are they comfortable letting it handle routine situations?

  • When something is wrong, do they understand why?

  • Do they want to use it on more workflows?

  • Would they complain if you took it away?

And eventually:

Are they willing to let it act without reviewing every step?

There is a progression here:

Usage → Reliance → Autonomy

Seeing users log in tells you that you've deployed software.

Seeing operators increasingly rely on it tells you that you might actually have something scalable.

3. Customer Impact

Finally, ask a very simple question:

Did anything get better for the customer?

That doesn't always mean NPS.

For a motor carrier, it might mean:

  • Faster status responses

  • Fewer missed emails

  • Quicker quote turnaround

  • Fewer service failures

  • Faster exception resolution

For a freight forwarder, it might be:

  • Fewer documentation errors

  • Faster shipment updates

  • Fewer delays caused by missing information

  • Quicker responses to customers and overseas offices

The metric should fit the workflow.

Interestingly, the same Quarterly Journal of Economics study found that the impact wasn't limited to employee productivity. Customer sentiment also improved, and customers became less likely to ask to escalate a conversation to a manager.

That matters because a workflow can become faster internally without necessarily becoming better externally. You want both.

The three-question AI scorecard

So before scaling an AI pilot across more people, terminals, branches, or workflows, I would ask three questions.

Capacity Created

Did we give meaningful time back to the operation?

Operator Trust

Do the people responsible for the work trust the system enough to rely on it?

Customer Impact

Did the customer or the business experience a better outcome?

You don't need a complicated formula.

A simple red, yellow, green assessment may tell you plenty.

High Capacity + Low Trust + High Customer Impact

Valuable, but difficult to scale.

High Capacity + High Trust + Low Customer Impact

Efficient, but possibly improving the wrong thing.

Low Capacity + High Trust + High Customer Impact

People like it, but the economic case is weak.

High Capacity + High Trust + High Customer Impact

Scale it.

There will, of course, be other considerations. Accuracy, risk, cost, integrations, and implementation effort all matter.

But these three questions tell you something fundamental:

Is the AI actually making the operation better?

Measure what happens around the AI

That question is becoming more important as AI adoption in logistics accelerates.

McKinsey recently reported that nearly 90% of surveyed shippers had already adopted at least one transportation AI use case.

Their conclusion was that adoption itself is becoming less of a differentiator. The real challenge is turning AI capabilities into measurable and repeatable operational and financial outcomes.

I think that's exactly right.

Today, saying “we use AI” will be about as remarkable as saying “we use software.” And, saying the AI performed two million tasks won't tell us much more.

Instead, look at what happened around the AI.

Did your people get meaningful time back?

Do they trust it enough to rely on it?

Did your customers experience something better?

If the answer to all three is yes, you probably have something worth scaling.

What would make this easier for the people doing the work?

Join our newsletter

Sign up to our mailing list below and be the first to know about new updates. Don't worry, we hate spam too.

© 2026 Ibis Laboratories, Inc.

Join our newsletter

Sign up to our mailing list below and be the first to know about new updates. Don't worry, we hate spam too.

© 2026 Ibis Laboratories, Inc.

Join our newsletter

Sign up to our mailing list below and be the first to know about new updates. Don't worry, we hate spam too.

© 2026 Ibis Laboratories, Inc.