Zero-Gravity.ai
FILE 012

Field Notes

5 min read

The pilot worked. The Tuesday morning didn't change.

The demo was convincing, the output made sense, and the pilot was called a success. Then normal work resumed — same spreadsheet, same queue, same report. A pilot proves that something can work. It does not prove that the organisation can work differently because of it.

Written by Bart Van Mulders

The signal

Monday afternoon, the demo works.

The model returns something useful. The output is explainable enough. Someone asks a hard question and the answer holds. The room is genuinely impressed — not politely, but properly. A slide is produced. A next phase is discussed, and sometimes funded.

Tuesday morning, the planner opens the same application. Downloads the same report. Copies the numbers into the same spreadsheet. Applies the same three rules of thumb that live in nobody's documentation. Makes the decision the way it has always been made.

The service agent still works the same queue in the same order. The manager still waits for the same weekly report. The analyst still prepares the same manual correction before anyone sees the figures.

Nothing failed. That is the strange part. The pilot did exactly what a pilot is designed to do. It changed what was technically possible, and it did not change what the organisation actually does.

The pattern

Pilots are optimised for the thing they are asked to prove.

What a pilot is usually asked to prove is that the capability exists: that the data can be assembled, that the model performs acceptably, that the output is coherent, that the idea survives contact with a real dataset. Those are legitimate questions and they deserve a serious answer.

What a pilot is rarely asked to prove is that a recurring operational moment can become different. Nobody sets the success criterion as the planner stops rebuilding the spreadsheet, or the intervention happens on the day the signal appears instead of eleven days later. The criterion is model quality, demo quality, and stakeholder enthusiasm.

So the pilot is designed backwards from the demo rather than backwards from the work. It runs beside the process instead of inside it. It uses an extract rather than the live operational context. It is evaluated by the people who built it and the people who sponsored it, rather than by the person who would have to change their Tuesday.

Then it succeeds — on its own terms — and everyone reasonably expects the rest to follow.

What is actually stuck

It helps to separate three things that get compressed into one word.

AI output. Something is generated, predicted, classified, ranked or recommended. This is what the pilot proved.

Decision input. That output reaches the moment where a choice is made, in a form the decider can act on, with enough context to be trusted, at the time the choice is actually made. Not a day later, and not in a separate tool.

Operational action. The decision changes something in the operating reality — a schedule, a price, a route, a priority, an intervention — and something downstream records that it happened.

Most stalled pilots have solved the first and neither of the other two.

The chain breaks in fairly ordinary places. There is no owner for the moment of work, only an owner for the model. The output lives in a separate interface, so using it costs the user an extra step they were not given time for. Exceptions have no logic, so the first odd case sends the user back to the old method permanently. Decision rights are unclear, so a recommendation that contradicts the usual answer needs an escalation nobody scheduled. Trust was never built, because the user saw the demo but never saw the model handle their difficult week. And nothing measures whether behaviour changed, so the question is never asked.

None of these are AI problems. All of them determine whether the AI is used.

This is not an argument that model quality is irrelevant. A weak model fails, and it should. The point is narrower: a good model is necessary and, on its own, insufficient.

Production is not operationalisation

The most useful distinction here is also the one most often skipped.

Moving a model from prototype infrastructure into a production environment is real engineering work. It matters. It is also frequently mistaken for the finish line, because it is the last step that is fully within the delivery team's control.

But the sequence continues, and each step can fail independently:

Deployed does not mean used. A running endpoint with no user is an operating cost.

Used does not mean decision-changing. People will happily look at an output, find it interesting, and then decide exactly what they would have decided anyway.

Decision-changing does not mean outcome-producing. A different decision that never reaches an action, or reaches it too late, leaves the outcome where it was.

Every one of these transitions has an owner, a workflow implication and something that could be measured. Production only clears the first.

What changed

The shift is in the opening question.

Instead of asking how do we productionise the model?, the more productive question is: which recurring operational moment should become different because this exists?

That question forces a specific moment into view — a Tuesday morning, a shift handover, a weekly prioritisation, a customer call. From there the reasoning runs backwards rather than forwards:

outcome → action → decision → required intelligence → data and context → AI capability.

Working in this direction changes what gets built. The action defines who must be able to act, and therefore who must trust the output. The decision defines the moment, the latency budget and the authority required. The intelligence requirement defines what the model actually needs to be good at — which is often narrower and more tractable than the pilot assumed.

It also changes what gets proved. The success criterion stops being the model performs and becomes this moment now behaves differently, and here is the before and after.

That is a smaller claim than most pilots make, and a considerably more useful one.

What we learned

The last mile of AI is not deployment. It is changed behaviour inside a real process.

A pilot that ends in a convincing demonstration has proved capability. A pilot that ends with someone doing their work differently — and being able to show what moved — has proved something else entirely.

The question worth asking at the end of a successful pilot is not whether the AI worked.

It is whether Tuesday morning looks any different.

Key takeaways

  • 01A pilot proves technical feasibility; it does not prove operational adoption.
  • 02Deployed is not used, used is not decision-changing, and decision-changing is not outcome-producing.
  • 03Value appears when a recurring moment of work becomes measurably different.

From recognition to movement

  1. 01

    Test it

    Pick one pilot that was called successful.

    Do not start with the technology. Start with the moment of work it was supposed to change, then compare before and after:

    • 01Which recurring moment of work was supposed to become different?
    • 02Before the pilot: what did someone see, decide and do?
    • 03After the pilot: what do they now see, decide and do differently?
    • 04Did any action or timing change — or only the availability of an output?

    The important question is not whether the AI works. It is whether the work changed.

  2. 02

    Understand it

    Automation Readiness

    Automation readiness is not simply technical readiness. It depends on sufficient clarity in information, meaning, decisions, actions, controls and operational context.

    Explore the concept →
  3. 03

    Move it

    Build Momentum

    Help the organisation sustain movement, learning and capability after the initial intervention.

    See how it works →

    Dealing with something like this? Bring us the challenge →

    How intelligence becomes part of the operating routine →

ShareLinkedInXEmail