INDUSTRIAL INSIGHTS

Physical AI has a local problem

An operator's gloved hands loading a tray of empty cans into a filling machine, seen from their own point of view

Hire the best maintenance engineer you can find. Twenty years of experience, excellent instincts, has seen more failure modes than most people will in a career. Put them on your site on a Monday morning.

It will still take six to nine months before they are genuinely effective.

Not because they lack skill — they have more of it than almost anyone. It is because none of their skill is about your plant. They do not yet know that line four runs hot in summer. They do not know which pump was rebuilt two years ago and has behaved oddly since. They do not know that the fault code on that machine is misleading, and everyone here learned long ago to check something else first.

They arrive with extraordinary general capability and zero local knowledge. The gap between those two things is measured in months, and every company pays it, for every senior hire, every time.

FIG 1The gap every company already pays
General capability — full, from day one Local knowledge — earned month by month the gap you pay for every senior hire month 9: finally effective MONDAY MORNING SIX TO NINE MONTHS YEAR ONE
Skill arrives whole. Knowledge of this plant does not. The shaded area is the cost of the distance between them — paid again for every hire, and soon for every machine.

Hold that thought, because it explains almost everything about where physical AI is heading.

The same gap, arriving again

There is an architecture forming in robotics, and it has two layers that are getting a great deal of attention.

At the bottom are the foundation models for physical work — general models for manipulation, movement and perception. They are improving quickly and they are getting cheaper.

At the top are the platforms — the humanoids and robots themselves. Capable hardware, advancing fast, increasingly available to buy.

Both layers are general by design. That is what makes them valuable, and it is also what makes them insufficient. Between them sits a layer nobody is building, and it is exactly the layer your new maintenance engineer spends nine months acquiring: how the work is actually done here.

FIG 2Two layers everyone is building, one nobody is
Site-specific harness How the work is actually done here: your equipment, your materials, your exceptions, your tolerances. The layer your new engineer spends nine months acquiring. LOCAL — MISSING Foundation models for physical work Manipulation, movement, perception. Improving quickly, getting cheaper. GENERAL Platforms Humanoids and robots. Capable hardware, advancing fast, available to buy. GENERAL
The two solid layers are general by design — which is what makes them valuable, and what makes them insufficient. The harness on top is the only layer that is yours.

Consider what a humanoid robot will be like when it arrives. Out of the box it will be genuinely impressive — 1X and others are pushing hard toward machines that handle generic tasks in a home, and folding laundry is essentially a solved problem. Then it will stand in your hallway holding a folded shirt, with no idea which drawer it goes in.

At home, that is solvable. You let it wander around for a few weeks, opening drawers, making mistakes, learning where the cutlery lives. The cost of exploration is a bit of mess and some patience.

You cannot let a robot wander around your plant pushing buttons to find out what they do.

That is where the household analogy breaks, and it breaks completely. On an industrial site the cost of a wrong action is scrapped product, damaged equipment, an unplanned stop, or someone getting hurt. Trial and error is not a learning strategy available to you.

So the local knowledge has to come from somewhere else. It has to come from the people who already have it.

FIG 3The data pyramid: abundant at the bottom, decisive at the top
Site-specific ego-centric video How the work is done on your line, by the people who know it. Scarce, irreplaceable, and not for sale — at any price. General ego-centric video How work is done somewhere. Good priors, wrong specifics. Teleoperation Real contact and real forces — but only along the paths that someone actually drove. Simulation Effectively unlimited, cheap, and confident about a plant that is not yours. LOCAL RELEVANCE VOLUME, AVAILABILITY, LOW COST
Everything below the tip can be bought, generated or scraped — which means your competitors have it too. The tip can only be captured, on site, from your own people.

What we keep seeing

We have spent years around robots in industrial settings, and the same two failures come up repeatedly. Both have the same root cause.

The first is adapting to findings. A robot is set up ad hoc, taught one task along one path, and performs that path well. Then something is slightly off — the material behaves differently, a part sits at an angle, a fixture has shifted two millimetres. An experienced person handles this without thinking, because they have seen it before and know what it means. The robot stops, or worse, carries on. It was never taught the variation, because nobody ever captured the variation.

The second is adapting to flexible operations. This week you run one product, next week another. A human operator carries the changeover in their head — what is different, what to watch for, what usually goes wrong on the first run. If the robot must be re-taught at every changeover, it is not an asset. It is another thing to maintain, competing for the time of the people it was meant to free up.

Both failures come from the same place: the system learned a task, not the work.

The operating envelope

This is the layer we think is missing, and it is worth being concrete about what it contains.

Take lockout-tagout. As a procedure it is universal, and you can find a perfectly good general description in any safety handbook. But the general description is not what keeps people safe. What keeps people safe is knowing which isolation points apply to this machine, in this configuration, on this site — and that is entirely local. The principle generalises. The execution never does.

Or take something as simple as pulling a lever. A general model knows how to pull a lever. It does not know that this lever, on this machine, normally takes a certain amount of force — that below it nothing engages, and above it something is wrong and you should stop rather than push harder. That range is written nowhere. It exists because people have pulled that lever thousands of times and know how it should feel.

FIG 4One lever, one envelope
NORMAL RANGE how it should feel nothing engages stop, do not push harder FORCE APPLIED TO THIS LEVER, ON THIS MACHINE → Written nowhere. Learned by hand, a thousand pulls at a time. A general model knows the lever. It does not know this lever.

That is an operating envelope: the normal range of how work is done here, what falls inside it, and what should trigger a stop. It covers safety, quality and efficiency at once, because it is not a separate policy layer — it is simply a description of how the work is really done.

FIG 5The harness: where general capability meets your site
General capability model + platform GENERIC THE HARNESS Your operating envelope Safe these isolation points, this configuration Correct the variation, the changeover, the exception Efficient the path that works, not the tidy one INSIDE THE ENVELOPE Work done your way this line, this machine OUTSIDE THE ENVELOPE Stop, and ask the person who knows Confirmed by the people who do it. Derived from real work. Continuously updated. Yours, and portable.
The harness is not a policy layer bolted on afterwards. It is the description of how the work is really done — which is what makes a general robot perform safely, correctly and efficiently on your site.

Four things matter about how it is built.

It is derived from real work, then confirmed by people who know it. The system proposes the envelope from what it actually observes on site. Someone who does the job reviews and confirms it. Neither half works alone: observation without expert judgement captures habits as if they were standards, and expert judgement without observation produces the document nobody follows.

It learns from what worked and what did not. Most knowledge systems only capture the clean answer, because that is what gets written down afterwards. But troubleshooting is rarely clean. Symptoms mislead, the obvious cause turns out to be a consequence of something else, and a component gets replaced that was not the problem. That is not a failure of the person — it is the nature of diagnosing a complex machine. What the work reveals about the machine is as valuable as the eventual fix, and it is exactly what nobody writes down. Capturing both means the next person skips the paths that lead nowhere.

FIG 6What gets written down, and what actually happened
THE REPORT fault reported bearing replaced, line running fixed THE ACTUAL WORK sensor suspected — not the problem coupling swapped — symptom, not cause fixed misleading fault code — check the drive first
The dead ends are the expensive part of the knowledge, and the part nobody records. Capture them and the next person skips the paths that lead nowhere.

It is continuously updated. An envelope is not a snapshot. Processes change, equipment is replaced, methods improve. The capture has to be continuous so the envelope describes how you work now, not how you worked two years ago. Knowledge capture is not a project that finishes.

It belongs to you, and it is portable. The envelope is your operational knowledge, and it is not tied to any robot vendor or platform. Capturing it today does not require betting on which hardware you will buy in three years — whatever arrives, your knowledge is already there waiting for it.

And an envelope is useful the moment it exists, long before any robot arrives. It is what lets your new maintenance engineer know the history of a machine on day one instead of month nine. It is what lets a new operator learn from the best person who ever did the job, rather than from whoever happens to be on shift that week.

The part that stays yours

Here is what follows.

General capability is going to arrive cheaply. The foundation models and the platforms will be broadly available, to you and to everyone you compete with. That is a commodity in the making, and it is not where advantage sits.

Advantage sits in the local layer — in site-specific knowledge: how work is done at your site, with your equipment, your materials, your constraints and your exceptions. That knowledge exists today, in your people, and it is not written down. It is also leaving, a little at a time, as they retire.

Capture it and it becomes yours — an asset that makes your people effective faster now, and that any future system starts from rather than relearning at your expense. Fail to capture it and you pay the nine-month curve over and over, for every hire and eventually for every machine.

The generic part is coming for free.
The local part never will.

That is what we are building at dotspot: the means to capture how skilled work is actually done, from the people who still know how to do it, and to keep that knowledge where it belongs — with the company whose people created it.

— Marcus Horn, CEO & Co-founder, dotspot

← All insights