The core imbalance behind industrial AI’s stalled progress is structural: physical‑operations AI is being asked to run real‑time, safety‑critical control workloads on a fraction of the compute maturity the digital economy already built. This is a compute‑maturity lag — the architectural deficit.
OECD data on AI adoption across G7 economies show a clear industry discrepancy: ICT firms reach nearly 45% AI adoption while manufacturing and transportation remain below 10%. This adoption gap does not quantify the compute‑architecture problem directly, but compute scarcity is a structural contributor to it — physical‑operations AI is constrained by limited edge compute, real‑time latency requirements, and safety‑critical workloads.
Layered on top of that structural gap, execution risk remains high. RAND Corporation found that more than 80% of AI projects fail, twice the rate of traditional IT projects. In comparison, only 14% of organizations consider themselves fully prepared to integrate AI despite widespread expectations of business impact.
Stanford’s AI Index documents a large gap between frontier-model capabilities and real-world robotic performance. This gap likely reflects, among other factors, constraints in deployable robotics hardware — including on-device power, latency, and compute limits — though the AI Index itself emphasizes data scarcity, robustness, safety, and real-world generalization rather than explicitly attributing the gap to edge compute.
Infrastructure constraints compound the challenge. The U.S. Department of Energy reports that data-center load has tripled over the past decade and could double or triple again by 2028, while power demand continues to rise across AI workloads. Organizations pursuing AI-driven physical operations must therefore close a readiness gap as compute and power resources become increasingly constrained.
Emerj’s Daniel Faggella recently hosted a conversation with Drew Henry, Executive Vice President of the Physical AI Business Unit at Arm, which designs the compute architecture used across mobile devices, data centers, and — increasingly — industrial and robotic systems. Henry’s vantage point spans mobile, cloud, and now industrial compute — giving leaders a rare lens on how AI shifts from digital to physical operations.
At the heart of the conversation was a critical question for infrastructure and AI leaders across transportation, logistics, manufacturing, and construction: as AI shifts physical operations from automated systems into intelligently controlled ones, how should leaders rethink risk, capital investment, and compute infrastructure?
For operations and infrastructure leaders in these industries, the trade-offs involved differ in specific, practical ways from the ones most executives have already learned to weigh for AI in purely digital settings.
This article examines three core insights that matter most for leaders as AI shifts from automating physical operations to directly controlling them:
- Assured AI outputs for physical equipment control: Bound AI predictions to guaranteed outcomes before granting systems control over equipment, where a wrong call can stop a manufacturing line or disrupt a logistics operation rather than produce a bad chatbot answer.
- Digital-twin simulation for de-risked capital investment: Test AI-driven operational changes virtually, running thousands to millions of trial scenarios, before committing capital to costly physical infrastructure.
- Power-efficient chip architecture for constrained AI deployment: Design compute systems to extract more output per watt, rather than simply adding more silicon, as electricity availability becomes the limiting factor from mobile devices to gigawatt-scale data centers.
Listen to the full episode below:
Episode: Building Compute Foundations for the Physical Economy – with Drew Henry of ARM
Guest: Drew Henry, EVP, Physical AI Business Unit at Arm
Expertise: Physical AI, Technology Strategy, Ecosystem Strategy, Semiconductor Technology
Brief Recognition: Henry previously led Arm’s Strategy and Ecosystems organization and was the founding general manager (GM) of its Infrastructure Business, helping establish Arm-based central processing units (CPUs) in hyperscale cloud data centers. He holds a Master of Science in Electrical Engineering (MSEE) from the University of Southern California and a BS in Engineering Physics from the University of the Pacific.
Assuring AI Outputs Before Connecting Them to Physical Equipment
Companies running logistics centers, factories, and transportation networks have automated physical operations for decades, but Henry draws a sharp line between automating a process and putting AI in control of it.
Automation means the system executes a sequence a human already defined. Intelligent control means an AI model — similar to the large language models (LLMs) that power tools like chatbots, but applied to a physical operation — is making real-time decisions about that sequence, and the cost of a wrong decision changes accordingly.
Henry captures why that shift raises the standard for AI reliability, particularly when model outputs move beyond information and begin influencing physical operations:
“When an LLM hallucinates, that’s a bad answer. When an AI system makes the wrong call on a manufacturing or logistics line, that’s lines down. You’ve got to be incredibly confident in exactly what the outcomes are going to be. It can’t just be something you guess is going to happen — it’s got to be assured that it’s going to happen.”
— Drew Henry, Executive Vice President, Physical AI Business Unit at Arm
That distinction changes how infrastructure and operations teams evaluate an AI system before deployment. A model that is 95% accurate might be acceptable for a recommendation engine; it’s a liability in a system that can stop a production line or create a safety incident.
Henry also pointed to a shared-vocabulary problem underneath the technical one: infrastructure leaders and their AI vendors need a common language for what a given system is bounded to do, or evaluation conversations stay vague.
Before an AI system is given control over physical equipment, the operational cost of a wrong output should determine how tightly its outputs are bounded and monitored, more than average accuracy alone:
- Define the failure mode before deployment. Ask what a wrong decision costs — a stopped line, a safety incident, a shipment error — and size the required confidence level to that cost, not to a generic accuracy benchmark.
- Build a shared technical vocabulary with vendors and partners. Henry noted that the most advanced companies arrive at conversations with infrastructure providers “incredibly informed,” able to discuss specific algorithms and computing platforms rather than general AI capabilities — which speeds up the process of agreeing what a system is bounded to do.
- Treat governance as a prerequisite, not a follow-up. Bounding a system’s use cases before it goes live is what allows teams to expand its authority later with evidence, rather than retrofitting guardrails after an incident.
Henry pointed to Amazon as an example of a company operating at this level of assurance-driven adoption, describing it as “incredibly famous for the use of really, really intelligent robotics” and a company that continually pushes chip and infrastructure partners toward their latest compute roadmaps rather than settling for proven-but-dated systems.
Simulating Operational Change Before Committing Capital
The second insight addresses a capital-allocation problem: once a company decides to change how a physical operation runs, reversing that decision is expensive. Henry described digital twins — virtual replicas of a physical operation — as the mechanism leading companies use to de-risk that decision before spending on new hardware or reconfigured infrastructure.
“If I’m going to get a computing system that’s optimized for the way I operate, I better have a pretty good view of how I operate,” Henry said, describing how companies build a digital representation of a logistics center or manufacturing line specifically so they can test changes there first.
“We’re seeing ratios that are thousands, if not tens of thousands to millions to one, where you’re running simulations of different ways you might do it, before you turn it into a physical representation of how it gets done,” he added.
Henry also distinguished between two systems that need to work together in these environments: the physical automation layer that moves goods or operates machinery, and an optimization layer above it — increasingly built on neural-network-based prediction rather than the fixed-form programming that connected the two layers in the past.
Validating that the optimization layer’s decisions translate correctly to the physical layer in simulation is now part of what digital twins are used to test.
Capital committed to physical infrastructure is difficult to reverse, so the AI-driven changes that will run on that infrastructure should be validated in simulation first — at a volume of test cycles that would be impossible to run against the physical system directly:
- Build the digital twin before the deployment plan. A working virtual representation of the operation is what makes high-volume simulation possible in the first place.
- Test the interface between optimization and execution, not just each layer alone. Henry’s distinction between the “optimization system” and the “physical system” suggests failures often occur at the handoff between the two, not within either one.
- Use simulation volume as a risk-reduction lever, not just a validation step. Running a wide range of scenarios in simulation allows a company to commit capital to a single physical configuration with more confidence than a single pilot could provide.
Engineering Chip Architecture Around Power Efficiency
The third insight reframes what infrastructure leaders should treat as their primary constraint. For most of the last two decades, more compute was simply a matter of adding more processors — a dynamic Henry linked to the Moore’s Law era of the 1990s and early 2000s. That assumption no longer holds.
“A huge manufacturing line or an AI cloud data center is as power-constrained as a mobile phone is, which is a really crazy thing to think about,” Henry said.
Electricity availability, not the number of available processors, is increasingly the ceiling on how much compute an operation can deploy — a shift borne out at the industry level, where AI-focused data center power use is projected to triple by 2030 even as new capacity struggles to keep pace with demand, per the IEA.
That constraint is pushing chip architecture away from generic, general-purpose designs and toward hardware built for specific workloads — one accelerator design for a given class of computation, a different system architecture for another.
Henry connected this shift directly to how companies should approach adoption: the infrastructure leaders getting the most value are the ones who start from a specific operational problem — a manufacturing line’s throughput relative to competitors, for example — and then ask what compute architecture solves that problem, rather than starting from an AI capability and searching for somewhere to apply it.
As an industry-level example of what happens when companies move at different speeds on this kind of transition, Henry pointed to the shift from automated guided vehicles (AGVs) to autonomous mobile robots (AMRs) in logistics centers.
Henry noted that companies that adopted AMRs early captured efficiency gains quickly, while companies that waited are still running the older AGV systems their operations were built around, with switching costs now higher than at the outset.
With power now the limiting resource in AI infrastructure, the compute architecture question shifts from “how much more can we add” to “how much output can we extract per watt for this specific workload” — and starting from the operational problem, rather than the AI capability, is what makes that question answerable:
- Treat power budget as a primary infrastructure constraint, not a secondary cost line, when planning AI-driven compute expansion.
- Match hardware to workload rather than defaulting to general-purpose compute — a bespoke accelerator for one class of computation and a different architecture for another will typically outperform a single generic system tasked with both.
- Start from the operational problem, not the AI capability. Henry described the companies getting the most value as those that bring a specific throughput or efficiency problem to their infrastructure partners, rather than asking generically, “How do we apply AI?”





















































































































































































































































































































































































































































































































































































































































































































































































































































































































































































































































































































