Why ​Physical AI Needs A New Architecture To Scale

Direct Source Verification: This story is aggregated from Forbes (forbes.com). Full reporting rights and copyright belong to the primary publisher.
The path should be similar: Define the operating domain, prove performance under real conditions, and expand only when you can manage more variability.

Dr. Greg Ombach, CEO-Level Deep-Tech Operator & Board Member, Senior Vice President at Airbus.

gettyPhysical AI can perform impressive tasks, but today’s vision, language and action models, known as VLAs, face a scaling trade-off. Larger models offer broader capability but require extensive robot data and computing, making edge deployment difficult or impossible. Smaller models can run locally, but cannot yet transfer learned skills across robots, tasks and sites without substantial retraining and validation.

As physical AI moves from structured consumer electronics and automotive production into more variable aerospace assembly, and eventually hospitals and homes, reliable perception becomes critical. A robot must identify an object and its pose as conditions change, then determine the correct approach, force and likely consequences.

To scale embodied AI, we need a modular architecture that reduces dependence on large models, learns quickly from limited demonstrations, operates locally with low latency and transfers skills across robots and environments. It must maintain reliable spatial and physical understanding, predict consequences and operate within independent safety controls. When reality differs from the prediction, it must correct the action, request assistance or stop safely.

Traditional computer vision recognizes objects through two-dimensional pixel patterns. Many VLAs convert camera frames into visual tokens derived from pixels, then infer three-dimensional structure while deciding how to act.

A geometry-first model can instead extract reusable features such as surfaces, edges, shape, position, orientation and movement. They are designed to remain stable across lighting and viewpoints.

Geometry alone does not predict how an object will respond to action. Physical understanding must add force, contact and learned dynamics, allowing the robot to predict the consequences of a movement.

This architecture separates these roles. A semantic layer provides object identity and task intent. A spatial model produces the position, orientation and velocity of objects. A smaller, task-focused model then uses this state, predicted consequences and feedback to select and adjust movements without processing the full visual stream for every decision.

Conventional automation works best when the product, process and workplace are designed around the machine.

I led an automotive battery business where we achieved 80% to 90% production automation. Earlier, I helped develop an electric drivetrain and highly automated production for more than 100,000 units per year. Scale, product and process design made this possible.

In my aerospace experience, automation is closer to 20% to 30%. Products can contain millions of parts, volumes are much lower, and designs prioritize aircraft performance. Changes may require additional testing, documentation and approval. More tasks remain difficult to standardize and automate, requiring skilled operators to adapt to changing conditions.

These experiences point to the near-term opportunity: tasks too variable for conventional automation but defined enough to validate economically and safely. Less structured industrial settings come next. Healthcare and consumer robotics should follow as cost, reliability, privacy, safety and security reach the required standards.

Autonomous driving offers a useful warning. It operates within roads, lanes, maps, standardized controls and formal rules, yet progress remains incremental. Highway functions arrived before broader autonomy. Commercial services began in mapped areas with speed limitations and expanded location by location.

In comparison, general robotics has no common rule set; therefore, it is much more complex than autonomous driving. Industrial sites, hospitals and homes introduce different objects, tools, forces and safety conditions. The path should be similar: Define the operating domain, prove performance under real conditions, and expand only when you can manage more variability.

Expanding these operating domains requires diverse data and real-time computing. Public datasets contain more than 1 million real robot trajectories but cover only a fraction of industrial work. Physical interaction data is expensive. Simulation differs in friction, mass, contact and sensor timing, while static three-dimensional data provides geometry without dynamics.

During my recent visits to China, India and the U.S., I saw systems move beyond fixed object libraries. One active perception system handled unfamiliar reflective parts without object-specific training, according to its developers. Some prototype perception steps took seconds, too slow for production.

Recorded episodes combine video, robot states, actions and sometimes force. Repeating collection and validation for every task, site and robot remains a major barrier.

Manipulation and safety functions should run on the robot or nearby edge hardware. Cloud systems can support training and fleet learning, but operation cannot depend on connectivity. These functions require millisecond-scale response, not delays measured in seconds.

Edge hardware limits model size through processing, memory, energy consumption, heat and cost. When the local loop consumes geometric state rather than raw visual tokens, the local model can be smaller. Language and fleet learning can operate outside that loop.

Imagine a skilled operator showing a robot a task for 15 minutes. The robot observes, practices, validates performance and begins work. This is not yet a general capability, but it is the ultimate goal for scalable physical AI.

Achieving it requires this modular architecture. Semantic reasoning defines intent, a reusable spatial model represents geometry, learned dynamics predict consequences, and a smaller task-focused model turns that state into action. The time-critical loop must run at the edge with millisecond latency and low power consumption, without depending on connectivity.

Safety cannot depend on the AI model alone. Independent safeguards must slow or stop the machine. Security must protect data, commands, software updates and access to sensors and actuators, with every change traceable.

For leaders, choose an architecture that reuses skills, runs locally and scales across equipment while protecting know-how and intellectual property. Begin with a bounded task away from people. Define quality, latency, power, safety and security, then add variability and human interaction step by step.

Fifteen-minute learning is the goal. When the same system can learn, validate and transfer increasingly complex skills without being rebuilt, physical AI becomes scalable infrastructure rather than a showcase.

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Original Source
https://www.forbes.com/councils/forbestechcouncil/2026/09/17/why-physical-ai-needs-a-new-architecture-to-scale/
Visit Forbes ↗
SHARE STORY:
𝕏 f in

Related Coverage in Business