For most of the modern AI boom, intelligence has been built around things machines can see, read, and hear. Large language models learned from text. Multimodal models added images, video, and audio. Robotics followed the same trajectory: cameras became the dominant sensor, and increasingly powerful vision-language-action models learned to turn visual observations and natural-language instructions into physical movements.
But the physical world contains information that cameras cannot see.
A robot can watch its fingers close around a glass, but vision alone may not tell it whether the glass is beginning to slip. It can see a plug approaching a socket, but not necessarily whether the pins are properly seated. It can grasp a piece of fabric without knowing how much tension is building between its fingers. It can hold an egg perfectly in view while still crushing it.
Humans solve these problems almost unconsciously because our hands are not merely mechanical actuators. They are dense sensory systems. Thousands of receptors continuously report pressure, vibration, deformation, texture, temperature, and slip. Much of this information never requires deliberate thought; our nervous system reacts before conscious reasoning becomes involved.
Robots are beginning to acquire something similar.
A wave of robotics research released in 2026 suggests that tactile sensing is moving from a secondary sensor input to a fundamental component of robot intelligence. Systems such as T-Rex, Dream-Tac, TouchWorld, and ViTacWorld are exploring different versions of the same idea: if robots are going to manipulate the real world reliably, they must learn not only to see what is happening, but to understand what physical interaction feels like.
That shift could become one of the most important developments in Physical AI.
The dominance of vision in robotics is understandable. Cameras are inexpensive, information-rich, standardized, and relatively easy to scale. They also fit naturally with the enormous progress made in computer vision and multimodal foundation models.
Modern Vision-Language-Action models, or VLAs, can take an instruction such as “put the red cup in the sink,” interpret a camera image, locate the relevant objects, and generate robot actions. Increasingly large datasets have enabled these models to generalize across objects, environments, and tasks that would once have required individually programmed control systems.
That is an enormous achievement.
But vision has a structural limitation: it observes the world primarily from the outside.
Many of the most important variables in manipulation occur exactly where two surfaces meet. Is an object slipping inside the gripper? How much force is being applied? Has a zipper engaged correctly? Has a key made contact with the inside of a lock? Is a screw cross-threaded? Is a soft object compressed enough to hold but not enough to damage?
Those states may produce almost no visible change.
This becomes particularly problematic as robots move from structured environments toward the messy environments humans inhabit. Moving a rigid box across a warehouse is largely a geometry problem. Folding a shirt, opening packaging, plugging in a cable, preparing food, or manipulating a medical instrument are contact problems.
And contact generates its own kind of data.
One of the most interesting developments in recent research is the realization that tactile intelligence may require a fundamentally different architecture from visual reasoning.
A large vision-language model can spend hundreds of milliseconds reasoning about a scene without creating a serious problem. A hand trying to catch a slipping object cannot.
Human motor control works in a similar hierarchy. We can deliberately decide to pick up a cup, but once the cup begins slipping, our fingers automatically adjust their grip. That correction happens through fast sensorimotor loops rather than conscious planning.
T-Rex, introduced in June, explicitly targets this problem. Instead of simply feeding tactile information into the same slow processing loop used by a conventional VLA, the system uses a variable-rate Mixture-of-Transformers architecture designed to handle touch at a higher temporal frequency. The researchers also collected a roughly 100-hour tactile-rich manipulation dataset focused on reusable motor primitives.
The architectural implication is arguably more important than the benchmark results.
Robotic intelligence may not ultimately be a single giant model processing every sensory modality at the same rate. It may look more like a nervous system, with slower layers responsible for semantic understanding and planning while faster control loops handle immediate physical interaction.
In other words, a robot might use AI to decide what to do, while another layer continuously determines how the contact should feel while doing it.
That begins to resemble biological intelligence much more closely than the conventional “camera plus neural network plus motor” model of robotics.
The next step is even more interesting.
Touch does not have to be purely reactive.
Imagine picking up a ceramic mug. Before your fingers make contact, your brain already has expectations about what should happen. You anticipate where resistance will appear, roughly how heavy the mug will be, and how your grip will change once the object leaves the table.
If reality differs from that expectation, your motor system reacts.
Several new robotics systems are beginning to explore the computational equivalent of this mechanism.
Dream-Tac, released in June, introduces what its researchers call a Tactile-World Action Model. Instead of predicting actions from visual and tactile observations alone, it jointly models actions, future visual observations, and future tactile dynamics. The model therefore attempts to learn how physical contact evolves as a consequence of robot actions. Across six contact-rich manipulation tasks, the researchers report an average 31.7% improvement in action accuracy, alongside techniques intended to make the model practical enough for real-time deployment.
Conceptually, this is a major shift.
A conventional policy asks:
Given what I see and feel now, what action should I take next?
A tactile world model can ask:
If I take this action, what should I expect to see and feel next?
That difference introduces prediction into physical interaction.
And prediction is one of the foundations of intelligence.
A system released in July pushes this concept further.
TouchWorld separates robot manipulation into multiple layers operating at different levels of abstraction. A high-level component uses vision and language to reason about tasks and generate tactile subgoals. A visuo-tactile policy then produces the primary robot actions. Finally, a faster tactile-conditioned refinement policy continuously adjusts those actions using recent touch and proprioceptive feedback.
The structure is important because it separates anticipation from reaction.
The robot can predict the contact state it wants to reach while simultaneously monitoring whether reality matches that prediction. If something slips, becomes misaligned, or encounters unexpected resistance, the fast tactile loop can intervene without requiring the entire high-level reasoning process to start again.
Across six long-horizon contact-rich manipulation tasks, TouchWorld achieved a reported 65% success rate in its clean evaluation setting and 53.7% when humans deliberately perturbed the robot during execution. Those results exceeded the strongest baseline in the study by 15.7 and 18.5 percentage points respectively.
The important idea is not simply that tactile sensors improve performance. Robotics researchers have known that touch matters for decades.
The change is that touch is becoming part of the learned intelligence architecture itself.
Instead of treating tactile data as another sensor value, these models are beginning to use it to represent physical state, predict future interaction, and create fast feedback loops.
That is much closer to how animals use touch.
Another emerging line of research suggests an even broader role for tactile data.
Earlier this year, researchers introduced Visuo-Tactile World Models, showing that incorporating touch could help world models represent contact physics that vision-only systems sometimes get wrong. Vision-based predictions can generate physically implausible outcomes when contact is ambiguous or occluded. Touch provides direct evidence about whether objects are actually colliding, sliding, gripping, or separating.
More recently, ViTacWorld, published in July, explored scaling this idea by training an action-conditioned world model using both real and simulated visuo-tactile trajectories. Given robot actions, the model predicts synchronized future visual observations and tactile feedback. The generated trajectories can then be used both to augment robot training data and to evaluate possible actions before executing them on physical hardware.
This points toward an intriguing possibility. The most useful robot world models may not simply generate videos of what the future will look like. They may generate multisensory futures. A robot considering an action could internally simulate:
What the scene should look like, where contact should occur, how pressure should develop, whether an object should begin to slip, and what tactile signal its fingers should receive.
The robot would not just imagine the future. It would imagine what the future feels like.
There is, however, a major obstacle.
The AI industry became powerful partly because it discovered enormous existing datasets. Language models inherited the internet. Vision models inherited billions of images and videos.
There is no equivalent internet of touch.
Tactile data must usually be generated through physical interaction. Sensors differ dramatically between robot hands. Data collection can damage hardware. Contact events happen at high frequency and produce large streams of information. A dataset captured using one sensor may not transfer cleanly to another.
This makes tactile intelligence expensive.
T-Rex's roughly 100-hour tactile dataset is significant precisely because large, diverse datasets of this kind are still relatively unusual. ViTacWorld's attempt to combine public real-world tactile datasets with simulated interaction addresses the same fundamental bottleneck: how do you scale physical interaction data without operating millions of robots?
This may become one of the defining infrastructure problems in robotics. The previous AI scaling race was about collecting tokens. The next one may be about collecting interactions. Every grasp, slip, collision, insertion, deformation, and correction contains information about how the physical world behaves. A sufficiently large dataset of those interactions could become the equivalent of a physical internet for robot learning.
This leads to a broader question about Physical AI.
For the last several years, much of robotics has attempted to import the scaling logic of language models into the physical world. Build larger models. Collect more demonstrations. Train across more tasks. Add language supervision. Turn robot policies into foundation models.
That approach will almost certainly continue. But physical intelligence may have an additional scaling dimension that language models never needed.
Contact.
Knowing what a cup is and knowing how to pick up a cup are fundamentally different forms of knowledge. The first can be learned from millions of images and descriptions. The second requires understanding friction, force, compliance, weight distribution, geometry, uncertainty, and the consequences of actions over time.
Much of that knowledge is difficult to infer from vision alone. Touch turns abstract physical concepts into direct observations. Friction becomes slip. Force becomes deformation. Stability becomes pressure distribution. Contact becomes a signal. The robot no longer needs to infer all of physics from pixels. Some of it can be felt.
This may ultimately change how we think about robot intelligence. The popular image of Physical AI is essentially a large multimodal model placed inside a humanoid body. The model sees the world, understands an instruction, reasons about what to do, and controls the robot. That architecture may prove incomplete.
Biological intelligence does not work through a single universal reasoning loop. Human behavior emerges from multiple systems operating at different speeds. Vision provides global information. The brain establishes goals and plans. Proprioception tracks the body. Touch measures contact. Reflex pathways make extremely fast corrections. Recent robotics research is beginning to reproduce pieces of this hierarchy.
The evolution increasingly looks something like this:
Vision → Action
became
Vision + Language → Action
and is now becoming
Vision + Language + Touch → Action
But the more interesting destination may be:
Vision + Language + Touch + World Prediction → Action + Reflex
At that point, the robot is no longer simply executing instructions generated by an AI model. It is continuously predicting the physical world, sensing the difference between expectation and reality, and correcting itself through multiple feedback loops. That begins to look less like traditional automation. It begins to look like an artificial nervous system. And that may be the real reason tactile robotics matters.
The next major breakthrough in Physical AI might not be a robot that can think much better. It might be a robot that can finally feel what it is doing.
07/12/2026