Robots are no longer just “seeing” the world—they’re starting to feel it with sensing that looks more like engineering than AI. The clearest signal came in April 2026, when UltraSense Systems unveiled an ultrasound-based tactile intelligence platform for Physical AI, built around a protected sub-surface sensing architecture and paired with evaluation kits planned to ship starting June. That shift matters because real-world manipulation fails not when a robot lacks a camera, but when it can’t reliably infer how an object deforms, slips, or resists.
At the same time, the intelligence that drives those robots is evolving beyond perception-only systems. A Nature Machine Intelligence analysis published on November 16, 2025 focused on how to build vision–language–action (VLA) models for generalist robots—pinning down design decisions like what backbone to use, how to structure the VLA architecture, and when to fuse modalities. And in September 2025, Tohoku University researchers described “TactileAloha,” a multimodal Physical AI approach that combines vision and tactile sensing to manipulate objects with a robotic arm, reporting improved task success compared with methods that rely on vision alone. Together, these developments suggest a near-term arc for Physical AI: make the sensing physically reliable, then make the action policy modality-aware and robust.

Ultrasound touch: the next sensing bottleneck for real manipulation
UltraSense’s ultrasound approach targets a problem tactile sensors have wrestled with for years: contact signals are noisy, degrade under use, and can be hard to protect on a robot that’s expected to handle objects repeatedly. By using a protected, sub-surface sensing architecture, the company is effectively moving critical measurement away from the most failure-prone surface layer. The result is a tactile channel that’s designed to survive the “messy middle” of robotics—scrapes, dust, repeated impacts, and varying contact geometries.
What’s especially telling is timing and deployment intent. The announcement didn’t position the system as a lab curiosity; it stated evaluation kits would be available starting June. That implies a validation path where researchers and integrators can test tactile intelligence against common manipulation benchmarks—under controlled grips, controlled slips, and controlled object variations—before scaling up. In Physical AI, the gap between a compelling demo and dependable behavior often comes down to whether a sensing modality can be trusted across conditions, not whether it can produce a pretty visualization.
From a systems perspective, ultrasound tactile sensing can also complement vision in a way that’s mechanically aligned with the problem. Vision tells you what’s where; it struggles when the object’s surface texture changes, when lighting shifts, or when occlusion hides the relevant region. Ultrasound, by contrast, can measure phenomena closer to contact physics—helping the robot infer deformation and internal changes that are invisible to cameras. That matters for tasks that require finesse: aligning a fragile part, controlling insertion force, or preventing slippage during grasp-and-lift cycles.
Why VLA models alone aren’t enough: the modality-matching challenge
VLA models are the flagship approach for generalist robotics because they turn instructions and sensory observations into action. But designing them isn’t just about picking a strong backbone network; it’s about deciding when and how to connect perception, language grounding, and action generation. The Nature Machine Intelligence work dated November 16, 2025 emphasizes exactly those decisions—especially how to formulate the VLA architecture and when to add specific components that affect multi-modal alignment.
In practice, the VLA design question becomes: what part of the model is responsible for handling uncertainty from the physical world? Vision-only policies can learn “likely” outcomes, but they struggle when the same visual scene yields different contact dynamics—like a soft object compressing more than expected, a smooth surface slipping on the third attempt, or a hidden defect changing how forces transmit. Without a tactile channel, the model’s uncertainty has nowhere to go; it remains internal, and the action policy still commits to a grasping trajectory.
This is where the ultrasound platform and tactile multimodality experiments become more than parallel threads. A physically grounded tactile channel can serve as an external correction signal for VLA policies. If your model is instruction-following, it still needs a feedback loop tied to the outcome: “did the gripper actually achieve the intended contact state?” Tactile intelligence provides a measurable proxy for those outcomes, helping reduce the gap between language-level intent (“pick up the cup without spilling”) and low-level motor control.
TactileAloha and the case for multimodal Physical AI
Tohoku University’s “TactileAloha,” reported on September 3, 2025, offers an instructive example of how Physical AI benefits from combining vision and tactile sensing in a single manipulation framework. By building a multimodal approach around a robotic arm, the researchers reported improved task success—an important detail because success rates, not just qualitative trajectories, determine whether a method scales beyond curated demonstrations.
What tends to happen in the field is that tactile cues are treated as an auxiliary input: useful but optional. TactileAloha flips that framing by making tactile sensing a first-class contributor to the action. That typically changes the failure mode. Instead of “the robot didn’t see the slip soon enough,” failures become “the policy didn’t interpret tactile signals correctly,” which is a more learnable problem—especially when the learning system has enough data to map tactile patterns to recovery actions like re-grasping, adjusting grip force, or changing approach angles.
When you connect this back to ultrasound’s sub-surface architecture, a compelling hypothesis emerges: robust tactile sensing isn’t just about adding another sensor—it’s about making tactile data stable enough that VLA-style policies can learn consistent relationships. If tactile signals drift or are too noisy, the learning system may either overfit to sensor quirks or discount tactile inputs entirely. A protected architecture is therefore not a hardware footnote; it’s a prerequisite for reliable multimodal learning.
From demos to dependable systems: what to build next
The three developments—UltraSense’s June-available ultrasound tactile evaluation kits announced April 15, 2026; the VLA architecture guidance from Nature Machine Intelligence dated November 16, 2025; and TactileAloha’s September 3, 2025 multimodal manipulation results—point toward a practical roadmap for teams building Physical AI systems.
First, treat tactile sensing as a core modality with explicit reliability targets. Before optimizing policies, define how you will measure tactile signal stability across repeated contact: does the signal preserve task-relevant features after dozens of trials, do readings remain consistent under varying object textures, and does the system degrade gracefully when surfaces are dusty or partially occluded? If evaluation kits are available starting June, use them early to build a tactile dataset that captures these physical variations.
Second, align your VLA architecture choices with your modality fusion strategy. If the November 2025 analysis emphasizes specific design and backbone decisions, the actionable takeaway is to avoid “late fusion by default.” Early experiments should test whether tactile should enter the network before or after instruction grounding, and whether it should modulate grasp parameters directly or only influence higher-level action selection. The goal is not maximum model complexity; it’s ensuring that tactile evidence can actually steer decisions when vision becomes ambiguous.
Third, choose evaluation tasks that force contact, not just perception. Physical AI needs benchmarks where vision can mislead and tactile must correct. Insertions, squeezes, and slip-prone grasps are ideal because they create measurable physical outcomes: success can be quantified as insertion depth achieved without jamming, slip prevented across a lift cycle, or force/torque thresholds maintained while following an instruction.
In other words, the next wave of Physical AI won’t be defined by whether robots can answer prompts or follow trajectories—it will be defined by whether tactile intelligence is robust enough to make generalist action policies dependable in the real world.
Key takeaways
- Ultrasound-based tactile intelligence with protected sub-surface sensing is moving toward practical evaluation, with kits planned starting June after an April 15, 2026 announcement by UltraSense Systems.
- Vision–language–action (VLA) success for generalist robots depends on concrete architecture choices—backbone selection, VLA formulation, and the timing of modality fusion—highlighted in a Nature Machine Intelligence paper dated November 16, 2025.
- Multimodal tactile+vision manipulation approaches like TactileAloha (Tohoku University, September 3, 2025) show improved task success, reinforcing that tactile inputs must be more than optional features.
- The near-term strategy: validate tactile reliability first, then design VLA fusion so tactile signals can correct action decisions under uncertainty.