A vision system can estimate an object’s position, orientation and visible shape before a robot reaches it. That is enough for many structured tasks that pick and place objects. It is not enough to reveal whether a grasp is beginning to slip, a connector is binding, a flexible part has folded or a tool has made uneven contact. In those moments, the most important information is produced by the interaction itself.

Research on multisensory dexterity highlighted by UC Berkeley EECS in August 2025 points toward a broader approach: give robotic systems a shared understanding of vision, touch and proprioception. Each sensing mode observes a different part of the physical state. Together, they can make manipulation less like following a precomputed path and more like continuously responding to the world.

Three senses, three different questions

Vision provides context across a wide area. It can identify the target, estimate free space, locate obstacles and guide the hand toward an initial contact pose. Its advantage is reach: the robot can observe before touching. Its limitation is that fingers, tools and the object itself can occlude the exact interaction that matters.

Touch begins where vision becomes uncertain. Tactile sensors can reveal where contact occurs, how pressure is distributed and whether the surface is moving relative to the gripper. Depending on the sensor, they may also capture local geometry, vibration or shear. These signals help distinguish a stable grasp from an object that is rotating or about to slip.

Proprioception describes the robot’s own body. Joint positions, velocities, motor currents and estimated forces show how the mechanism is moving and resisting motion. This makes it possible to detect that an insertion has met unexpected resistance or that the commanded trajectory and physical response no longer agree.

Design principle

Do not ask one sensor to explain the whole interaction. Fuse complementary evidence around the physical state the task actually needs.

The important event may begin at contact

Many manipulation tasks look simple until the gripper touches the object. Inserting a part requires alignment with a tolerance that may be smaller than the uncertainty in a visual estimate. Twisting a cap changes friction and torque through the motion. Handling cloth, cable or soft packaging changes the object’s shape. Gripping a smooth component requires enough force to prevent slip without marking or damaging it.

In each case, a fixed motion without feedback assumes that the estimate before contact is correct. A multisensory policy can instead treat contact as new evidence. A pressure shift can trigger a grip adjustment. Unexpected joint effort can stop an insertion and initiate a small search motion. Incipient slip can cause the hand to increase force or support the object differently.

Fusion needs a shared physical frame

Adding sensors does not automatically create understanding. Camera images, tactile arrays and signals about robot state arrive at different frequencies, with different delays and noise characteristics. They must be synchronized closely enough to associate a visual motion with the contact event and joint response that produced it.

Geometry matters too. The system needs to relate a tactile location on a fingertip to the hand, tool and camera coordinate frames. Calibration drift, compliant fingertips and changing tools can all weaken this relationship. The fusion layer should preserve uncertainty rather than turning every estimate into false precision.

A practical representation may combine object pose, hand pose, contact location, normal and shear estimates, slip probability and task phase. The right representation depends on the task. A system that picks objects from bins, a connector insertion cell and a handler for deformable material do not need identical sensory detail.

Dexterity depends on a closed control loop

Multisensory perception creates value only when the robot can act on it at the required speed. Contact events can evolve faster than a remote planning service can respond. Force, slip and motion correction at the control level therefore belongs close to the robot, with bounded latency and deterministic safety behavior. Models at a higher level can select strategies and interpret task context without owning every control cycle.

This suggests a layered architecture. Fast local controllers enforce motion and force limits. A learned manipulation policy combines recent sensory history to choose bounded adjustments. A task layer tracks the goal, product variant and process state. Safety functions remain independent and can stop motion regardless of model output.

→Define which hidden physical state must be estimated, such as slip, alignment, deformation or contact force.
→Synchronize and calibrate sensor streams across the actual operating envelope.
→Keep rapid contact responses close to the robot and inside explicit force and motion limits.
→Validate with variation in surface, geometry, lighting, wear and placement rather than one scripted demonstration.

Better sensing changes the data problem

Multisensory systems also create a richer learning signal. A successful trajectory can include not only video and final task outcome, but the sequence of contacts, force changes, joint motion and corrective actions that led there. Failed attempts become more informative because teams can inspect where perception and physical response diverged.

The data pipeline must preserve those relationships. Logs need common timestamps, robot and tool configuration, sensor calibration versions, task phase and outcome. If a tactile sensor is replaced or a gripper pad wears, that change may alter the signal distribution even when the task appears unchanged. Monitoring sensor health is therefore part of monitoring model performance.

The next leap may be sensory

Larger models can improve generalization and planning, but work involving contact exposes a basic constraint: a model cannot reason over physical evidence the system never measures. Cameras will remain essential because they provide awareness of the scene. Touch and proprioception add the local evidence needed once the robot and world begin to influence each other.

The practical opportunity is not to imitate every aspect of human sensation. It is to identify the uncertainty that prevents a task from being reliable, instrument that interaction and close the loop. For insertion, twisting, deforming, gripping and adjustment, better sensing may advance dexterity as much as a better model.

Source and scope

This perspective is informed by UC Berkeley EECS research on multisensory dexterity published in August 2025. It develops the engineering implications of combining vision, touch and proprioception; it does not claim that every architecture or implementation detail described here is part of a single Berkeley system.