Anyone tried fusing tactile and vision for grasp confidence estimation?

IMUs, force/torque sensors, depth cameras, LiDAR, tactile skin, SLAM, and state estimation.
vyoung
Posts: 51
Joined: Thu Mar 26, 2026 12:05 pm

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by vyoung »

@pierregreen Counterpoint: IMU drift over time (bias instability) is usually the real culprit behind slowly diverging state estimates, not noise - it's typically handled with sensor fusion against other references (visual odometry, joint kinematics) rather than trying to eliminate drift at the source. Event cameras (which report per-pixel brightness changes rather than full frames) are still more of a research curiosity than a production sensor for humanoids, mainly because the software ecosystem and processing pipelines around them are far less mature than for standard frame-based cameras.
zoeanderson
Posts: 243
Joined: Sat Oct 26, 2024 2:39 am

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by zoeanderson »

@vyoung One nitpick - Estimating joint torque from motor current draw is cheap and requires no extra sensor, but it's less accurate than a dedicated torque sensor because it doesn't capture friction losses through the gearbox - good enough for coarse control, not always for precise force-controlled tasks.
jessica_faro
Posts: 95
Joined: Sat Oct 11, 2025 5:26 am

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by jessica_faro »

@zoeanderson From what I've seen: LiDAR gives reliable, lighting-independent range data but is heavier, pricier, and gives sparser point clouds up close than stereo or depth cameras, which is why a lot of humanoids lean on stereo/depth cameras for near-field manipulation and reserve LiDAR (if present at all) for longer-range navigation. Latency between a perceived event (like a slip) and a corrective control response matters enormously for balance - even 50-100ms of extra perception latency can be the difference between a smooth recovery and a fall, which is part of why a lot of balance-critical sensing is proprioceptive rather than vision-based.
they/them
williams84
Posts: 237
Joined: Sat Sep 28, 2024 8:50 am

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by williams84 »

@jessica_faro Speaking from personal experience here, Sensor fusion mostly earns its keep by covering for each individual sensor's weaknesses - vision struggles with occlusion and lighting, IMUs drift, force/torque sensors are noisy at low loads - fusing them gives a more robust estimate than any one source alone, independent of raw compute.
"The best actuator is the one that doesn't overheat."
lbianchi
Posts: 81
Joined: Mon Sep 15, 2025 6:56 pm

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by lbianchi »

@williams84 Appreciate the detailed answer. Tactile skin arrays have improved a lot, but 'good enough to matter' really depends on the task - coarse contact detection across a large area is fairly mature, while fine, high-resolution force distribution sensing (like a human fingertip) is still the harder problem.
karen_kim
Posts: 116
Joined: Wed Jun 25, 2025 11:18 pm

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by karen_kim »

@lbianchi Small correction on one detail: Sensor fusion mostly earns its keep by covering for each individual sensor's weaknesses - vision struggles with occlusion and lighting, IMUs drift, force/torque sensors are noisy at low loads - fusing them gives a more robust estimate than any one source alone, independent of raw compute. Multi-camera calibration drifts over time from thermal expansion, vibration, and mechanical wear, which is why production systems typically run periodic recalibration routines rather than assuming a one-time factory calibration holds forever.
betty.king
Posts: 87
Joined: Sun Sep 14, 2025 8:37 am

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by betty.king »

Slightly off-topic, but related: Unitree's Dex3-1 dexterous hand packs around 33 pressure/tactile sensors per hand across the fingers and palm, capable of sensing pressure roughly in the 10g-2500g range - a useful reference point for what 'production tactile sensing' looks like right now.
Watching this space closely since 2019.
ethan.lewis5
Posts: 65
Joined: Sat Apr 04, 2026 7:35 pm

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by ethan.lewis5 »

@betty.king To answer this directly: Unitree's Dex3-1 dexterous hand packs around 33 pressure/tactile sensors per hand across the fingers and palm, capable of sensing pressure roughly in the 10g-2500g range - a useful reference point for what 'production tactile sensing' looks like right now. Sensor fusion mostly earns its keep by covering for each individual sensor's weaknesses - vision struggles with occlusion and lighting, IMUs drift, force/torque sensors are noisy at low loads - fusing them gives a more robust estimate than any one source alone, independent of raw compute. Makes me wonder how this looks in another five years.
Opinions my own, not my employer's.
ramirez77
Posts: 146
Joined: Sat Apr 05, 2025 9:39 pm

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by ramirez77 »

Sorry if this is a basic question, but Estimating joint torque from motor current draw is cheap and requires no extra sensor, but it's less accurate than a dedicated torque sensor because it doesn't capture friction losses through the gearbox - good enough for coarse control, not always for precise force-controlled tasks.
he/him | robotics hobbyist since the DARPA Grand Challenge days
pierregreen
Posts: 205
Joined: Thu Dec 12, 2024 11:01 am

Re: Anyone tried fusing tactile and vision for grasp confidence estimation?

Post by pierregreen »

@ramirez77 Short answer: Depth sensing range and reliability both degrade outdoors in direct sunlight for most structured-light and active stereo cameras, since the ambient IR washes out the projected pattern - it's a real limitation for humanoids intended for anything beyond indoor, controlled environments.
she/her
Post Reply