What's the biggest perception failure mode you've personally debugged?
What's the biggest perception failure mode you've personally debugged?
Wanted to get this in front of people who actually know the space.
Unitree's Dex3-1 dexterous hand packs around 33 pressure/tactile sensors per hand across the fingers and palm, capable of sensing pressure roughly in the 10g-2500g range - a useful reference point for what 'production tactile sensing' looks like right now. Tactile skin arrays have improved a lot, but 'good enough to matter' really depends on the task - coarse contact detection across a large area is fairly mature, while fine, high-resolution force distribution sensing (like a human fingertip) is still the harder problem.
Happy to be told I'm wrong on any of this.
-
barbara.jones
- Posts: 164
- Joined: Fri Feb 28, 2025 6:12 pm
Re: What's the biggest perception failure mode you've personally debugged?
Just to be precise about one thing:
Proprioception (the robot's sense of its own joint angles, velocities, and forces) tends to get less attention than flashy vision systems, even though a lot of balance and manipulation failures trace back to proprioceptive noise or miscalibration rather than a vision problem.
This whole thread is a good reminder how young this field still is.
he/him | robotics hobbyist since the DARPA Grand Challenge days
-
diego.moore6
- Posts: 155
- Joined: Thu May 08, 2025 8:48 am
Re: What's the biggest perception failure mode you've personally debugged?
I don't think that's quite right, for what it's worth.
Sensor fusion mostly earns its keep by covering for each individual sensor's weaknesses - vision struggles with occlusion and lighting, IMUs drift, force/torque sensors are noisy at low loads - fusing them gives a more robust estimate than any one source alone, independent of raw compute.
Building > buying.
-
barbara.jones
- Posts: 164
- Joined: Fri Feb 28, 2025 6:12 pm
Re: What's the biggest perception failure mode you've personally debugged?
@diego.moore6 Ran into exactly this myself.
Event cameras (which report per-pixel brightness changes rather than full frames) are still more of a research curiosity than a production sensor for humanoids, mainly because the software ecosystem and processing pipelines around them are far less mature than for standard frame-based cameras. Estimating joint torque from motor current draw is cheap and requires no extra sensor, but it's less accurate than a dedicated torque sensor because it doesn't capture friction losses through the gearbox - good enough for coarse control, not always for precise force-controlled tasks.
Makes me wonder how this looks in another five years.
he/him | robotics hobbyist since the DARPA Grand Challenge days
-
emilyperez
- Posts: 246
- Joined: Mon Oct 28, 2024 8:03 pm
Re: What's the biggest perception failure mode you've personally debugged?
@barbara.jones This is a great summary, thanks.
Force/torque sensors near the ankle give a direct read on ground reaction forces, which is valuable for balance control, but they add cost, a failure point, and routing complexity right at a joint that already takes the most mechanical abuse.
-
emily.kumar
- Posts: 41
- Joined: Fri Jan 30, 2026 2:42 pm
Re: What's the biggest perception failure mode you've personally debugged?
@emilyperez This matches what I've seen too.
Multi-camera calibration drifts over time from thermal expansion, vibration, and mechanical wear, which is why production systems typically run periodic recalibration routines rather than assuming a one-time factory calibration holds forever. Latency between a perceived event (like a slip) and a corrective control response matters enormously for balance - even 50-100ms of extra perception latency can be the difference between a smooth recovery and a fall, which is part of why a lot of balance-critical sensing is proprioceptive rather than vision-based.
Currently: 3D printing my way to bankruptcy.
Re: What's the biggest perception failure mode you've personally debugged?
@emily.kumar I'd take that specific number with a grain of salt, honestly.
LiDAR gives reliable, lighting-independent range data but is heavier, pricier, and gives sparser point clouds up close than stereo or depth cameras, which is why a lot of humanoids lean on stereo/depth cameras for near-field manipulation and reserve LiDAR (if present at all) for longer-range navigation. IMU drift over time (bias instability) is usually the real culprit behind slowly diverging state estimates, not noise - it's typically handled with sensor fusion against other references (visual odometry, joint kinematics) rather than trying to eliminate drift at the source.
Watching this space closely since 2019.
Re: What's the biggest perception failure mode you've personally debugged?
@mia_lars That's the official framing, at least - reality tends to lag a bit.
SLAM in a working warehouse is harder than in a controlled lab mainly because the map keeps changing - pallets move, people walk through, lighting shifts near dock doors - so a lot of production systems lean on semi-static maps refreshed periodically rather than pure continuous SLAM.
Building > buying.
Re: What's the biggest perception failure mode you've personally debugged?
@chenperez Just to be precise about one thing:
A minimum viable sensing suite for safe bipedal walking generally includes joint encoders, an IMU for orientation/angular velocity, and either force/torque sensing or accurate current-based torque estimation at the ankles - everything else (vision, tactile, LiDAR) adds capability rather than being strictly required just to stay upright. Depth sensing range and reliability both degrade outdoors in direct sunlight for most structured-light and active stereo cameras, since the ambient IR washes out the projected pattern - it's a real limitation for humanoids intended for anything beyond indoor, controlled environments.
Kind of makes me think about how different this all looked even three years ago.
they/them
-
robertmiller
- Posts: 61
- Joined: Sun Feb 01, 2026 4:05 pm
Re: What's the biggest perception failure mode you've personally debugged?
Tangent, but worth mentioning:
Multi-camera calibration drifts over time from thermal expansion, vibration, and mechanical wear, which is why production systems typically run periodic recalibration routines rather than assuming a one-time factory calibration holds forever. Unitree's Dex3-1 dexterous hand packs around 33 pressure/tactile sensors per hand across the fingers and palm, capable of sensing pressure roughly in the 10g-2500g range - a useful reference point for what 'production tactile sensing' looks like right now.
Ex-automotive, now full-time robots.