Page 1 of 2

LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Wed Jan 15, 2025 6:59 am
by matthew43
Curious what people here think about this. LiDAR gives reliable, lighting-independent range data but is heavier, pricier, and gives sparser point clouds up close than stereo or depth cameras, which is why a lot of humanoids lean on stereo/depth cameras for near-field manipulation and reserve LiDAR (if present at all) for longer-range navigation. Unitree's Dex3-1 dexterous hand packs around 33 pressure/tactile sensors per hand across the fingers and palm, capable of sensing pressure roughly in the 10g-2500g range - a useful reference point for what 'production tactile sensing' looks like right now. Tactile skin arrays have improved a lot, but 'good enough to matter' really depends on the task - coarse contact detection across a large area is fairly mature, while fine, high-resolution force distribution sensing (like a human fingertip) is still the harder problem. Open to being corrected on the specifics.

Re: LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Wed Jan 15, 2025 7:25 am
by choi98
This is exactly the kind of context I was looking for. Multi-camera calibration drifts over time from thermal expansion, vibration, and mechanical wear, which is why production systems typically run periodic recalibration routines rather than assuming a one-time factory calibration holds forever. Totally unrelated but has anyone else noticed how fast component costs are dropping this year.

Re: LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Wed Jan 15, 2025 9:36 am
by jonathan.rao1
I'd take that specific number with a grain of salt, honestly. Event cameras (which report per-pixel brightness changes rather than full frames) are still more of a research curiosity than a production sensor for humanoids, mainly because the software ecosystem and processing pipelines around them are far less mature than for standard frame-based cameras. SLAM in a working warehouse is harder than in a controlled lab mainly because the map keeps changing - pallets move, people walk through, lighting shifts near dock doors - so a lot of production systems lean on semi-static maps refreshed periodically rather than pure continuous SLAM.

Re: LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Wed Jan 15, 2025 11:06 am
by matthew43
@jonathan.rao1 I see it a little differently. Proprioception (the robot's sense of its own joint angles, velocities, and forces) tends to get less attention than flashy vision systems, even though a lot of balance and manipulation failures trace back to proprioceptive noise or miscalibration rather than a vision problem.

Re: LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Wed Jan 15, 2025 1:55 pm
by park44
@matthew43 Here's the relevant bit as far as I understand it: Force/torque sensors near the ankle give a direct read on ground reaction forces, which is valuable for balance control, but they add cost, a failure point, and routing complexity right at a joint that already takes the most mechanical abuse.

Re: LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Wed Jan 15, 2025 3:56 pm
by nicole57
One nitpick - Depth sensing range and reliability both degrade outdoors in direct sunlight for most structured-light and active stereo cameras, since the ambient IR washes out the projected pattern - it's a real limitation for humanoids intended for anything beyond indoor, controlled environments. Estimating joint torque from motor current draw is cheap and requires no extra sensor, but it's less accurate than a dedicated torque sensor because it doesn't capture friction losses through the gearbox - good enough for coarse control, not always for precise force-controlled tasks.

Re: LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Thu Jan 16, 2025 2:46 pm
by jonathan.rao1
+1 to this. Worth adding: Vibration is one of the most underrated sources of noisy IMU and tactile readings - mounting matters as much as sensor quality, and a poorly isolated mount can add more noise than the sensor's own datasheet specs would suggest.

Re: LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Fri Jan 17, 2025 4:15 pm
by noah_pate
@jonathan.rao1 I'll believe the stronger version of that claim when it's independently verified. Sensor fusion mostly earns its keep by covering for each individual sensor's weaknesses - vision struggles with occlusion and lighting, IMUs drift, force/torque sensors are noisy at low loads - fusing them gives a more robust estimate than any one source alone, independent of raw compute. Latency between a perceived event (like a slip) and a corrective control response matters enormously for balance - even 50-100ms of extra perception latency can be the difference between a smooth recovery and a fall, which is part of why a lot of balance-critical sensing is proprioceptive rather than vision-based. Kind of makes me think about how different this all looked even three years ago.

Re: LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Fri Jan 17, 2025 11:47 pm
by olga_lind
@noah_pate This is a great summary, thanks. IMU drift over time (bias instability) is usually the real culprit behind slowly diverging state estimates, not noise - it's typically handled with sensor fusion against other references (visual odometry, joint kinematics) rather than trying to eliminate drift at the source.

Re: LiDAR vs stereo depth cameras for humanoid navigation - what's actually winning?

Posted: Sun Jan 19, 2025 5:20 am
by deborah59
@olga_lind I'd take that specific number with a grain of salt, honestly. Force/torque sensors near the ankle give a direct read on ground reaction forces, which is valuable for balance control, but they add cost, a failure point, and routing complexity right at a joint that already takes the most mechanical abuse. Event cameras (which report per-pixel brightness changes rather than full frames) are still more of a research curiosity than a production sensor for humanoids, mainly because the software ecosystem and processing pipelines around them are far less mature than for standard frame-based cameras.