Anyone using imitation learning purely from human video for grasp policies?

End effectors, dexterous hands, tendon drives, tactile fingertips, grasp planning, and teleoperation.
emily.kumar
Posts: 41
Joined: Fri Jan 30, 2026 2:42 pm

Anyone using imitation learning purely from human video for grasp policies?

Post by emily.kumar »

Genuinely split on this one, wanted outside opinions. Cable routing through a wrist joint with multiple degrees of freedom is a genuinely tricky mechanical design problem - tendons and wiring both need enough slack to avoid binding through the full range of motion without tangling or fraying over thousands of cycles. Vision-based grasp confidence estimation (predicting success before attempting a grasp) and tactile-based confirmation (confirming after contact) are complementary rather than competing - vision helps you choose a grasp, tactile tells you if it actually worked. Feel free to tell me I'm overthinking this.
Currently: 3D printing my way to bankruptcy.
emilyperez
Posts: 246
Joined: Mon Oct 28, 2024 8:03 pm

Re: Anyone using imitation learning purely from human video for grasp policies?

Post by emilyperez »

Minor factual note: A lot of the manipulation shown in production demos still leans heavily on teleoperation, particularly for anything involving fine force control or novel objects - autonomous grasp success rates on genuinely unstructured, previously-unseen clutter are still well below what teleoperation can achieve.
harmonicjen60
Posts: 64
Joined: Sat Feb 07, 2026 7:12 am

Re: Anyone using imitation learning purely from human video for grasp policies?

Post by harmonicjen60 »

@emilyperez Short answer: Underactuated hands (fewer actuators than joints, using mechanical coupling to shape the grasp) are a reasonable engineering compromise for robust power grasps on a budget, but they generally can't do fine in-hand manipulation the way a fully actuated hand can.
barbara50
Posts: 178
Joined: Thu Dec 19, 2024 12:19 pm

Re: Anyone using imitation learning purely from human video for grasp policies?

Post by barbara50 »

@harmonicjen60 Related question - Figure's fourth-generation Dexterous Hand (on Figure 02/03) reportedly offers 16 degrees of freedom per hand with sensors integrated into each finger, aimed at fine force control tasks like handling small electronic components without crushing them.
Opinions my own, not my employer's.
mary.taylor6
Posts: 82
Joined: Tue Nov 11, 2025 10:51 pm

Re: Anyone using imitation learning purely from human video for grasp policies?

Post by mary.taylor6 »

Just to be precise about one thing: Payload-to-hand-weight ratio varies a lot across current dexterous hands, and it's a meaningful tradeoff - more DoF and finer sensing generally means more actuators and mass in the hand itself, which eats into the arm's usable payload budget. Imitation learning from human demonstration video (without robot teleoperation data) is an appealing way to scale up training data cheaply, but it runs into the embodiment gap - human hand kinematics and force profiles don't map directly onto a robot hand's very different mechanism.
dubois35
Posts: 280
Joined: Mon Sep 09, 2024 12:01 pm

Re: Anyone using imitation learning purely from human video for grasp policies?

Post by dubois35 »

@mary.taylor6 I can speak to this a bit. Wet, oily, or otherwise low-friction objects are still a genuine edge case for most current hands, since tactile sensing and grasp-force controllers are typically tuned and validated on dry, higher-friction test objects.
they/them
matthew43
Posts: 199
Joined: Wed Nov 06, 2024 9:18 am

Re: Anyone using imitation learning purely from human video for grasp policies?

Post by matthew43 »

@dubois35 Pretty much this. One thing to add: In-hand reorientation (repositioning a grasped object without setting it down) is one of the more advanced manipulation skills, requiring either a highly dexterous hand with enough DoF or clever use of gravity and controlled slipping - it's an active research area rather than a solved problem. Palm sensing gets less attention than fingertip sensing, but a lot of power grasps (holding a box, a tool handle) rely more on palm and lateral finger contact than fingertip contact, so under-sensing the palm can leave a real blind spot in grasp confidence.
he/him | robotics hobbyist since the DARPA Grand Challenge days
kim37
Posts: 109
Joined: Mon Jul 21, 2025 11:21 pm

Re: Anyone using imitation learning purely from human video for grasp policies?

Post by kim37 »

Slight correction, though the overall point stands: Compliant wrists that absorb impact during a bad approach or misjudged contact reduce mechanical stress on the whole arm, which matters a lot for long-term reliability even though it's a less visible feature than the hand itself.
karen.chen3
Posts: 189
Joined: Mon Mar 10, 2025 1:30 pm

Re: Anyone using imitation learning purely from human video for grasp policies?

Post by karen.chen3 »

Slight correction, though the overall point stands: Object occlusion by the robot's own hand during the final approach to a grasp is a common and annoying perception problem - the closer the hand gets to a good grasp position, the more it blocks the camera's view of exactly what it's about to grab. Tendon-driven fingers let you move the heavier actuators back into the palm or forearm, keeping the fingers themselves light and fast, but they introduce cable routing, tensioning, and long-term wear problems that direct-actuated fingers don't have.
they/them
carlossanchez
Posts: 155
Joined: Mon Feb 17, 2025 1:44 am

Re: Anyone using imitation learning purely from human video for grasp policies?

Post by carlossanchez »

Yeah, this tracks with what I've read as well. Grasp planning for deformable or non-rigid objects (bags, cables, cloth) remains one of the genuinely unsolved problems in manipulation - rigid-body grasp models simply don't capture how the object will behave once contact starts.
"Torque is a lifestyle."
Post Reply