I work on learning manipulation policies that run on real hardware in contact-rich industrial settings: cable routing, connector insertion and precision assembly, tasks that are still largely done by hand in factories and datacenters.
Diffusion policies. In the Industrial Dexterity Benchmark (arXiv:2607.14021) we introduce AG-iDP3, a multimodal diffusion-policy framework that fuses wrist and scene RGB, scene point clouds, joint positions and wrist-frame wrench data. On the datacenter cable-manipulation board, the best multimodal configuration reaches a 78% combined grasp-and-insert success rate against 36% for a single-camera RGB diffusion-policy baseline, using roughly 100 teleoperated demonstrations per task phase. The benchmark boards, the DAG-ROS imitation-learning framework and the scoring rubrics are open source.
Vision-language-action models. I also work on force-conditioned extensions of vision-language-action models, integrating six-axis end-effector wrench into the pretrained vision-language-action stack so that a policy can respond to contact rather than to vision alone.
Perception for manipulation. Earlier work built an end-to-end RGB-D perception pipeline combining semantic segmentation, human tracking and semantics-guided point-cloud fusion, exporting structured scenes in OpenUSD for downstream robotics and simulation (arXiv:2410.17988).