We present In-Context Imitation Learning with Visual Reasoning (ICLR), a framework that augments demonstration prompts with structured visual reasoning traces representing anticipated future robot trajectories in image space
Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation
Shaocong Xu, Songlin Wei, Qizhe Wei, Zheng Geng, Hong Li, Licheng Shen, Qianpu Sun, Shu Han, Bin Ma, Bohan Li, Chongjie Ye, Yuhang Zheng, Nan Wang, Saining Zhang, and Hao Zhao†
“Diffusion knows transparency.” Generative video priors can be repurposed, efficiently and label-free, into robust, temporally coherent perception for challenging real-world manipulation.
GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
We present Uni-NaVid, the first video-based vision-language-action (VLA) model designed to unify diverse embodied navigation tasks and enable seamless navigation for mixed long-horizon tasks in unseen real-world environments.
RoboHanger: Learning Generalizable Robotic Hanger Insertion for Diverse Garments
we introduced a large-scale part-centric dataset for articulated object manipulation that features both photo-realistic material randomizations and detailed annotations of part-oriented, scene-level actionable interaction poses.
D3RoMa: Disparity Diffusion-based Depth Sensing for Material-Agnostic Robotic Manipulation
In this work, we introduce a demonstration-free hierarchical planning approach capable of tackling intricate long-horizon tasks without necessitating any training
Open6DOR: Benchmarking Open-instruction 6-DoF Object Rearrangement and A VLM-based Approach
We present Open6DOR, a challenging and comprehensive benchmark for open-instruction 6-DoF object rearrangement tasks. Following this, we propose a zero-shot and robust method, Open6DORGPT, which proves effective in demanding simulation environments and real-world scenarios.
SAGE🌿: Bridging Semantic and Actionable Parts for Generalizable Manipulation of Articulated Objects
We propose an independence-assumption-free probabilistic neural radiance field based on Flow-GAN. By combining the generative capability of adversarial learning and the powerful expressivity of normalizing flow, our method explicitly models the density-radiance distribution of the whole scene.
3D Object Aided Self-Supervised Monocular Depth Estimation
Songlin Wei, Guodong Chen, Wenzheng Chi, Zhenhua Wang and Lining Sun
Self-supervised depth estimation methods rely on static world assumption, which produce inaccurate depths of dynamic objects.
In this work, we propose to address dynamic object movements through monocular 3D object detection.
Object Clustering with Dirichlet Process Mixture Model for Data Association in Monocular SLAM
Songlin Wei, Guodong Chen, Wenzheng Chi, Zhenhua Wang and Lining Sun
We propose a novel data association method for cuboid landmarks based on Dirichlet Process Mixture Model. By jointly considering object class, position, and size, our method can perform data association robustly.