Task-oriented grasping from point clouds

Finding a good grasping region for a task directly from a partial point cloud, using a neural network trained on the grasp metric instead of labeled grasps.

Code Project page

My PhD research focused on task-oriented grasping from sensor data. The problem has three parts: defining the task mathematically, using that definition to say what makes a grasp good for the task, and working with noisy sensor data in the form of partial point clouds from an RGB-D camera.

We define a task as a constant screw motion, or a sequence of them, to be imparted to the object after grasping. The task-dependent grasp metric evaluates whether a pair of contact locations can impart that motion. Here we solve the inverse problem: given a partial point cloud of the object and a task screw, find a good region to grasp.

We use the simplest geometric representation of the point cloud, a bounding box, and train a neural network to predict the grasp metric on a cuboid. The predicted grasping region on the bounding box is then used to compute an antipodal grasp on the actual object. The approach does not use any manually labeled data or grasping simulator, and because the task is a screw motion, it couples directly with screw linear interpolation-based motion planners. Follow-up work (ICRA 2025) extends this with point cloud decomposition.

Conditioner bottle: simulated point cloud and task screw, the pure rotation about the screw, and the predicted grasping region (yellow).

For more results, see the lab project page.

Papers