Diving Deeper: Taking a closer look into the scientific content of the paper on tracking of deformable objects
Looking into the key innovations of the paper on Tracking of Deformable Objects
Related Paper: RGB-D Tracking and Optimal Perception of Deformable Objects
The tracking of deformable objects is a more challenging task than tracking rigid objects. The reason is that as force is exerted onto a deformable object it can move in more unpredictable ways. There is also an increased possibility of occlusions as the object, for example a rope, overlaps on itself, which make tracking difficult. Having a great understanding of deformable objects will be necessary in the field of robotics with various important applications such as for medical surgeries or automating routine tasks such as folding laundry.
The main challenge in terms of using supervised ML approaches is that there is not a lot of labelled dataset available for learning. The authors are proposing a learning free method by essentially preprocessing the RBG and Depth data into point clouds and performing a graph-based super voxel tracking approach.
Key Scientific Contributions
- No learning required: Since it is a computer vision based approach no Deep learning / training is required
- Graph-Based Super voxel Tracking: The objects are tracked in 3D space as supervoxel graphs
- Real-time Camera optimisation: The authors propose a method of constantly achieving best camera view point onto the object by dynamically computing and moving camera
How the method works
I will explain the methods overall first to the get the full picture and dive deeper into some of the technical terms that are domain specific
The input into the pipeline is a video that contains RGB and Depth information. It is captured by an Intel RealSense D435 camera.
Each RGB+Depth Frame gets processed in this manner:
- Use modified ORB-SLAM2 to estimate camera pose
- Create a dense map using RGB-Depth info and downsample to point clouds
- Perform supervoxel segmentation and object selection using LCCP
- Keep track of the selected object state and update per Frame
- Solve Next Best View problem to reposition camera. This includes avoids occlusions of the object
The entire pipeline is illustrated here
Info on the technical terms
- ORB-SLAM2: SLAM stands for simultaneous localisation and mapping and it is a common method used in robotics that allow a robot/camera to map an environment and track its own position within that map. ORB-SLAM2 is an updated version that works with RGB-D cameras and achieves real-time performance even if the camera is moving. The authors of this paper modified ORB-SLAM2 to also generate a Dense Map of the object instead of the default sparse map. this will later be downsampled and used the segmentation
- Super voxelization: It is a 3D version of super pixels which is just a group of pixels with similar color or geometry or texture. In this case instead of pixels it is a group of point clouds with similar patterns
- LCCP: It stands for locally convex connected patches and it is a 3d segmentation method that groups points into meaningful objects. The user can then select the object, like a balloon, and the state of that specific object is then tracked per frame
Experimental Setup and Results
The authors experimented on 4 different objects: a box, a show, a balloon and a cloth. The authors mention that the relevant evaluation is essentially qualitative. So whether the tracking works well is observed with snapshots of the tracking. The authors also collect some relevant quantitative measures such as number of supervoxels belonging to the object over time compared to total supervoxels as well as processing time etc..
The experiment is performed by moving the objects around but also applying force and in the case of the balloon, they are popping the balloon.
Here are some of the results:
Here with the shoe you can see for example that the tracking works really well and the LCCP method is nicely separating the shoe for the scene. At the bottom you also see the camera adjusting its angle to always maintain the best view.
Here is also the results of the object tracking of a cloth piece. This is a much more deformable object and the tracking works relatively well.
The authors found that in the case of the cloth there were many more number of supervoxels in the object which directly also translates to longer computation time for object segmentation and tracking. This can also be seen in the following illustration as its the number of supervoxels in object vs processing time
Value of the work and How it fits into the field
The value of the work lies in the practical and general applicability. Since there is no learning involved the pipeline or parts of it can be used and adapted as needed. For example for my VT project, I will look into this LCCP segmentation as well as the graph based tracking approach but not put too much focus on the Best Next View problem as I am dealing with a static camera.
While the evaluation is not without its limitations as it mainly relies on qualitative metrics - the real-time tracking of deformable objects using graph based approach segmentation approach is definitely a useful finding for the further progress in the field of robotics and deformable objects.
There have been several studies done on tracking of deformable objects such as 'Tracking Deformable Objects by Point Cloud' by Schulman et al. The study separates itself from other tasks by also including camera view point optimisation and combining it with a learning- free tracking method and this is particularly useful in robotic manipulation of tasks since getting a clear view of the object is just as important as tracking and handling it.
Conclusion
Overall, this paper contributes a significant solution to a hard problem in robotics and computer vision and it does so without requiring a heavy load of training data or deep networks. That is pretty valuable step forward




Comments
Post a Comment