Object Detection for Understanding Assembly Instruction Using Context-aware Data Augmentation and Cascade Mask R-CNN

01/07/2021
by   Joosoon Lee, et al.
0

Understanding assembly instruction has the potential to enhance the robot s task planning ability and enables advanced robotic applications. To recognize the key components from the 2D assembly instruction image, We mainly focus on segmenting the speech bubble area, which contains lots of information about instructions. For this, We applied Cascade Mask R-CNN and developed a context-aware data augmentation scheme for speech bubble segmentation, which randomly combines images cuts by considering the context of assembly instructions. We showed that the proposed augmentation scheme achieves a better segmentation performance compared to the existing augmentation algorithm by increasing the diversity of trainable data while considering the distribution of components locations. Also, we showed that deep learning can be useful to understand assembly instruction by detecting the essential objects in the assembly instruction, such as tools and parts.

READ FULL TEXT

page 2

page 3

research
06/01/2021

Assembly Planning by Recognizing a Graphical Instruction Manual

This paper proposes a robot assembly planning method by automatically re...
research
11/20/2022

Context-Aware Data Augmentation for LIDAR 3D Object Detection

For 3D object detection, labeling lidar point cloud is difficult, so dat...
research
07/06/2023

BrickPal: Augmented Reality-based Assembly Instructions for Brick Models

The assembly instruction is a mandatory component of Lego-like brick set...
research
08/14/2023

Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents

Accomplishing household tasks requires to plan step-by-step actions cons...
research
04/19/2018

Defining Pathway Assembly and Exploring its Applications

How do we estimate the probability of an abundant objects' formation, wi...
research
09/18/2023

Instruction-Following Speech Recognition

Conventional end-to-end Automatic Speech Recognition (ASR) models primar...
research
09/17/2022

Bilevel Optimization for Just-in-Time Robotic Kitting and Delivery via Adaptive Task Segmentation and Scheduling

Kitting refers to the task of preparing and grouping necessary parts and...

Please sign up or login with your details

Forgot password? Click here to reset