Future-State Predicting LSTM for Early Surgery Type Recognition

11/28/2018
by   Siddharth Kannan, et al.
0

This work presents a novel approach for the early recognition of the type of a laparoscopic surgery from its video. Early recognition algorithms can be beneficial to the development of 'smart' OR systems that can provide automatic context-aware assistance, and also enable quick database indexing. The task is however ridden with challenges specific to videos belonging to the domain of laparoscopy, such as high visual similarity across surgeries and large variations in video durations. To capture the spatio-temporal dependencies in these videos, we choose as our model a combination of a Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) network. We then propose two complementary approaches for improving early recognition performance. The first approach is a CNN fine-tuning method that encourages surgeries to be distinguished based on the initial frames of laparoscopic videos. The second approach, referred to as 'Future-State Predicting LSTM', trains an LSTM to predict information related to future frames, which helps in distinguishing between the different types of surgeries. We evaluate our approaches on a large dataset of 425 laparoscopic videos containing 9 types of surgeries (Laparo425), and achieve on average an accuracy of 75 minutes of a surgery. These results are quite promising from a practical standpoint and also encouraging for other types of image-guided surgeries.

READ FULL TEXT

page 1

page 5

page 9

research
02/06/2020

An Information-rich Sampling Technique over Spatio-Temporal CNN for Classification of Human Actions in Videos

We propose a novel scheme for human action recognition in videos, using ...
research
04/12/2021

Predicting the Accuracy of Early-est Earthquake Magnitude Estimates with an LSTM Neural Network: A Preliminary Analysis

This report presents a preliminary analysis of an LSTM neural network de...
research
09/19/2017

Predicting Video Saliency with Object-to-Motion CNN and Two-layer Convolutional LSTM

Over the past few years, deep neural networks (DNNs) have exhibited grea...
research
07/15/2020

Proof of Concept: Automatic Type Recognition

The type used to print an early modern book can give scholars valuable i...
research
02/19/2019

Predicting tongue motion in unlabeled ultrasound videos using convolutional LSTM neural network

A challenge in speech production research is to predict future tongue mo...
research
07/19/2017

Recognizing and Curating Photo Albums via Event-Specific Image Importance

Automatic organization of personal photos is a problem with many real wo...
research
09/05/2020

Player Identification in Hockey Broadcast Videos

We present a deep recurrent convolutional neural network (CNN) approach ...

Please sign up or login with your details

Forgot password? Click here to reset