Learning to see like children: proof of concept

08/11/2014
by   Marco Gori, et al.
0

In the last few years we have seen a growing interest in machine learning approaches to computer vision and, especially, to semantic labeling. Nowadays state of the art systems use deep learning on millions of labeled images with very successful results on benchmarks, though it is unlikely to expect similar results in unrestricted visual environments. Most learning schemes essentially ignore the inherent sequential structure of videos: this might be a critical issue, since any visual recognition process is remarkably more complex when shuffling video frames. Based on this remark, we propose a re-foundation of the communication protocol between visual agents and the environment, which is referred to as learning to see like children. Like for human interaction, visual concepts are acquired by the agents solely by processing their own visual stream along with human supervisions on selected pixels. We give a proof of concept that remarkable semantic labeling can emerge within this protocol by using only a few supervised examples. This is made possible by exploiting a constraint of motion coherent labeling that virtually offers tons of supervisions. Additional visual constraints, including those associated with object supervisions, are used within the context of learning from constraints. The framework is extended in the direction of lifelong learning, so as our visual agents live in their own visual environment without distinguishing learning and test set. Learning takes place in deep architectures under a progressive developmental scheme. In order to evaluate our Developmental Visual Agents (DVAs), in addition to classic benchmarks, we open the doors of our lab, allowing people to evaluate DVAs by crowd-sourcing. Such assessment mechanism might result in a paradigm shift in methodologies and algorithms for computer vision, encouraging truly novel solutions within the proposed framework.

READ FULL TEXT

page 5

page 16

page 17

page 18

page 20

page 21

research
07/07/2022

Predicting Word Learning in Children from the Performance of Computer Vision Systems

For human children as well as machine learning systems, a key challenge ...
research
01/16/2018

Convolutional Networks in Visual Environments

The puzzle of computer vision might find new challenging solutions when ...
research
08/09/2021

Probabilistic annotations for protocol models

We describe how a probabilistic Hoare logic with localities can be used ...
research
05/16/2017

Cooperative Learning with Visual Attributes

Learning paradigms involving varying levels of supervision have received...
research
10/14/2020

Vision-Aided Radio: User Identity Match in Radio and Video Domains Using Machine Learning

5G is designed to be an essential enabler and a leading infrastructure p...
research
09/01/2019

Learning Visual Features Under Motion Invariance

Humans are continuously exposed to a stream of visual data with a natura...
research
12/11/2020

Risk returns around FOMC press conferences: a novel perspective from computer vision

I propose a new tool to characterize the resolution of uncertainty aroun...

Please sign up or login with your details

Forgot password? Click here to reset