InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

09/07/2023
by   Zigang Geng, et al.
0

We present InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions. Unlike existing approaches that integrate prior knowledge and pre-define the output space (e.g., categories and coordinates) for each vision task, we cast diverse vision tasks into a human-intuitive image-manipulating process whose output space is a flexible and interactive pixel space. Concretely, the model is built upon the diffusion process and is trained to predict pixels according to user instructions, such as encircling the man's left shoulder in red or applying a blue mask to the left car. InstructDiffusion could handle a variety of vision tasks, including understanding tasks (such as segmentation and keypoint detection) and generative tasks (such as editing and enhancement). It even exhibits the ability to handle unseen tasks and outperforms prior methods on novel datasets. This represents a significant step towards a generalist modeling interface for vision tasks, advancing artificial general intelligence in the field of computer vision.

READ FULL TEXT

page 1

page 7

page 9

page 10

page 11

page 12

research
06/15/2022

A Unified Sequence Interface for Vision Tasks

While language tasks are naturally expressed in a single, unified, model...
research
09/20/2023

RMT: Retentive Networks Meet Vision Transformers

Transformer first appears in the field of natural language processing an...
research
12/05/2022

Images Speak in Images: A Generalist Painter for In-Context Visual Learning

In-context learning, as a new paradigm in NLP, allows the model to rapid...
research
05/20/2022

UViM: A Unified Modeling Approach for Vision with Learned Guiding Codes

We introduce UViM, a unified approach capable of modeling a wide range o...
research
10/04/2017

Semantic 3D Reconstruction with Finite Element Bases

We propose a novel framework for the discretisation of multi-label probl...
research
12/06/2019

Gaussian Process Priors for View-Aware Inference

We derive a principled framework for encoding prior knowledge of informa...
research
08/19/2021

Neural TMDlayer: Modeling Instantaneous flow of features via SDE Generators

We study how stochastic differential equation (SDE) based ideas can insp...

Please sign up or login with your details

Forgot password? Click here to reset