EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments

07/09/2021
by   Jacob Donley, et al.
16

Augmented Reality (AR) as a platform has the potential to facilitate the reduction of the cocktail party effect. Future AR headsets could potentially leverage information from an array of sensors spanning many different modalities. Training and testing signal processing and machine learning algorithms on tasks such as beam-forming and speech enhancement require high quality representative data. To the best of the author's knowledge, as of publication there are no available datasets that contain synchronized egocentric multi-channel audio and video with dynamic movement and conversations in a noisy environment. In this work, we describe, evaluate and release a dataset that contains over 5 hours of multi-modal data useful for training and testing algorithms for the application of improving conversations for an AR glasses wearer. We provide speech intelligibility, quality and signal-to-noise ratio improvement results for a baseline method and show improvements across all tested metrics. The dataset we are releasing contains AR glasses egocentric multi-channel microphone array audio, wide field-of-view RGB video, speech source pose, headset microphone audio, annotated voice activity, speech transcriptions, head bounding boxes, target of speech and source identification labels. We have created and are releasing this dataset to facilitate research in multi-modal AR solutions to the cocktail party problem.

READ FULL TEXT

page 1

page 4

page 6

research
09/03/2022

Deceiving Audio Design in Augmented Environments : A Systematic Review of Audio Effects in Augmented Reality

Recently, a lot of works show promising directions for audio design in a...
research
01/16/2023

OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

Inspired by humans comprehending speech in a multi-modal manner, various...
research
07/15/2022

Direction-Aware Adaptive Online Neural Speech Enhancement with an Augmented Reality Headset in Real Noisy Conversational Environments

This paper describes the practical response- and performance-aware devel...
research
09/29/2021

Here To Stay: Measuring Hologram Stability in Markerless Smartphone Augmented Reality

Markerless augmented reality (AR) has the potential to provide engaging ...
research
04/13/2023

The future of hearing aid technology

Background. Hearing aid technology has proven successful in the rehabili...
research
08/24/2023

Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Egocentric, multi-modal data as available on future augmented reality (A...
research
04/24/2020

Binaural Audio Source Remixing with Microphone Array Listening Devices

Augmented listening devices, such as hearing aids and augmented reality ...

Please sign up or login with your details

Forgot password? Click here to reset