Finding Fallen Objects Via Asynchronous Audio-Visual Integration

Chuang Gan; Yi Gu; Siyuan Zhou; Jeremy Schwartz; Seth Alter; James Traer; Dan Gutfreund; Joshua B Tenenbaum; Josh McDermott; Antonio Torralba

doi:10.48550/arXiv.2207.03483

Back

Finding Fallen Objects Via Asynchronous Audio-Visual Integration

Preprint

Open access

Finding Fallen Objects Via Asynchronous Audio-Visual Integration

Chuang Gan, Yi Gu, Siyuan Zhou, Jeremy Schwartz, Seth Alter, James Traer, Dan Gutfreund, Joshua B Tenenbaum, Josh McDermott and Antonio Torralba

ArXiv.org

07/07/2022

DOI: 10.48550/arXiv.2207.03483

Files and links (1)

url

https://doi.org/10.48550/arXiv.2207.03483View

Preprint (Author's original)This preprint has not been evaluated by subject experts through peer review. Preprints may undergo extensive changes and/or become peer-reviewed journal articles. Open Access

Abstract

The way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and then must find it. In this paper, we introduce a setting in which to study multi-modal object localization in 3D virtual environments. An object is dropped somewhere in a room. An embodied robot agent, equipped with a camera and microphone, must determine what object has been dropped -- and where -- by combining audio and visual signals with knowledge of the underlying physics. To study this problem, we have generated a large-scale dataset -- the Fallen Objects dataset -- that includes 8000 instances of 30 physical object categories in 64 rooms. The dataset uses the ThreeDWorld platform which can simulate physics-based impact sounds and complex physical interactions between objects in a photorealistic setting. As a first step toward addressing this challenge, we develop a set of embodied agent baselines, based on imitation learning, reinforcement learning, and modular planning, and perform an in-depth analysis of the challenge of this new task.

Details

Title: Subtitle: Finding Fallen Objects Via Asynchronous Audio-Visual Integration
Creators: Chuang Gan
Yi Gu
Siyuan Zhou
Jeremy Schwartz
Seth Alter
James Traer
Dan Gutfreund
Joshua B Tenenbaum
Josh McDermott
Antonio Torralba
Resource Type: Preprint
Publication Details: ArXiv.org
DOI: 10.48550/arXiv.2207.03483
ISSN: 2331-8422
Language: English
Date posted: 07/07/2022
Academic Unit: Psychological and Brain Sciences; Iowa Neuroscience Institute
Record Identifier: 9984272136002771

Metrics

13 Record Views