IP Library Granted Patent US 12,288,567
Granted Patent B2
US 12,288,567 · App. 17/792,073 · Granted Apr 29, 2025

Method for training a neural network to describe an environment on the basis of an audio signal, and the corresponding neural network

Inventors: Wim Abbeloos (Brussels, BE); Arun Balajee Vasudevan (Zurich, CH); Dengxin Dai (Zurich, CH); Luc Van Gool (Zurich, CH)
Assignees: TOYOTA JIDOSHA KABUSHIKI KAISHA; ETH ZÜRICH
G10L25/30G06V10/26G06V10/774G06V10/82G06V10/95G06V20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,288,567
App. No.
17/792,073
Granted
Apr 29, 2025
Kind
B2
Abstract

A neural network, a system using this neural network and a method for training a neural network to output a description of the environment in the vicinity of at least one sound acquisition device on the basis of an audio signal acquired by the sound acquisition device, the method including: obtaining audio and image training signals of a scene showing an environment with objects generating sounds, obtaining a target description of the environment seen on the image training signal, inputting the audio training signal to the neural network so that the neural network outputs a training description of the environment, and comparing the target description of the environment with the training description of the environment.

Claims (32)

1. A method for training a neural network to output a description of the environment in the vicinity of at least one sound acquisition device on the basis of an audio signal acquired by the sound acquisition device, the method comprising:

obtaining audio and image training signals of a scene showing an environment with objects generating sounds,

obtaining a target description of the environment seen on the image training signal,

inputting the audio training signal to the neural network so that the neural network outputs a training description of the environment, and

comparing the target description of the environment with the training description of the environment, wherein

the description of the environment, the target description of the environment, and the training description of the environment include at least one of a semantic segmentation of a frame of the image training signal or a depth map of a frame of the image training signal.

2. The method of claim 1 , wherein the audio training signal is acquired with a plurality of sound acquisition devices.

3. The method of claim 2 , wherein the sound acquisition devices of the plurality of sound acquisition devices are all spaced apart from each other.

4. The method of claim 2 , wherein

at least one additional sound acquisition device is used to acquire an audio signal at a location which differs from the location of any one of the sound acquisition devices of the plurality of sound acquisition devices,

the neural network being further configured to determine at least one predicted audio signal representative of the audio signal that is acquired by the at least one additional sound acquisition device, and

the method further comprising comparing the predicted audio signal with an audio signal acquired by the at least one additional sound acquisition device.

5. The method of claim 1 , wherein the audio training signal is acquired using at least one binaural sound acquisition device.

6. The method of claim 1 , wherein the image training signal is acquired using a 360 degrees camera.

7. The method of claim 1 , wherein the target description is obtained using at least one pre-trained neural network configured to receive an image signal as input and to output the target description.

8. A neural network trained using the method of claim 1 .

9. The neural network of claim 8 , comprising, for each possible audio signal to be used as input, four convolutional layers, a concatenation module for concatenating the outputs of every four convolutional layers, and an ASPP module.

10. A system comprising at least one sound acquisition device and

a neural network in accordance with claim 8 .

11. A vehicle comprising a system according to claim 10 .

12. A system for training a neural network to output a description of the environment in the vicinity of at least one sound acquisition device on the basis of an audio signal acquired by the sound acquisition device, the system comprising:

a module for obtaining audio and image training signals of a scene showing an environment with objects generating sounds,

a module for obtaining a target description of the environment seen on the image training signal,

a module for inputting the audio training signal to the neural network so that the neural network outputs a training description of the environment, and

a module for comparing the target description of the environment with the training description of the environment, wherein

the description of the environment, the target description of the environment, and the training description of the environment include at least one of a semantic segmentation of a frame of the image training signal or a depth map of a frame of the image training signal.

13. A non-transitory recording medium readable by a computer and having recorded thereon a computer program including instructions for executing a method for training a neural network to output a description of the environment in the vicinity of at least one sound acquisition device on the basis of an audio signal acquired by the sound acquisition device, the method comprising:

obtaining audio and image training signals of a scene showing an environment with objects generating sounds,

obtaining a target description of the environment seen on the image training signal,

inputting the audio training signal to the neural network so that the neural network outputs a training description of the environment, and

comparing the target description of the environment with the training description of the environment, wherein

the description of the environment, the target description of the environment, and the training description of the environment include at least one of a semantic segmentation of a frame of the image training signal or a depth map of a frame of the image training signal.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2024
From: TOYOTA MOTOR EUROPE
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 068226/0776 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2023
From: ABBELOOS, WIM; BALAJEE VASUDEVAN, ARUN; DAI, DENGXIN; VAN GOOL, LUC
To: TOYOTA MOTOR EUROPE; ETH ZURICH
Reel/Frame 063192/0972 →
Continuity (1)
Related Publication 20230047017A1 · Feb 16, 2023
References Cited (25)
US 10372991B1 · Niemasik · 2019 [cited by examiner]
US 10977872B2 · Krishnamurthy · 2021 [cited by examiner]
US 11210523B2 · Geng · 2021 [cited by examiner]
US 11375293B2 · Kumar · 2022 [cited by examiner]
US 11381888B2 · Krishnamurthy · 2022 [cited by examiner]
US 11727913B2 · Verma · 2023 [cited by examiner]
US 11756260B1 · Ponce · 2023 [cited by examiner]
US 11756551B2 · Moritz · 2023 [cited by examiner]
US 12075187B2 · Wang · 2024 [cited by examiner]
US 20180096632A1 · Florez · 2018 [cited by examiner]
US 20190174237A1 · Lunner · 2019 [cited by examiner]
US 20230047017A1 · Abbeloos · 2023 [cited by examiner]
Gan et al.; “Self-supervised Moving Vehicle Tracking with Stereo Sound;” 2019 IEEE/CVF International Conference on Computer Vision (ICCV); 2019; pp. 7052-7061. [cited by applicant]
Aytar et al.; “SoundNet: Learning Sound Representations from Unlabeled Video;” 30th Conference on Neural Information Processing Systems (NIPS); 2016; pp. 1-9. [cited by applicant]
Gao et al.; “2.5D Visual Sound;” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2019; pp. 324-333. [cited by applicant]
Chen et al.; “Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation;” Annual International Conference on the Theory and Applications of Cryptographic Techniques, Eurocrypt 2018; 2018; pp. 833… [cited by applicant]
Marchegiani et al.; “Leveraging the Urban Soundscape: Auditory Perception for Smart Vehicles;” 2017 IEEE International Conference on Robotics and Automation (ICRA); 2017; pp. 6547-6554. [cited by applicant]
Zhao et al.; “The Sound of Pixels;” Arvix:1804.03160; 2018; pp. 1-17. [cited by applicant]
Cordts et al.; “The Cityscapes Dataset for Semantic Urban Scene Understanding;” In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016; pp. 1-29. [cited by applicant]
Godard et al.; “Digging Into Self-Supervised Monocular Depth Estimation;” In Proceedings of the IEEE International Conference on Computer Vision; 2019; pp. 3828-3838. [cited by applicant]
Geiger et al.; “Vision meets Robotics: The KITTI Dataset;” International Journal of Robotics Research (IJRR); 2013; pp. 1-6. [cited by applicant]
Chen et al.; “Semantic Image Segmentation With Deep Convolutional Nets and Fully Connected CRFS;” IEEE transactions on pattern analysis and machine intelligence; 2017; pp. 834-848; vol. 40, No. 4. [cited by applicant]
Anonymous; “Auditory Semantic and Depth Perception with Binaural Sounds;” CVPR 2020 Submission #1702; 2020;. pp. 1-10. [cited by applicant]
Sep. 21, 2020 Search Report issued in International Patent Application No. PCT/EP2020/050605. [cited by applicant]
Sep. 21, 2020 Written Opinion of the International Searching Authority issued in International Patent Application No. PCT/EP2020/050605. [cited by applicant]
Cited By (2)
US 12,452,463 US 12,694,650