IP Library Granted Patent US 12,417,070
Granted Patent B2
US 12,417,070 · App. 18/327,134 · Granted Sep 16, 2025

Video-informed spatial audio expansion

Inventors: Marcin Gorzel (Dublin, IE); Balineedu Adsumilli (Sunnyvale, CA)
Assignee: GOOGLE LLC
G06F3/165G06V20/41G10L25/51H04S5/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,070
App. No.
18/327,134
Granted
Sep 16, 2025
Kind
B2
Abstract

First video frames that include a visual object and a non-spatialized first audio segment that includes an auditory event are received. If that second video frames do not include the visual object and a first time difference between the first video frames and the second video frames does not exceed a certain time, a motion vector of the visual object is used to assign a spatial location to the auditory event in at least one of the second video frames. A second audio segment that includes the auditory event and third video frames are received. If the third video frames do not include the visual object and a second time difference between the first video frames and the third video frames exceeds the certain time, the auditory event is assigned to a diffuse sound field. An audio output that conveys spatial locations of the visual object is output.

Claims (76)

1. A method of assigning spatial information to audio segments, comprising:

receiving first video frames that include a visual object;

receiving a first audio segment, wherein the first audio segment includes an auditory event associated with the visual object, and wherein the first audio segment is non-spatialized;

upon determining that second video frames do not include the visual object and that a first time difference between the first video frames and the second video frames are not longer than a certain time, using a motion vector of the visual object to assign a spatial location to the auditory event in at least one of the second video frames;

receiving a second audio segment, wherein the second audio segment includes the auditory event;

receiving third video frames, wherein the third video frames do not include the visual object;

upon determining that the third video frames do not include the visual object and that a second time difference between the first video frames and the third video frames are longer than the certain time, assigning the auditory event to a diffuse sound field; and

generating an audio output that includes spatial locations of the visual object to a listener.

2. The method of claim 1 , further comprising:

identifying visual objects including the visual object in the first video frames by steps comprising:

assigning respective visual labels to the visual objects in the first video frames;

identifying auditory events including the auditory event in the first audio segment by steps comprising:

separating the first audio segment into multiple tracks; and

assigning respective audio labels to the multiple tracks; and

identifying matches between some of the auditory events and some of the visual objects by automatically matching some of the respective audio labels to some of the respective visual labels.

3. The method of claim 1 , further comprising:

identifying an unmatched auditory event, wherein the unmatched auditory event is not matched to an identified visual object in the first video frames; and

presenting the unmatched auditory event in a user interface.

4. The method of claim 3 , further comprising:

receiving, from a user, an assignment of the unmatched auditory event to another visual object identified in the first video frames.

5. The method of claim 3 , further comprising:

receiving, from a user, an indication to assign the unmatched auditory event as diffuse sound.

6. The method of claim 3 , further comprising:

receiving, from a user, an indication to assign the unmatched auditory event as directional sound and a spatial direction for the unmatched auditory event.

7. The method of claim 1 , wherein the first audio segment is monophonic.

8. The method of claim 1 , further comprising:

using blind source separation to identify auditory events that include the auditory event in the first audio segment by decomposing the first audio segment into multiple tracks, each corresponding to a respective auditory event.

9. The method of claim 1 , further comprising:

using image recognition to identify visual objects including the visual object in the first video frames.

10. A system, comprising:

a memory; and

a processor, the processor configured to execute instructions stored in the memory to:

receive first video frames that include a visual object;

receive an audio segment, wherein the audio segment includes an auditory event associated with the visual object;

receive second video frames, wherein the second video frames do not include the visual object;

upon to determining that the second video frames do not include the visual object and that a first time difference between the first video frames and the second video frames are not longer than a certain time, use a motion vector of the visual object to assign a second spatial location to the auditory event in at least one of the second video frames;

receive a third audio segment, wherein the third audio segment includes the auditory event;

receive third video frames, wherein the third video frames do not include the visual object;

upon determining that the third video frames do not include the visual object and that a second time difference between the first video frames and the third video frames are longer than the certain time, assign the auditory event to a diffuse sound field; and

generate an audio output that includes spatial locations of the visual object to a listener.

11. The system of claim 10 , wherein the processor is further configured to execute instructions stored in the memory to:

identify visual objects including the visual object in the first video frames, wherein to identify the visual objects comprises to:

assign respective visual labels to the visual objects in the first video frames;

identify auditory events including the auditory event in the audio segment, wherein to identify the auditory events comprises to:

separate the audio segment into multiple tracks; and

assign respective audio labels to the multiple tracks; and

identify matches between some of the auditory events and some of the visual objects by automatically matching some of the respective audio labels to some of the respective visual labels.

12. The system of claim 10 , wherein the processor is further configured to execute instructions stored in the memory to:

identify an unmatched auditory event, wherein the unmatched auditory event is not matched to an identified visual object in the first video frames; and

present the unmatched auditory event in a user interface.

13. The system of claim 12 , wherein the processor is further configured to execute instructions stored in the memory to:

receive, from a user, an assignment of the unmatched auditory event to another visual object identified in the first video frames.

14. The system of claim 12 , wherein the processor is further configured to execute instructions stored in the memory to:

receive, from a user, an indication to assign the unmatched auditory event as diffuse sound.

15. The system of claim 12 , wherein the processor is further configured to execute instructions stored in the memory to:

receive, from a user, an indication to assign the unmatched auditory event as directional sound and a spatial direction for the unmatched auditory event.

16. The system of claim 10 , wherein the audio segment is monophonic.

17. The system of claim 10 , wherein the processor is further configured to execute instructions stored in the memory to:

use blind source separation to identify auditory events that include the auditory event in the audio segment by decomposing the audio segment into multiple tracks, each corresponding to a respective auditory event.

18. The system of claim 10 , wherein the processor is further configured to execute instructions stored in the memory to:

use image recognition to identify visual objects including the visual object in the first video frames.

19. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, perform operations for assigning spatial information to audio segments, the operations comprising:

receiving first video frames that include a visual object;

receiving a first audio segment, wherein the first audio segment includes an auditory event associated with the visual object, and wherein the first audio segment is non-spatialized;

upon determining that second video frames do not include the visual object and that a first time difference between the first video frames and the second video frames are not longer than a certain time, using a motion vector of the visual object to assign a spatial location to the auditory event in at least one of the second video frames;

receiving a second audio segment, wherein the second audio segment includes the auditory event;

receiving third video frames, wherein the third video frames do not include the visual object;

upon determining that the third video frames do not include the visual object and that a second time difference between the first video frames and the third video frames are longer than the certain time, assigning the auditory event to a diffuse sound field; and

generating an audio output that includes spatial locations of the visual object to a listener.

20. The non-transitory computer-readable storage medium of claim 19 , the operations further comprising:

identifying visual objects including the visual object in the first video frames by steps comprising:

assigning respective visual labels to the visual objects in the first video frames;

identifying auditory events including the auditory event in the first audio segment by steps comprising:

separating the first audio segment into multiple tracks; and

assigning respective audio labels to the multiple tracks; and

identifying matches between some of the auditory events and some of the visual objects by automatically matching some of the respective audio labels to some of the respective visual labels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2023
From: GORZEL, MARCIN; ADSUMILLI, BALINEEDU
To: GOOGLE LLC
Reel/Frame 063819/0698 →
Continuity (2)
Continuation 16779921 · Feb 3, 2020
Related Publication 20230305800A1 · Sep 28, 2023
References Cited (36)
US 9888333B2 · Zurek et al. · 2018 [cited by applicant]
US 20030053680A1 · Lin et al. · 2003 [cited by applicant]
US 20050281410A1 · Grosvenor · 2005 [cited by examiner]
US 20060044419A1 · Ozawa · 2006 [cited by examiner]
US 20120307048A1 · Abrahamsson · 2012 [cited by examiner]
US 20140233917A1 · Xiang · 2014 [cited by applicant]
US 20140314391A1 · Kim · 2014 [cited by examiner]
US 20140316543A1 · Sharma · 2014 [cited by examiner]
US 20150245133A1 · Kim et al. · 2015 [cited by applicant]
US 20160005435A1 · Campbell · 2016 [cited by examiner]
US 20160337776A1 · Breebaart · 2016 [cited by examiner]
US 20170215005A1 · Hsu et al. · 2017 [cited by applicant]
US 20170249973A1 · Antonellis et al. · 2017 [cited by applicant]
US 20170293461A1 · McCauley · 2017 [cited by examiner]
US 20190243530A1 · De Ridder et al. · 2019 [cited by applicant]
US 20200312347A1 · Mate · 2020 [cited by examiner]
EP 1643769A1 · 2006 [cited by applicant]
JP 2006123161A · 2006 [cited by applicant]
JP 2007272733A · 2007 [cited by applicant]
JP 2010117946A · 2010 [cited by applicant]
JP 2011071683A · 2011 [cited by applicant]
JP 2015032001A · 2015 [cited by applicant]
JP 2016062071A · 2016 [cited by applicant]
JP 2016513410A · 2016 [cited by applicant]
JP 2019050482A · 2019 [cited by applicant]
JP 2019078864A · 2019 [cited by applicant]
JP 2019523902A · 2019 [cited by applicant]
Weitnaur et al; “Object-based Audio, the future of audio production, delivery and consumption”; https://lab.irt.de/demos/object-based-audio/; 2018; pp. 1-5. [cited by applicant]
https://en.wikipedia.org/wiki/Surround_sound; Jan. 2020; pp. 1-12. [cited by applicant]
https://en.wikipedia.org/wiki/Ambisonics; Sep. 2019; pp. 1-12. [cited by applicant]
“Immersive Sound and Object-Based Audio and Microphones”; Nov. 2019; 18 Pages. [cited by applicant]
“Multichannel sound technology in home and broadcasting applications”; Report ITU-R BS.2159-4; May 2012; pp. 1-54. [cited by applicant]
Makino et al: “Blind Audio Source Separation based on Independent Component Analysis”; NTT Communication Science Laboratories, Kyoto, Japan; Apr. 2003; pp. 1-35. [cited by applicant]
Hawksford, Malcolm J.; “Diffuse Signal Processing and Acoustic Source Characterization for Applications in Synthetic Loudspeaker Arrays”; http://www.aes.org/e-lib/browse.cfm?elib=11405; Apr. 2002; 4 Pages. [cited by applicant]
Vision AI; https://cloud.google.com/vision; Feb. 2020; pp. 1-15. [cited by applicant]
International Search Report and Written Opinion of International Application No. PCT/US2020/055964 dated Jan. 29, 2021. [cited by applicant]