IP Library Granted Patent US 12,273,697
Granted Patent B2
US 12,273,697 · App. 18/042,258 · Granted Apr 8, 2025

Systems and methods for upmixing audiovisual data

Inventors: Aren Jansen (Mountain View, CA); Manoj Plakal (New York, NY); Dan Ellis (New York, NY); Shawn Hershey (Kirkland, WA); Richard Channing Moore, III (Brooklyn, NY)
Assignee: GOOGLE LLC
H04S7/301H04S2400/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,273,697
App. No.
18/042,258
Granted
Apr 8, 2025
Kind
B2
Abstract

A computer-implemented method for upmixing audiovisual data can include obtaining audiovisual data including input audio data and video data accompanying the input audio data. Each frame of the video data can depict only a portion of a larger scene. The input audio data can have a first number of audio channels. The computer-implemented method can include providing the audiovisual data as input to a machine-learned audiovisual upmixing model. The audiovisual upmixing model can include a sequence-to-sequence model configured to model a respective location of one or more audio sources within the larger scene over multiple frames of the video data. The computer-implemented method can include receiving upmixed audio data from the audiovisual upmixing model. The upmixed audio data can have a second number of audio channels. The second number of audio channels can be greater than the first number of audio channels.

Claims (34)

1. A computer-implemented method for upmixing audiovisual data, the computer-implemented method comprising:

obtaining, by a computing system comprising one or more computing devices, audiovisual data comprising input audio data and video data accompanying the input audio data, wherein each frame of the video data depicts only a portion of a larger scene, and wherein the input audio data has a first number of audio channels;

providing, by the computing system, the audiovisual data as input to a machine-learned audiovisual upmixing model, the audiovisual upmixing model comprising a sequence-to-sequence model configured to model a respective location of one or more audio sources within the larger scene over multiple frames of the video data; and

receiving, by the computing system, upmixed audio data from the audiovisual upmixing model, the upmixed audio data having a second number of audio channels, the second number of audio channels greater than the first number of audio channels.

2. The computer-implemented method of claim 1 , wherein the audiovisual upmixing model comprises an encoder-decoder model.

3. The computer-implemented method of claim 1 , wherein the audiovisual upmixing model comprises a transformer model.

4. The computer-implemented method of claim 1 , wherein the audiovisual upmixing model comprises an attention mechanism.

5. The computer-implemented method of claim 4 , wherein the attention mechanism comprises a plurality of context vectors and an alignment model.

6. The computer-implemented method of claim 1 , wherein the audiovisual upmixing model comprises a plurality of input streams, each of the plurality of input streams corresponding to a respective audio channel of the input audio data, and a plurality of output streams, each of the plurality of output streams corresponding to a respective audio channel of the upmixed audio data.

7. The computer-implemented method of claim 1 , wherein the video data comprises two-dimensional video data.

8. The computer-implemented method of claim 1 , wherein the input audio data comprises mono audio data, the mono audio data having a single audio channel.

9. The computer-implemented method of claim 1 , wherein the upmixed audio data comprises stereo audio data, the stereo audio data having a left audio channel and a right audio channel.

10. The computer-implemented method of claim 1 , wherein the input audio data comprises stereo audio data, the stereo audio data having a left audio channel and a right audio channel.

11. The computer-implemented method of claim 1 , wherein the upmixed audio data comprises surround sound audio data, the surround sound audio data having three or more audio channels.

12. The computer-implemented method of claim 1 , wherein training the machine-learned audiovisual upmixing model comprises:

obtaining, by the computing system, audiovisual training data comprising video training data and audio training data having the second number of audio channels;

downmixing, by the computing system, the audio training data to produce downmixed audio training data comprising the first number of audio channels;

providing, by the computing system, the video training data and corresponding downmixed audio training data to the audiovisual upmixing model;

obtaining, by the computing system, a predicted upmixed audio data output comprising the second number of audio channels from the audiovisual upmixing model;

determining, by the computing system, a difference between the predicted upmixed audio data and the audio training data; and

updating one or more parameters of the model based the difference.

13. A computing system configured for upmixing audiovisual data, the computing system comprising:

one or more processors; and

one or more memory devices storing computer-readable data comprising instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:

obtaining audiovisual data comprising input audio data and video data accompanying the input audio data, the input audio data having a first number of audio channels;

providing the audiovisual data as input to a machine-learned audiovisual upmixing model, the audiovisual upmixing model comprising a sequence-to-sequence model; and

receiving upmixed audio data from the audiovisual upmixing model, the upmixed audio data having a second number of audio channels, the second number of audio channels greater than the first number of audio channels.

14. The computing system of claim 13 , wherein the audiovisual upmixing model comprises an encoder-decoder model.

15. The computing system of claim 13 , wherein the audiovisual upmixing model comprises a transformer model.

16. The computing system of claim 13 , wherein the audiovisual upmixing model comprises an attention mechanism.

17. The computing system of claim 16 , wherein the attention mechanism comprises a plurality of context vectors and an alignment model.

18. The computing system of claim 13 , wherein the audiovisual upmixing model comprises a plurality of internal state vectors.

19. The computing system of claim 13 , wherein the audiovisual upmixing model comprises a plurality of input streams, each of the plurality of input streams corresponding to a respective audio channel of the input audio data, and a plurality of output streams, each of the plurality of output streams corresponding to a respective audio channel of the upmixed audio data.

20. The computing system of claim 13 , wherein the video data comprises two-dimensional video data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2023
From: PLAKAL, MANOJ; ELLIS, DAN; HERSHEY, SHAWN; MOORE, RICHARD CHANNING, III; JANSEN, AREN
To: GOOGLE LLC
Reel/Frame 062744/0287 →
Continuity (1)
Related Publication 20230308823A1 · Sep 28, 2023
References Cited (17)
US 11321542B2 · Kalchbrenner et al. · 2022 [cited by applicant]
US 20190306451A1 · Wang et al. · 2019 [cited by applicant]
US 20200258496A1 · Yang et al. · 2020 [cited by applicant]
US 20200372359A1 · Shaked et al. · 2020 [cited by applicant]
US 20210021949A1 · Sridharan · 2021 [cited by examiner]
US 20210232705A1 · Chandelier et al. · 2021 [cited by applicant]
CN 111428015 · 2020 [cited by applicant]
WO WO2018100244 · 2018 [cited by applicant]
WO WO2019229199 · 2019 [cited by examiner]
International Preliminary Report on Patentability for Application No. PCT/US2020/047930, mailed Mar. 9, 2023, 9 pages. [cited by applicant]
Anonymous, “Seq2seq—Wikipedia”, Aug. 1, 2020, https://em/wikipedia.org/w/index.php?title=Seq2seq&oldid=970695786, retrieved Apr. 4, 2021, XP055797487. [cited by applicant]
International Search Report for Application No. PCT/US2020/047930, mailed on May 3, 2021, 3 pages. [cited by applicant]
Machine Translated Chinese Search Report Corresponding to Application No. 2020801024806 on Jan. 23, 2025. [cited by applicant]
Gao et al., “2.5D Visual Sound”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 10 pages. [cited by applicant]
Jianming Wu, “2.5D Visual Sound”, 2019, https://www.cnblogs.com/wujianming-110117/p/12677563.html, retrieved on Mar. 5, 2025, 18 pages. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, arXiv:1706.03762v7, 15 pages. [cited by applicant]
Zhang et al., “Comprehensive 3D game design, game engine and game development case analysis”, China Railway Publishing House, 5 pages. [cited by applicant]