IP Library Granted Patent US 12,425,797
Granted Patent B2
US 12,425,797 · App. 18/335,730 · Granted Sep 23, 2025

Three-dimensional (3D) sound rendering with multi-channel audio based on mono audio input

Inventors: Vijendra Raj Apsingekar (San Jose, CA); Akash Sahoo (San Jose, CA); Anil S. Yadav (San Jose, CA); Sivakumar Balasubramanian (Sunnyvale, CA)
Assignee: Samsung Electronics Co., Ltd.
H04S7/304G06F3/165G10L19/008H04S3/008H04S2400/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,425,797
App. No.
18/335,730
Granted
Sep 23, 2025
Kind
B2
Abstract

A method includes obtaining video content and associated substantially mono audio content. The method also includes determining at least one of a position or a motion trajectory of each of one or more objects detected in the video content and classifying each of the one or more objects into one of multiple object classes. The method further includes separating audio streams within the audio content based on the video content. Each of the audio streams is associated with one of multiple audio sources. The method also includes classifying each of the audio sources into one of the object classes. In addition, the method includes, for each audio source classified into the same object class as one of the one or more objects, distributing the audio stream associated with that audio source into multiple audio channels based on at least one of the position or the motion trajectory of that object.

Claims (51)

1. A method comprising:

obtaining video content and associated substantially mono audio content;

determining at least one of a position or a motion trajectory of each of one or more objects detected in the video content;

classifying each of the one or more objects into one of multiple object classes;

separating audio streams within the audio content based on the video content, each of the audio streams associated with one of multiple audio sources;

classifying each of the audio sources into one of the object classes; and

for each of the audio sources classified into the same object class as one of the one or more objects, distributing the audio stream associated with that audio source into multiple audio channels based on at least one of the position or the motion trajectory of that object.

2. The method of claim 1 , wherein, for each of the audio sources classified into the same object class as one of the one or more objects, distributing the audio stream associated with that audio source into the multiple audio channels comprises:

determining an amplitude and a phase of the audio stream associated with that audio source for each of the multiple audio channels, the amplitudes and the phases of the audio stream based on at least one of the position or the motion trajectory of that object.

3. The method of claim 2 , wherein:

each of multiple ones of the audio sources is classified into the same object class as one of the one or more objects; and

the method further comprises combining distributed portions of the audio streams associated with those audio sources within each of the multiple audio channels with one another.

4. The method of claim 3 , wherein combining the distributed portions of the audio streams comprises:

combining the distributed portions of the audio streams within each of the multiple audio channels with one another and with one or more of the audio streams associated with one or more of the audio sources not classified into the same object class as any of the one or more objects.

5. The method of claim 4 , wherein an amplitude and a phase of each of the audio streams not classified into the same object class as any of the one or more objects are unchanged during combination with the distributed portions of the audio streams within each of the multiple audio channels.

6. The method of claim 1 , further comprising:

generating a multi-channel audio output, the multi-channel audio output comprising different audio data in different audio channels.

7. The method of claim 1 , wherein separating the audio streams within the audio content and classifying each of the audio sources are based on a classification of each of the one or more objects into one of the object classes.

8. An electronic device comprising:

at least one processing device configured to:

obtain video content and associated substantially mono audio content;

determine at least one of a position or a motion trajectory of each of one or more objects detected in the video content;

classify each of the one or more objects into one of multiple object classes;

separate audio streams within the audio content based on the video content, each of the audio streams associated with one of multiple audio sources;

classify each of the audio sources into one of the object classes; and

for each of the audio sources classified into the same object class as one of the one or more objects, distribute the audio stream associated with that audio source into multiple audio channels based on at least one of the position or the motion trajectory of that object.

9. The electronic device of claim 8 , wherein, for each of the audio sources classified into the same object class as one of the one or more objects, to distribute the audio stream associated with that audio source into the multiple audio channels, the at least one processing device is configured to:

determine an amplitude and a phase of the audio stream associated with that audio source for each of the multiple audio channels, the amplitudes and the phases of the audio stream based on at least one of the position or the motion trajectory of that object.

10. The electronic device of claim 9 , wherein:

the at least one processing device is configured to classify each of multiple ones of the audio sources into the same object class as one of the one or more objects; and

the at least one processing device is further configured to combine distributed portions of the audio streams associated with those audio sources within each of the multiple audio channels with one another.

11. The electronic device of claim 10 , wherein, to combine the distributed portions of the audio streams, the at least one processing device is configured to combine the distributed portions of the audio streams within each of the multiple audio channels with one another and with one or more of the audio streams associated with one or more of the audio sources not classified into the same object class as any of the one or more objects.

12. The electronic device of claim 11 , wherein the at least one processing device is configured to leave an amplitude and a phase of each of the audio streams not classified into the same object class as any of the one or more objects unchanged during combination with the distributed portions of the audio streams within each of the multiple audio channels.

13. The electronic device of claim 8 , wherein the at least one processing device is further configured to generate a multi-channel audio output, the multi-channel audio output comprising different audio data in different audio channels.

14. The electronic device of claim 8 , wherein the at least one processing device is configured to separate the audio streams within the audio content and classify each of the audio sources based on a classification of each of the one or more objects into one of the object classes.

15. A non-transitory machine readable medium containing instructions that when executed cause at least one processor of an electronic device to:

obtain video content and associated substantially mono audio content;

determine at least one of a position or a motion trajectory of each of one or more objects detected in the video content;

classify each of the one or more objects into one of multiple object classes;

separate audio streams within the audio content based on the video content, each of the audio streams associated with one of multiple audio sources;

classify each of the audio sources into one of the object classes; and

for each of the audio sources classified into the same object class as one of the one or more objects, distribute the audio stream associated with that audio source into multiple audio channels based on at least one of the position or the motion trajectory of that object.

16. The non-transitory machine readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor, for each of the audio sources classified into the same object class as one of the one or more objects, to distribute the audio stream associated with that audio source into the multiple audio channels comprise:

instructions that when executed cause the at least one processor to determine an amplitude and a phase of the audio stream associated with that audio source for each of the multiple audio channels, the amplitudes and the phases of the audio stream based on at least one of the position or the motion trajectory of that object.

17. The non-transitory machine readable medium of claim 16 , wherein:

the instructions when executed cause the at least one processor to classify each of multiple ones of the audio sources into the same object class as one of the one or more objects; and

the non-transitory machine readable medium further contains instructions that when executed cause the at least one processor to combine distributed portions of the audio streams associated with those audio sources within each of the multiple audio channels with one another.

18. The non-transitory machine readable medium of claim 17 , wherein the instructions that when executed cause the at least one processor to combine the distributed portions of the audio streams comprise:

instructions that when executed cause the at least one processor to combine the distributed portions of the audio streams within each of the multiple audio channels with one another and with one or more of the audio streams associated with one or more of the audio sources not classified into the same object class as any of the one or more objects.

19. The non-transitory machine readable medium of claim 18 , wherein the instructions when executed cause the at least one processor to leave an amplitude and a phase of each of the audio streams not classified into the same object class as any of the one or more objects unchanged during combination with the distributed portions of the audio streams within each of the multiple audio channels.

20. The non-transitory machine readable medium of claim 15 , further containing instructions that when executed cause the at least one processor to generate a multi-channel audio output, the multi-channel audio output comprising different audio data in different audio channels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2023
From: APSINGEKAR, VIJENDRA RAJ; SAHOO, AKASH; YADAV, ANIL S.; BALASUBRAMANIAN, SIVAKUMAR
To: SAMSUNG ELECTRONICS CO., LTD
Reel/Frame 063965/0857 →
Continuity (2)
Provisional Application 63396840 · Aug 10, 2022
Related Publication 20240056761A1 · Feb 15, 2024
References Cited (20)
US 6697564B1 · Toklu et al. · 2004 [cited by applicant]
US 6829018B2 · Lin et al. · 2004 [cited by applicant]
US 9654895B2 · Breebaart et al. · 2017 [cited by applicant]
US 10993063B2 · Yan · 2021 [cited by applicant]
US 11902704B2 · Honma · 2024 [cited by examiner]
US 12100416B2 · Charantimath · 2024 [cited by examiner]
US 20150131966A1 · Zurek · 2015 [cited by examiner]
US 20200381003A1 · Gerrard · 2020 [cited by applicant]
US 20220214858A1 · Karri · 2022 [cited by examiner]
US 20230043122A1 · Shin et al. · 2023 [cited by applicant]
US 20230305800A1 · Gorzel · 2023 [cited by examiner]
US 20230360665A1 · Kim · 2023 [cited by examiner]
EP 2727383B1 · 2021 [cited by applicant]
WO WO2014127019A1 · 2014 [cited by examiner]
MIT Technology Review, “Deep learning turns mono recordings into immersive sound,” Emerging Technology from the arXivarchive, Dec. 2018, 6 pages. [cited by applicant]
Wisdom et al., “Unsupervised Sound Separation Using Mixtures of Mixtures,” NeurIPS 2020, Jun. 2020, 14 pages. [cited by applicant]
Tzinis et al., “Into the Wild With Audioscope: Unsupervised Audio-Visual Separation of on-Screen Sounds,” ICLR 2021 Conference Paper, May 2021, 27 pages. [cited by applicant]
Tzinis et al., “Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds,” ICLR 2021 Poster Session 2, May 2021, 18 pages. [cited by applicant]
Tzinis et al., “AudioScopeV2: Audio-Visual Attention Architectures for Calibrated Open-Domain On-Screen Sound Separation,” ECCV 2022, Jul. 2022, 33 pages. [cited by applicant]
Morgado et al., “Self-Supervised Generation of Spatial Audio for 360 Video,” NeurIPS 2018, Sep. 2018, 11 pages. [cited by applicant]