IP Library Granted Patent US 12,445,799
Granted Patent B2
US 12,445,799 · App. 18/476,172 · Granted Oct 14, 2025

Surround sound to immersive audio upmixing based on video scene analysis

Inventors: Allan Devantier (Newhall, CA); Sunil Bharitkar (Stevenson Ranch, CA); Seongnam Oh (Irvine, CA); Carlos Tejeda Ocampo (Tuxtla Gutiérrez, MX)
Assignee: Samsung Electronics Co., Ltd.
H04S7/305G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,799
App. No.
18/476,172
Granted
Oct 14, 2025
Kind
B2
Abstract

One embodiment provides a method of audio upmixing comprising performing video scene analysis by segmenting visual objects from video frames of a video, and performing audio analysis by extracting audio signals from an audio corresponding to the video. The method further comprises determining whether any of the audio signals correspond to any of the visual objects, and estimating a video-based trajectory of a visual object if the visual object is in motion and transitions from on-screen to off-screen, or vice versa, during the video. The method further comprises positioning an audio trajectory of an audio signal from at least one speaker associated with the display to at least one other speaker associated with providing surround sound. The audio trajectory is automatically matched with the video. The audio signal is delivered to the at least one speaker and the at least one other speaker for audio reproduction during the presentation.

Claims (46)

1. A method of audio upmixing, comprising:

performing video scene analysis by segmenting one or more visual objects from one or more video frames of a video;

performing audio analysis by extracting one or more audio signals from an audio corresponding to the video;

determining whether any of the audio signals correspond to any of the visual objects;

estimating a video-based trajectory of a visual object of the visual objects if the visual object is in motion and transitions from on-screen to off-screen, or vice versa, during the video; and

positioning an audio trajectory of an audio signal of the audio signals from at least one speaker associated with the display to at least one other speaker associated with providing surround sound, wherein the audio trajectory is automatically matched with the video, and the audio signal is delivered to the at least one speaker and the at least one other speaker for audio reproduction during presentation of the video.

2. The method of claim 1 , wherein each of the audio signals corresponds to either one of the visual objects or a non-visual object that is not visually present in the one or more video frames.

3. The method of claim 1 , wherein the positioning includes panning the audio trajectory of the audio signal between the at least one speaker and the at least one other speaker.

4. The method of claim 1 , wherein the visual trajectory correlates with the panning during the transitions if the audio signal corresponds to the visual object.

5. The method of claim 1 , wherein the extracting comprises:

for each of the audio signals:

classifying the audio signal as directional or diffuse; and

estimating a likelihood that the audio signal is assigned to a horizontal speaker channel or a height speaker channel based on the classifying.

6. The method of claim 1 , wherein the audio signals are extracted from the audio using one or more audio separation techniques.

7. The method of claim 1 , wherein the at least one other speaker comprises at least one of a surround sound speaker or a height speaker.

8. A system of audio upmixing, comprising:

at least one processor; and

a non-transitory processor-readable memory device storing instructions that when executed by the at least one processor causes the at least one processor to perform operations including:

performing video scene analysis by segmenting one or more visual objects from one or more video frames of a video;

performing audio analysis by extracting one or more audio signals from an audio corresponding to the video;

determining whether any of the audio signals correspond to any of the visual objects;

estimating a video-based trajectory of a visual object of the visual objects if the visual object is in motion and transitions from on-screen to off-screen, or vice versa, during the video; and

positioning an audio trajectory of an audio signal of the audio signals from at least one speaker associated with the display to at least one other speaker associated with providing surround sound, wherein the audio trajectory is automatically matched with the video, and the audio signal is delivered to the at least one speaker and the at least one other speaker for audio reproduction during presentation of the video.

9. The system of claim 8 , wherein each of the audio signals corresponds to either one of the visual objects or a non-visual object that is not visually present in the one or more video frames.

10. The system of claim 8 , wherein the positioning includes panning the audio trajectory of the audio signal between the at least one speaker and the at least one other speaker.

11. The system of claim 8 , wherein the visual trajectory correlates with the panning during the transitions if the audio signal corresponds to the visual object.

12. The system of claim 8 , wherein the extracting comprises:

for each of the audio signals:

classifying the audio signal as directional or diffuse; and

estimating a likelihood that the audio signal is assigned to a horizontal speaker channel or a height speaker channel based on the classifying.

13. The system of claim 8 , wherein the audio signals are extracted from the audio using one or more audio separation techniques.

14. The system of claim 8 , wherein the at least one other speaker comprises at least one of a surround sound speaker or a height speaker.

15. A non-transitory processor-readable medium that includes a program that when executed by a processor performs a method of audio upmixing, the method comprising:

performing video scene analysis by segmenting one or more visual objects from one or more video frames of a video;

performing audio analysis by extracting one or more audio signals from an audio corresponding to the video;

determining whether any of the audio signals correspond to any of the visual objects;

estimating a video-based trajectory of a visual object of the visual objects if the visual object is in motion and transitions from on-screen to off-screen, or vice versa, during the video; and

positioning an audio trajectory of an audio signal of the audio signals from at least one speaker associated with the display to at least one other speaker associated with providing surround sound, wherein the audio trajectory is automatically matched with the video, and the audio signal is delivered to the at least one speaker and the at least one other speaker for audio reproduction during presentation of the video.

16. The non-transitory processor-readable medium of claim 15 , wherein each of the audio signals corresponds to either one of the visual objects or a non-visual object that is not visually present in the one or more video frames.

17. The non-transitory processor-readable medium of claim 15 , wherein the positioning includes panning the audio trajectory of the audio signal between the at least one speaker and the at least one other speaker.

18. The non-transitory processor-readable medium of claim 15 , wherein the visual trajectory correlates with the panning during the transitions if the audio signal corresponds to the visual object.

19. The non-transitory processor-readable medium of claim 15 , wherein the extracting comprises:

for each of the audio signals:

classifying the audio signal as directional or diffuse; and

estimating a likelihood that the audio signal is assigned to a horizontal speaker channel or a height speaker channel based on the classifying.

20. The non-transitory processor-readable medium of claim 15 , wherein the at least one other speaker comprises at least one of a surround sound speaker or a height speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2023
From: DEVANTIER, ALLAN; BHARITKAR, SUNIL; OH, SEONGNAM; TEJEDA OCAMPO, CARLOS
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 065055/0122 →
Continuity (2)
Provisional Application 63431263 · Dec 8, 2022
Related Publication 20240196158A1 · Jun 13, 2024
References Cited (15)
US 10073607B2 · Kim et al. · 2018 [cited by applicant]
US 10200804B2 · Chen et al. · 2019 [cited by applicant]
US 10477311B2 · Vilkamo · 2019 [cited by applicant]
US 11026037B2 · Sridharan et al. · 2021 [cited by applicant]
US 11089423B2 · Kim et al. · 2021 [cited by applicant]
US 20140119581A1 · Tsingos et al. · 2014 [cited by applicant]
US 20160225377A1 · Miyasaka et al. · 2016 [cited by applicant]
US 20170013202A1 · Xu et al. · 2017 [cited by applicant]
US 20170013387A1 · Fersch et al. · 2017 [cited by applicant]
US 20200368616A1 · Delamont · 2020 [cited by applicant]
KR 102057393B1 · 2019 [cited by applicant]
WO 2018162803A1 · 2018 [cited by applicant]
Avendano, C. et al., “A Frequency-Domain Approach to Multichannel Upmix,” J. Audio Eng. Society, Jul. 2004, pp. 740-749, v. 52, No. 7/8, United States. [cited by applicant]
Nikunen, J., et al., “Multichannel Audio Upmixing by time-frequency filtering using non-negative tensor factorization,” J. Audio Eng. Society, Oct. 2012, pp. 794-806, v. 60, n. 10, United States. [cited by applicant]
International Search Report and Written Opinion dated Feb. 2, 2024 for International Application PCT/KR2023/015705, from Korean Intellectual Property Office, pp. 1-11, Republic of Korea. [cited by applicant]