IP Library › Granted Patent US 12,581,266
Granted Patent B2
US 12,581,266 · App. 18/399,148 · Granted Mar 17, 2026

Deep learning based voice extraction and primary-ambience decomposition for stereo to surround upmixing with dialog-enhanced center channel

Inventors: Sunil Bharitkar (Stevenson Ranch, CA); Ricardo Thaddeus Páez Amaro (Puebla, MX); Carlos Tejeda Ocampo (Tuxtla Gutierrez, MX); Luis Madrid Herrera (Chihuahua, MX)
Assignee: Samsung Electronics Co., Ltd.
H04S7/307H04S3/008H04S7/302H04S2400/01H04S2400/03H04S2400/05H04S2400/11H04S2400/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,581,266
App. No.
18/399,148
Granted
Mar 17, 2026
Kind
B2
Abstract

One embodiment provides a computer-implemented method that includes determining directional sounds from a content mix using a machine learning unmixing model. The directional sounds are panned in an upmixed signal. Signal-dependent upmixing gains for specific frequency bins are computed on a frame-basis using a machine learning model for the upmixed signal. Dedicated voice clarity gains are computed using a hearing impairment model for multiple hearing-impaired profiles for achieving dialog enhancement. The signal dependent upmixing gains and voice clarity gains are transmitted as metadata with a downmixed signal representing the content mix.

Claims (43)

1 . A computing method comprising:

determining directional sounds from a content mix using a machine learning unmixing model;

panning the directional sounds in an upmixed signal;

computing signal-dependent upmixing gains for specific frequency bins on a frame-basis using a machine learning model for the upmixed signal; and

computing dedicated voice clarity gains using a hearing impairment model for a plurality of hearing-impaired profiles for achieving dialog enhancement;

wherein the signal dependent upmixing gains and voice clarity gains are transmitted as metadata with a downmixed signal representing the content mix.

2 . The method of claim 1 , further comprising:

performing, by the computing device, a primary-ambience decomposition process for the upmixed signal.

3 . The method of claim 2 , further comprising:

applying the signal-dependent upmixing gains to downmixed signal components.

4 . The method of claim 2 , wherein the content mix comprises a voice content mix.

5 . The method of claim 2 , wherein during upmixing, the signal-dependent upmixing gains are applied to primary and ambient signals to generate a final output.

6 . The method of claim 2 , wherein the signal-dependent upmixing gains are embedded as audio-codec metadata.

7 . The method of claim 6 , wherein the audio-codec metadata is transmitted with encoded downmixed stereo signals.

8 . A non-transitory processor-readable medium that includes a program that when executed by a processor performs dialog enhancement of extracted sources of an unmixed signal, comprising:

determining, by the processor, directional sounds from a content mix using a machine learning unmixing model;

panning, by the processor, the directional sounds in an upmixed signal;

computing, by the processor, signal-dependent upmixing gains for specific frequency bins on a frame-basis using a machine learning model for the upmixed signal; and

computing, by the processor, dedicated voice clarity gains using a hearing impairment model for a plurality of hearing-impaired profiles for achieving dialog enhancement;

wherein the signal dependent upmixing gains and voice clarity gains are transmitted as metadata with a downmixed signal representing the content mix.

9 . The non-transitory processor-readable medium of claim 8 , further comprising:

performing, by the processor, a primary-ambience decomposition process for the upmixed signal.

10 . The non-transitory processor-readable medium of claim 9 , further comprising:

applying the signal-dependent upmixing gains to downmixed signal components.

11 . The non-transitory processor-readable medium of claim 9 , wherein the content mix comprises a voice content mix.

12 . The non-transitory processor-readable medium of claim 9 , wherein during upmixing, the signal-dependent upmixing gains are applied to primary and ambient signals to generate a final output.

13 . The non-transitory processor-readable medium of claim 9 , wherein the signal-dependent upmixing gains are embedded as audio-codec metadata.

14 . The non-transitory processor-readable medium of claim 13 , wherein the audio-codec metadata is transmitted with encoded downmixed stereo signals.

15 . An apparatus comprising:

a memory storing instructions; and

at least one processor executes the instructions including a process configured to:

determine directional sounds from a content mix using a machine learning unmixing model;

pan the directional sounds in an upmixed signal;

compute signal-dependent upmixing gains for specific frequency bins on a frame-basis using a machine learning model for the upmixed signal; and

compute dedicated voice clarity gains using a hearing impairment model for a plurality of hearing-impaired profiles for achieving dialog enhancement;

wherein the signal dependent upmixing gains and voice clarity gains are transmitted as metadata with a downmixed signal representing the content mix.

16 . The apparatus of claim 15 , further comprising:

performing, by the computing device, a primary-ambience decomposition process for the upmixed signal.

17 . The apparatus of claim 16 , further comprising:

applying the signal-dependent upmixing gains to downmixed signal components.

18 . The apparatus of claim 16 , wherein the content mix comprises a voice content mix.

19 . The apparatus of claim 16 , wherein during upmixing, the signal-dependent upmixing gains are applied to primary and ambient signals to generate a final output.

20 . The apparatus of claim 16 , wherein the signal-dependent upmixing gains are embedded as audio-codec metadata, and the audio-codec metadata is transmitted with encoded downmixed stereo signals.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2023
From: BHARITKAR, SUNIL; PAEZ AMARO, RICARDO THADDEUS; TEJEDA OCAMPO, CARLOS; MADRID HERRERA, LUIS
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 065974/0480 →
Continuity (2)
Provisional Application 63443769 · Feb 7, 2023
Related Publication 20240267701A1 · Aug 8, 2024
References Cited (22)
US 9053700B2 · Neusinger · 2015 [cited by examiner]
US 9093063B2 · Mlkamo et al. · 2015 [cited by applicant]
US 9393412B2 · Strahl et al. · 2016 [cited by applicant]
US 10362426B2 · Wang et al. · 2019 [cited by applicant]
US 10453464B2 · Wang et al. · 2019 [cited by applicant]
US 10650836B2 · Wang · 2020 [cited by examiner]
US 11089423B2 · Kim et al. · 2021 [cited by applicant]
US 11297454B2 · Said · 2022 [cited by applicant]
US 11470438B2 · Uhle et al. · 2022 [cited by applicant]
US 11564048B2 · Bramslow · 2023 [cited by applicant]
US 20220392461A1 · Giron et al. · 2022 [cited by applicant]
US 20230105623A1 · Fukui et al. · 2023 [cited by applicant]
CN 114203163A · 2002 [cited by applicant]
CN 112866896B · 2022 [cited by applicant]
CN 115315747A · 2022 [cited by applicant]
WO 2022200136A1 · 2022 [cited by applicant]
WO 2023118078A1 · 2023 [cited by applicant]
International Search Report and Written Opinion dated May 2, 2024 for International Application PCT/KR2024/001539, from Korean Intellectual Property Office, pp. 1-11, Republic of Korea. [cited by applicant]
{Grace Period Disclosure}: “Deep Learning Based Voice Extraction And Primary-Ambience Decomposition For Stereo To Surround Upmixing,” Ricardo Thaddeus Paez-Amaro, Carlos Tejeda-Ocampo, Ema Souza-Blanes, Sunil Bharitkar,… [cited by applicant]
World Health Organization, “Deafness and Hearing Loss”, Feb. 27, 2023, pp. 1-5, downloaded Dec. 15, 23 from: https://www.who.int/news-room/fact-sheets/detail/deafness-and-hearing-loss, United States. [cited by applicant]
Monson, BB et al., “The perceptual significance of high-frequency energy in the human voice”. Front Psychol. Jun. 16, 2014, 2014, pp. 1-11, Article 587, Switzerland. [cited by applicant]
Karjalainen, M. et al., “Warped filters and their audio applications,” Proc. 1997 IEEE Int. Workshop Applications of Signal Processing to Audio & Acoustics, Oct. 19, 1997, pp. 1-4, IEEE, United States. [cited by applicant]