IP Library › Granted Patent US 12,367,890
Granted Patent B2
US 12,367,890 · App. 17/925,025 · Granted Jul 22, 2025

Audio source separation and audio dubbing

Inventors: Stefan Uhlich (Stuttgart, DE); Giorgio Fabbro (Stuttgart, DE); Marc Ferras Font (Stuttgart, DE); Falk-Martin Hoffmann (Stuttgart, DE); Thomas Kemp (Stuttgart, DE)
Assignee: Sony Group Corporation
G10L21/028G10H1/361G10L13/04G10L15/08G10L15/22G10L25/18G10H2210/005G10L2015/088H04S3/008H04S2400/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,890
App. No.
17/925,025
Filed
Nov 14, 2022
Granted
Jul 22, 2025
Kind
B2
Examiner
AZAD, ABUL K
Art Unit
2656
USPC
704/251
Abstract

An electronic device having a circuitry configured to perform audio source separation on an audio input signal to obtain a separated source and configured to perform audio dubbing on the separated source based on replacement conditions to obtain a personalized separated source.

Claims (25)

1. An electronic device, comprising:

circuitry configured to

perform audio source separation on an audio input signal to obtain a separated source and a residual signal, wherein the residual signal includes audio content from the audio input signal other than the separated source,

delay the separated source by a first delay amount to obtain a delayed separated source, wherein the first delay amount corresponds to a processing latency of phrase detection,

delay the residual signal by a second delay amount to obtain a delayed residual signal, wherein the second delay amount corresponds to a combined processing latency of the phrase detection and audio dubbing,

perform audio dubbing on the separated source based on replacement conditions to obtain a personalized separated source, and

mix the personalized separated source with the delayed residual signal to obtain a personalized audio signal.

2. The electronic device of claim 1 , wherein the circuitry is further configured to perform lyrics recognition on the separated source to obtain lyrics and to perform lyrics replacement on the lyrics based on the replacement conditions to obtain personalized lyrics.

3. The electronic device of claim 2 , wherein the circuitry is further configured to perform text-to-vocals synthesis on the personalized lyrics based on the separated source to obtain the personalized separated source.

4. The electronic device of claim 2 , wherein the circuitry is further configured to apply a Seq2Seq Model on the personalized lyrics based on the separated source to obtain a Mel-Spectrogram, and to apply a MelGAN generator on the Mel-Spectrogram to obtain the personalized separated source.

5. The electronic device of claim 1 , wherein the circuitry is further configured to perform the source separation on the audio input signal to obtain the separated source and a residual signal, and to perform mixing of the personalized separated source with the residual signal, to obtain a personalized audio signal.

6. The electronic device of claim 1 , wherein the circuitry is further configured to perform the audio dubbing on the separated source based on a trigger signal to obtain the personalized separated source.

7. The electronic device of claim 6 , wherein the circuitry is further configured to perform phrase detection on the separated source based on the replacement conditions obtain the trigger signal.

8. The electronic device of claim 7 , wherein the circuitry is further configured to perform speech recognition on the separated source to obtain transcript/lyrics.

9. The electronic device of claim 8 , wherein the circuitry is further configured to perform target phrase detection on the transcript/lyrics based on the replacement conditions to obtain the trigger signal.

10. The electronic device of claim 1 , wherein the separated source comprises vocals and the residual signal comprises an accompaniment.

11. The electronic device of claim 1 , wherein the separated source comprises speech and the residual signal comprises background noise.

12. The electronic device of claim 1 , wherein the replacement conditions are age dependent replacement conditions.

13. A method comprising:

performing audio source separation on an audio input signal to obtain a separated source and a residual signal, wherein the residual signal includes audio content from the audio input signal other than the separated source;

delaying the separated source by a first delay amount to obtain a delayed separated source, wherein the first delay amount corresponds to a processing latency of phrase detection;

delaying the residual signal by a second delay amount to obtain a delayed residual signal, wherein the second delay amount corresponds to a combined processing latency of the phrase detection and audio dubbing;

performing dubbing on the separated source based on replacement conditions to obtain a personalized separated source; and

mixing the personalized separated source with the delayed residual signal to obtain a personalized audio signal.

14. A computer program comprising instructions, the instructions when executed on a processor causing the processor to perform the method of claim 13 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2022
From: UHLICH, STEFAN; FABBRO, GIORGIO; FONT, MARC FERRAS; HOFFMANN, FALK-MARTIN; KEMP, THOMAS
To: SONY GROUP CORPORATION
Reel/Frame 061932/0831 →
Priority Claims (1)
EP 20177598 · May 29, 2020 · regional
Continuity (1)
Related Publication 20230186937A1 · Jun 15, 2023
References Cited (24)
US 11514885B2 · Gabryjelski · 2022 [cited by examiner]
US 20100070274A1 · Cho · 2010 [cited by examiner]
US 20100280828A1 · Fein et al. · 2010 [cited by applicant]
US 20110219940A1 · Jiang · 2011 [cited by applicant]
US 20130151251A1 · Herz et al. · 2013 [cited by applicant]
US 20150149183A1 · Hennequin · 2015 [cited by examiner]
US 20150380014A1 · Le Magoarou · 2015 [cited by examiner]
US 20160093316A1 · Pacquier et al. · 2016 [cited by applicant]
US 20180122403A1 · Koretzky · 2018 [cited by examiner]
US 20190102138A1 · Valdez · 2019 [cited by applicant]
US 20210174781A1 · Chen · 2021 [cited by examiner]
US 20210224319A1 · Ingel · 2021 [cited by examiner]
US 20210352380A1 · Duncan · 2021 [cited by examiner]
CN 106157485A · 2016 [cited by examiner]
EP 3239981A1 · 2017 [cited by applicant]
EP 3201917B1 · 2021 [cited by applicant]
WO 2008132265A1 · 2008 [cited by applicant]
Kumar et al., “MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,” Advances in Neural Information Processing Systems. 2019, 14 pages. [cited by applicant]
Shen et al., “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018, 5 pages. [cited by applicant]
Prenger et al., “Waveglow: A flow-Based Generative Network for Speech Synthesis,” ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, 5 pages. [cited by applicant]
Jia et al., “Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis,” Advances in Neural Information Processing Systems, 2018, 15 pages. [cited by applicant]
Jin et al., “VoCo: Text-based Insertion and Replacement in Audio Narration,” ACM Transactions on Graphics 36(4), Article 96, Jul. 2017, 13 pages. [cited by applicant]
Uhlich et al. “Improving Music Source Separation Based on Deep Neural Networks Through Data Augmentation and Network Blending,” 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEE… [cited by applicant]
International Search Report and Written Opinion mailed on May 20, 2021, received for PCT Application PCT/EP2021/056828, filed on Mar. 17, 2021, 9 pages. [cited by applicant]