IP Library › Granted Patent US 12,633,230
Granted Patent B2
US 12,633,230 · App. 17/926,136 · Granted May 19, 2026

Electronic device, method and computer program

Inventors: Falk-Martin Hoffmann (Stuttgart, DE); Giorgio Fabbro (Stuttgart, DE); Marc Ferras Font (Stuttgart, DE); Thomas Kemp (Stuttgart, DE); Stefan Uhlich (Stuttgart, DE)
Assignee: Sony Group Corporation
G09B15/00G09B5/02G09B5/04G10H1/0008G10H1/361G10H1/46G10L21/0208G10L21/028G10L25/78G10H2210/005G10H2210/066G10H2210/071G10H2210/091G10H2220/011G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,230
App. No.
17/926,136
Granted
May 19, 2026
Kind
B2
Abstract

An electronic device having a circuitry configured to perform audio source separation on an audio input signal to obtain a vocals signal and an accompaniment signal and to perform a confidence analysis on a user's voice signal based on the vocals signal to provide guidance to the user.

Claims (29)

1 . An electronic device comprising:

circuitry configured to

perform audio source separation on an audio input signal (x(n)) to obtain a vocals signal (s O (n)) and an accompaniment signal (s A (n)); and

perform a confidence analysis on a user's voice signal (s U (n)) based on the vocals signal (s O (n)) to provide real-time guidance to the user by adjusting a level of the vocals signal in real-time based on the confidence analysis such that the level of the vocals signal is reduced as a confidence value resulting from the confidence analysis increases.

2 . The electronic device of claim 1 , wherein the circuitry is further configured to obtain an adjusted vocals signal (s′ O (n)) from the vocals signal (s O (n)) and to perform mixing of the adjusted vocals signal (s′ O (n)) with the accompaniment signal (s A (n)), to obtain an adjusted audio signal (s″ O (n)) for providing guidance to the user.

3 . The electronic device of claim 2 , wherein the circuitry is further configured to play back the adjusted audio signal (s″ O (n)) for providing guidance to the user.

4 . The electronic device of claim 2 , wherein the circuitry is further configured to perform a gain control on the vocals signal (s O (n)) based on the confidence analysis to obtain the adjusted vocals signal (s′ O (n)).

5 . The electronic device of claim 1 , wherein the circuitry is further configured to generate a guidance control signal based on the confidence analysis and to perform a visual or audio guidance based on the guidance control signal for providing guidance to the user.

6 . The electronic device of claim 1 , wherein the circuitry is further configured to perform pitch analysis on the vocals signal (s O (n)) to obtain a vocals pitch analysis result (ω fO (n)), to perform a pitch analysis ( 303 ) on the user's voice (s U (n)) to obtain a user's pitch analysis result (ω fU (n)), and to perform a vocals pitch comparison based on the vocals pitch analysis result (ω fO (n)) and the user's pitch analysis result (ω fU (n) to obtain a pitch error (e P (n)), wherein the pitch analysis includes calculating a frequency ration in cents and performing a modulo-600 operation to maintain the frequency ratio within half an octave.

7 . The electronic device of claim 1 , wherein the circuitry is further configured to perform rhythm analysis on the vocals signal (s O (n)) to obtain a vocals rhythm analysis result ({tilde over (d)} O (n)), to perform a rhythm analysis on the user's voice (s U (n)) to obtain a user's rhythm analysis result ({tilde over (d)} U (n)), and to perform a vocals rhythm comparison based on the vocals rhythm analysis result ({tilde over (d)} O (n)) and the user's rhythm analysis result ({tilde over (d)} U (n)) to obtain a rhythm error (e R (n)).

8 . The electronic device of claim 5 , wherein the circuitry is further configured to perform a confidence estimation based on the pitch analysis result (ω fO (n), ω fU (n)) and based on the rhythm analysis result ({tilde over (d)} O (n), {tilde over (d)} U (n)) to obtain a confidence value (e tot (n)).

9 . The electronic device of claim 8 , wherein the circuitry is further configured to perform a confidence-to-gain mapping based on the confidence value (e tot (n)) to obtain a gain control signal.

10 . The electronic device of claim 8 , wherein the circuitry is further configured to perform a guidance logic based on the confidence value (e tot (n)) to obtain a guidance control signal.

11 . An electronic device comprising:

circuitry configured to

perform audio source separation on an audio input signal (x(n)) to obtain a vocals signal (s O (n)) and an accompaniment signal (s A (n));

perform a confidence analysis on a user's voice signal (s U (n)) based on the vocals signal (s O (n)) to provide guidance to the user; and

perform voice activity detection on the user's voice signal (s U (n)) to obtain a trigger signal that triggers the confidence analysis.

12 . The electronic device of claim 8 , wherein the circuitry is configured to provide guidance to the user based on the confidence value (e tot (n)).

13 . The electronic device of claim 9 , wherein the confidence-to-gain mapping is configured to set the gain control signal in such a way that the user receives no guidance if the user is singing in perfect pitch.

14 . The electronic device of claim 1 , wherein the audio input signal (x(n)) comprises a mono and/or stereo audio input signal (x(n)) or the audio input signal (x(n)) comprises an object-based audio input signal.

15 . The electronic device of claim 1 , wherein the circuitry is further configured to perform echo cancellation on the user's voice (s U (n)) to obtain an echo free user's voice.

16 . The electronic device of claim 1 , wherein the circuitry is further configured to perform a vocal characteristic analysis on the vocals signal (s O (n)) to obtain a vocal characteristic analysis result, to perform a vocal characteristic analysis on the user's voice (s U (n)) to obtain a user's vocal characteristic analysis result ({tilde over (d)} U (n)).

17 . The electronic device of claim 1 , wherein the circuitry comprises a microphone configured to capture the user's vocals signal (s U (n)).

18 . A method comprising:

performing audio source separation on an audio input signal (x(n)) to obtain a vocals signal (s O (n)) and an accompaniment signal (s A (n)); and

performing confidence analysis on a user's voice (s U (n)) based on the vocals signal (s O (n)) to provide real-time guidance to the user by adjusting a level of the vocals signal in real-time based on the confidence analysis such that the level of the vocals signal is reduced as a confidence value resulting from the confidence analysis increases.

19 . A non-transitory computer-readable medium including computer program comprising instructions, the instructions when executed by circuitry causing the circuitry to perform the method of claim 18 .

20 . The electronic device of claim 1 , wherein the circuitry is configured to trigger the confidence analysis when the user starts singing above a predetermined threshold and stop the confidence analysis when the user stops singing or lowers voice below the predetermined threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2022
From: HOFFMANN, FALK-MARTIN; FABBRO, GIORGIO; FONT, MARC FERRAS; KEMP, THOMAS; UHLICH, STEFAN
To: SONY GROUP CORPORATION
Reel/Frame 061819/0389 →
Priority Claims (1)
EP 20178622 · Jun 5, 2020 · regional
Continuity (1)
Related Publication 20230186782A1 · Jun 15, 2023
References Cited (35)
US 5525062A · Ogawa et al. · 1996 [cited by applicant]
US 8138409B2 · Brennan · 2012 [cited by applicant]
US 8148621B2 · Bright · 2012 [cited by examiner]
US 20040177744A1 · Strasser et al. · 2004 [cited by applicant]
US 20060112812A1 · Venkataraman et al. · 2006 [cited by applicant]
US 20090038467A1 · Brennan · 2009 [cited by applicant]
US 20090038468A1 · Brennan · 2009 [cited by examiner]
US 20100107856A1 · Hetherington et al. · 2010 [cited by applicant]
US 20100192752A1 · Bright · 2010 [cited by examiner]
US 20160049915A1 · Wang · 2016 [cited by examiner]
US 20170140745A1 · Nayak et al. · 2017 [cited by applicant]
CN 103971674A · 2014 [cited by applicant]
CN 108257613A · 2018 [cited by applicant]
CN 109272975A · 2019 [cited by applicant]
CN 109300485A · 2019 [cited by applicant]
CN 110660383A · 2020 [cited by applicant]
CN 111046226A · 2020 [cited by applicant]
JP 2005242238A · 2005 [cited by applicant]
JP 4481225B2 · 2010 [cited by applicant]
WO 2005111997A1 · 2005 [cited by applicant]
Cheng et al., “Time-frequency Analysis of Musical Rhythm”, Notices of the American Mathematical Society, vol. 56, No. 3, Mar. 2009, pp. 356-372. [cited by applicant]
Li et al., “Separation of Singing Voice From Music Accompaniment For Monaural Recordings”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 15, No. 4, May 2007, pp. 1475-1487. [cited by applicant]
Ortiz P. et al., “A Simple but Efficient Voice Activity Detection Algorithm Through Hilbert Transform and Dynamic Threshold for Speech Pathologies”, Journal of Physics: Conference Series 705 (2016) 012037, 10 pages. [cited by applicant]
Cheveigne et al., “YIN, A Fundamental Frequency Estimator for Speech and Music”, J. Acoust. Soc. Am., vol. 111, No. 4, Apr. 2002, pp. 1917-1930. [cited by applicant]
Liu et al., “Fundamental Frequency Estimation Based on the Joint Time-Frequency Analysis of Harmonic Spectral Structure”, IEEE Transactions on Speech and Audio Processing, vol. 9, No. 6, Sep. 2001, pp. 609-621. [cited by applicant]
Uhlich et al., “Improving Music Source Separation Based on Deep Neural Networks Through Data Augmentation and Network Blending”, ICASSP 2017, pp. 261-265. [cited by applicant]
Bai et al., “Voice Activity Detection Based on Deep Neural Networks and Viterbi”, IOP Conference Series: Materials Science and Engineering 231 (2017) 012042, 7 pages. [cited by applicant]
Bello et al., “A Tutorial on Onset Detection in Music Signals”, IEEE Transactions on Speech and Audio Processing, 2005, pp. 1-13. [cited by applicant]
Jarne, “A Method for Estimation of Fundamental Frequency for Tonal Sounds Inspired on Bird Song Studies”, MethodsX, vol. 6, 2019, pp. 124-131. [cited by applicant]
Smith et al., “Time-Frequency Representation of Musical Rhythm by Continuous Wavelets”, Journal of Mathematics and Music, Jan. 2007, pp. 1-24. [cited by applicant]
Lane, John E., “Pitch Detection Using a Tunable IIR Filter”, Computer Music Journal, vol. 14, No. 3 (Autumn, 1990), pp. 46-59. [cited by applicant]
Tilsen et al., “Speech Rhythm Analysis with Decomposition of the Amplitude Envelope: Characterizing Rhythmic Patterns Within and Across Languages”, The Journal of the Acoustical Society of America—Jul. 2013, vol. 134, N… [cited by applicant]
Enzner et al., “Frequency-Domain Adaptive Kalman Filter for Acoustic Echo Control in Hands-Free Telephones”, Signal Processing, vol. 86, Issue 6, Jun. 2006, pp. 1140-1156. [cited by applicant]
Ikemiya et al., “Singing Voice Separation and Vocal F0 Estimation Based on Mutual Combination of Robust Principal Component Analysis and Subharmonic Summation”, IEEE/ACM Transactions On Audio, Speech, And Language Proce… [cited by applicant]
International Search Report and Written Opinion mailed on Sep. 7, 2021, received for PCT Application PCT/EP2021/065009, filed on Jun. 4, 2021, 12 pages. [cited by applicant]