IP Library › Granted Patent US 12,367,867
Granted Patent B2
US 12,367,867 · App. 17/973,266 · Granted Jul 22, 2025

System for generating voice in an ongoing call session based on artificial intelligent techniques

Inventors: Sandeep Singh Spall (Moga, IN); Tarun Gupta (Noida, IN); Narang Lucky Manoharlal (Noida, IN)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/16G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,867
App. No.
17/973,266
Granted
Jul 22, 2025
Kind
B2
Abstract

A method for generating voice in an ongoing call session based on artificial intelligent techniques is provided. The method includes extracting a plurality of features from a voice input through an artificial neural network (ANN); identifying one or more lost audio frames within the voice input; predicting by the ANN, for each of the one or more lost audio frames, one or more features of the respective lost audio frame; and superposing the predicted features upon the voice input to generate an updated voice input.

Claims (62)

1. A method of generating voice in a call session, the method comprising:

extracting a plurality of features from a voice input through an artificial neural network (ANN);

identifying one or more lost audio frames within the voice input, wherein the one or more lost audio frames are lost due to at least one of vocal issues or network issues;

predicting by the ANN, for each of the one or more lost audio frames, one or more features of the respective lost audio frame;

superposing the predicted features upon the voice input to generate an updated voice input; and

correcting the updated voice input by:

obtaining a confidence score of the updated voice input;

splitting the updated voice input into a plurality of phonemes based on the confidence score;

identifying one or more non-aligned phonemes out of the plurality of phonemes based on comparing the plurality of phonemes with language vocabulary knowledge;

generating a plurality of variant phonemes; and

updating the identified one or more non-aligned phonemes through one or more of:

replacing the identified one or more non-aligned phonemes with the plurality of variant phonemes;

adding additional phonemes to supplement the identified one or more non-aligned phonemes; or

deleting the identified one or more non-aligned phonemes.

2. The method of claim 1 , wherein the predicting by the ANN comprises:

operating a single-layer recurrent neural network for audio generation configured to predict raw audio samples.

3. The method of claim 2 , wherein the single-layer recurrent neural network comprises an architecture defined by WaveRNN and a dual softmax layer.

4. The method of claim 1 , further comprising:

updating the identified one or more non-aligned phonemes through regenerating the updated voice input defined by one or more of:

replacement of the identified one or more non-aligned phonemes with the variant phonemes;

removal of the identified one or more non-aligned phonemes; or

additional phonemes supplementing the identified one or more non-aligned phonemes.

5. The method of claim 1 , further comprising converting the updated voice input defined by a whisper into a converted updated voice input defined by a normal voice-input based on:

executing a time-alignment between whispered and a corresponding normal speech; and

learning by a generator adversarial network (GAN) model cross-domain relations between the whisper and the normal speech.

6. The method of claim 1 , further comprising:

improving voice quality of the updated voice input based on a nonlinear activation function.

7. A system for generating voice in a call session, the system comprising:

a microphone;

a memory storing one or more instructions; and

at least one processor configured to execute the one or more instructions to: extract a plurality of features from a voice input through an artificial neural network (ANN);

identify one or more lost audio frames within the voice input, wherein the one or more lost audio frames are lost due to at least one of vocal issues or network issues;

predict by the ANN the features of the one or more lost audio frames for each missing frame;

superpose the predicted features upon the voice input to generate an updated voice input; and

correct the updated voice input by:

obtaining a confidence score of the updated voice input;

splitting the updated voice input into a plurality of phonemes based on the confidence score;

identifying one or more non-aligned phonemes out of the plurality of phonemes based on comparing the plurality of phonemes with language vocabulary knowledge;

generating a plurality of variant phonemes; and

updating the identified one or more non-aligned phonemes through one or more of:

replacing the identified one or more non-aligned phonemes with the plurality of variant phonemes;

adding additional phonemes to supplement the identified one or more non-aligned phonemes; or

deleting the identified one or more non-aligned phonemes.

8. The system of claim 7 , wherein the processor is further configured to execute the one or more instructions to by the ANN by operating a single-layer recurrent neural network for audio generation configured to predict raw audio samples.

9. The system of claim 8 , wherein the single-layer recurrent neural network comprises an architecture defined by WaveRNN and a dual softmax layer.

10. The system of claim 7 , the processor is further configured to:

update the identified one or more non-aligned phonemes through regenerating the updated voice input defined by one or more of:

replacement of the identified one or more non-aligned phonemes with the variant phonemes;

removal of the identified one or more non-aligned phonemes; or

additional phonemes supplementing the identified one or more non-aligned phonemes.

11. The system of claim 7 , the processor is further configured to convert the updated voice input defined by a whisper into a converted updated voice input defined by a normal voice-input based on:

executing a time-alignment between whispered and a corresponding normal speech; and

learning by a generator adversarial network (GAN) model cross-domain relations between the whisper and the normal speech.

12. The system of claim 7 , the processor is further configured to improve voice quality of the updated voice input based on a nonlinear activation function.

13. The system of claim 7 , wherein the processor is further configured to:

extract one or more frames from the updated voice input;

encode a temporal information associated with the one or more frames through a convolution network;

execute a deconvolution over the encoded information to obtain a prediction score;

classify the extracted frames into key frames or non-key frames based on the prediction score; and

present a summary based on the classified key frames.

14. The system of claim 13 , the processor is configured to learn relation between raw audio frames (A) and a set of summary audios (S), wherein a distribution of resultant summary-audios F(A) is targeted to be similar to a distribution of the set of summary audios (S); and

train a summary discriminator to differentiate between the generated summary audio F(A) and a real summary audio.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2022
From: SPALL, SANDEEP SINGH; GUPTA, TARUN; MANOHARLAL, NARANG LUCKY
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061770/0213 →
Priority Claims (1)
IN 202111048934 · Oct 26, 2021 · national
Continuity (1)
Related Publication 20230130777A1 · Apr 27, 2023
References Cited (27)
US 10965806B1 · Haggerty · 2021 [cited by applicant]
US 20050027531A1 · Gleason · 2005 [cited by examiner]
US 20060129401A1 · Smith · 2006 [cited by examiner]
US 20150058005A1 · Khare · 2015 [cited by examiner]
US 20180315420A1 · Ash · 2018 [cited by examiner]
US 20190348021A1 · Trim · 2019 [cited by examiner]
US 20200175987A1 · Thomson et al. · 2020 [cited by applicant]
US 20220165280A1 · Liang · 2022 [cited by applicant]
US 20230206902A1 · Guo · 2023 [cited by examiner]
CN 111292768A · 2020 [cited by applicant]
CN 113053400A · 2021 [cited by applicant]
KR 1020140013606A · 2014 [cited by applicant]
WO 2018048934A1 · 2018 [cited by applicant]
Rochan, Mrigank, Linwei Ye, and Yang Wang. “Video summarization using fully convolutional sequence networks.” Proceedings of the European conference on computer vision (ECCV). 2018. (Year: 2018). [cited by examiner]
Dubey, Shiv Ram, Satish Kumar Singh, and Bidyut Baran Chaudhuri. “A comprehensive survey and performance analysis of activation functions in deep learning.” arXiv preprint arXiv:2109.14545 99 (2021). (Year: 2021). [cited by examiner]
Pascual, Santiago, Joan Serrá, and Jordi Pons. “Adversarial auto-encoding for packet loss concealment.” 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). IEEE, 2021. (Year: 2021). [cited by examiner]
Patel, Maitreya, et al. “Novel Inception-GAN for Whisper-to-Normal speech conversion.” Proceedings of 10th ISCA Speech Synthesis Workshop (SSW 10), Vienna, Austria. 2019. (Year: 2019). [cited by examiner]
Rochan, Mrigank, and Yang Wang. “Video summarization by learning from unpaired data.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019. (Year: 2019). [cited by examiner]
Stimberg, Florian, et al. “WaveNetEQ—Packet loss concealment with WaveRNN.” 2020 54th Asilomar Conference on Signals, Systems, and Computers. IEEE, 2020. (Year: 2020). [cited by examiner]
Munmi Dutta et al., “Closed-Set Text-Independent Speaker Identification System Using Multiple ANN Classifiers”, Proceedings of the 3rd International Conference on Frontiers of Intelligent Computing: Theory and Applicati… [cited by applicant]
Communication issued on Apr. 1, 2024 by the Intellectual Property India for Indian Patent Application No. 202111048934. [cited by applicant]
Harold, “Google Duo Has No Sound, Can't Hear Other Callers”, The Droid Guy, May 2020, https://thedroidguy.com/google-duo-has-no-sound-1130598, 17 pages total. [cited by applicant]
“How do I fix sound or video call problems on Google Duo?”, Google Duo Help, Oct. 2017, https://googleduohelp.com/question/how-do-i-fix-sound-or-video-call-problems-on-google-duo/, 2 pages total. [cited by applicant]
“How to improve the quality of your Microsoft Teams calls”, HR future, Jul. 2020, https://www.hrfuture.net/future-of-work/trending/how-to-improve-the-quality-of-your-microsoft-teams-calls, 5 pages total. [cited by applicant]
Mitul Patel, “How to improve WhatsApp call quality for voice calling feature”, KADVACORP, Oct. 2018, https://www.kadvacorp.com/technology/improve-whatsapp-call-quality-voice-calling, 4 pages total. [cited by applicant]
“Why am I experiencing issues with WhatsApp Calling?”, WhatsApp Web, Jan. 2021, https://faq.whatsapp.com/general/why-am-i-experiencing-issues-with-whatsapp-calling, 2 pages total. [cited by applicant]
Elena Malykhina, “Why Is Cell Phone Call Quality So Terrible?”, Scientific American, Springer Nature, Apr. 2015, https://www.scientificamerican.com/article/why-is-cell-phone-call-quality-so-terrible/, 3 pages total. [cited by applicant]