IP Library › Granted Patent US 12,444,414
Granted Patent B2
US 12,444,414 · App. 17/247,414 · Granted Oct 14, 2025

Dynamic virtual assistant speech modulation

Inventors: Shikhar Kwatra (San Jose, CA); Komminist Weldemariam (Ottawa, CA); Zachary A. Silverstein (Austin, TX); Victor Povar (Vancouver, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G10L15/22G06F9/453G06N3/08G10L25/63G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,414
App. No.
17/247,414
Granted
Oct 14, 2025
Kind
B2
Abstract

A method, computer system, and a computer program product for dynamic speech modulation is provided. The present invention may include transmitting a first response to a received command. The present invention may include determining the first response is not understood by a user. The present invention may include transmitting a second response to the received command.

Claims (37)

1. A method for dynamic speech modulation, the method comprising:

extracting, by a digital voice assistant, Mel-frequency cepstral coefficient (MFCC) features from a received audio command from a user to identify the user and a corresponding user profile based on the MFCC features, wherein the corresponding user profile stores a default language of the user and one or more user pronunciation models learned by the digital voice assistant based on analysis of user pronunciations;

generating, by the digital voice assistant, a first response to the received audio command, wherein the first response includes a first pronunciation associated with a different language than the default language of the user;

sampling, by the digital voice assistant, the first response to determine whether the first response will be understood by the user based on a comprehension level of the user associated with the different language in the first response, wherein the comprehension level of the user is identified in the corresponding user profile, and wherein the comprehension level is dynamically set by the digital voice assistant based on learning a user comprehension rate for one or more prior unmodified responses from the digital voice assistant; and

in response to determining, by the digital voice assistant, that the first response includes a lower confidence score than the comprehension level of the user, performing a cosine similarity to identify at least one speech feature in the first pronunciation of the first response that is different relative to the one or more user pronunciation models learned by the digital voice assistant, modifying the at least one speech feature in the first pronunciation of the first response to align the first pronunciation with the one or more user pronunciation models learned by the digital voice assistant, and transmitting to the user, by the digital voice assistant, a second response to the received audio command generated based on the one or more user pronunciation models, wherein the second response includes a higher confidence score than the comprehension level of the user.

2. The method of claim 1 , further comprising:

transmitting to the user, by the digital voice assistant, the first response to the received audio command to determine whether the first response will be understood by the user; and

determining that the first response transmitted to the user is not understood by the user based on determining, by the digital voice assistant executing a sentiment analysis application programming interface (API) and a tone analysis API on a user response to the first response, that a frustration level of the user exceeds a baseline frustration level.

3. The method of claim 1 , further comprising:

using a long short-term memory (LSTM) recurrent neural network (RNN), trained using historical user data including pronunciation data, accent data, and the corresponding user profile, to predict a comprehension difficulty of the user.

4. The method of claim 1 , further comprising:

updating the one or more user pronunciation models based on a perceived understanding of the second response by the user.

5. A computer system for dynamic speech modulation, comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more computer-readable tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories, wherein the computer system is capable of performing a method comprising:

extracting, by a digital voice assistant, Mel-frequency cepstral coefficient (MFCC) features from a received audio command from a user to identify the user and a corresponding user profile based on the MFCC features, wherein the corresponding user profile stores a default language of the user and one or more user pronunciation models learned by the digital voice assistant based on analysis of user pronunciations;

generating, by the digital voice assistant, a first response to the received audio command, wherein the first response includes a first pronunciation associated with a different language than the default language of the user;

sampling, by the digital voice assistant, the first response to determine whether the first response will be understood by the user based on a comprehension level of the user associated with the different language in the first response, wherein the comprehension level of the user is identified in the corresponding user profile, and wherein the comprehension level is dynamically set by the digital voice assistant based on learning a user comprehension rate for one or more prior unmodified responses from the digital voice assistant; and

in response to determining, by the digital voice assistant, that the first response includes a lower confidence score than the comprehension level of the user, performing a cosine similarity to identify at least one speech feature in the first pronunciation of the first response that is different relative to the one or more user pronunciation models learned by the digital voice assistant, modifying the at least one speech feature in the first pronunciation of the first response to align the first pronunciation with the one or more user pronunciation models learned by the digital voice assistant, and transmitting to the user, by the digital voice assistant, a second response to the received audio command generated based on the one or more user pronunciation models, wherein the second response includes a higher confidence score than the comprehension level of the user.

6. The computer system of claim 5 , further comprising:

transmitting to the user, by the digital voice assistant, the first response to the received audio command to determine whether the first response will be understood by the user; and

determining that the first response transmitted to the user is not understood by the user based on determining, by the digital voice assistant executing a sentiment analysis application programming interface (API) and a tone analysis API on a user response to the first response, that a frustration level of the user exceeds a baseline frustration level.

7. The computer system of claim 5 , further comprising:

using a long short-term memory (LSTM) recurrent neural network (RNN), trained using historical user data including pronunciation data, accent data, and the corresponding user profile, to predict a comprehension difficulty of the user.

8. The computer system of claim 5 , further comprising:

updating the one or more user pronunciation models based on a perceived understanding of the second response by the user.

9. A computer program product for dynamic speech modulation, comprising:

one or more non-transitory computer-readable storage media and program instructions stored on at least one of the one or more non-transitory computer-readable storage media, the program instructions executable by a processor to cause the processor to perform a method comprising:

extracting, by a digital voice assistant, Mel-frequency cepstral coefficient (MFCC) features from a received audio command from a user to identify the user and a corresponding user profile based on the MFCC features, wherein the corresponding user profile stores a default language of the user and one or more user pronunciation models learned by the digital voice assistant based on analysis of user pronunciations;

generating, by the digital voice assistant, a first response to the received audio command, wherein the first response includes a first pronunciation associated with a different language than the default language of the user;

sampling, by the digital voice assistant, the first response to determine whether the first response will be understood by the user based on a comprehension level of the user associated with the different language in the first response, wherein the comprehension level of the user is identified in the corresponding user profile, and wherein the comprehension level is dynamically set by the digital voice assistant based on learning a user comprehension rate for one or more prior unmodified responses from the digital voice assistant; and

in response to determining, by the digital voice assistant, that the first response includes a lower confidence score than the comprehension level of the user, performing a cosine similarity to identify at least one speech feature in the first pronunciation of the first response that is different relative to the one or more user pronunciation models learned by the digital voice assistant, modifying the at least one speech feature in the first pronunciation of the first response to align the first pronunciation with the one or more user pronunciation models learned by the digital voice assistant, and transmitting to the user, by the digital voice assistant, a second response to the received audio command generated based on the one or more user pronunciation models, wherein the second response includes a higher confidence score than the comprehension level of the user.

10. The computer program product of claim 9 , further comprising:

transmitting to the user, by the digital voice assistant, the first response to the received audio command to determine whether the first response will be understood by the user; and

determining that the first response transmitted to the user is not understood by the user based on determining, by the digital voice assistant executing a sentiment analysis application programming interface (API) and a tone analysis API on a user response to the first response, that a frustration level of the user exceeds a baseline frustration level.

11. The computer program product of claim 9 , further comprising:

using a long short-term memory (LSTM) recurrent neural network (RNN), trained using historical user data including pronunciation data, accent data, and the corresponding user profile, to predict a comprehension difficulty of the user.

12. The computer program product of claim 9 , wherein the one or more user pronunciation models is updated based on a perceived understanding of the second response by the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2020
From: KWATRA, SHIKHAR; WELDEMARIAM, KOMMINIST; SILVERSTEIN, ZACHARY A.; POVAR, VICTOR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054606/0518 →
Continuity (1)
Related Publication 20220189475A1 · Jun 16, 2022
References Cited (39)
US 8275621B2 · Alewine · 2012 [cited by applicant]
US 9812151B1 · Amini · 2017 [cited by examiner]
US 11302312B1 · Soni · 2022 [cited by examiner]
US 20030191645A1 · Zhou · 2003 [cited by applicant]
US 20080147408A1 · Da Palma · 2008 [cited by applicant]
US 20090006097A1 · Etezadi · 2009 [cited by examiner]
US 20120166180A1 · Au · 2012 [cited by applicant]
US 20120173241A1 · Li · 2012 [cited by examiner]
US 20130289998A1 · Eller · 2013 [cited by examiner]
US 20140122083A1 · Xiaojiang · 2014 [cited by applicant]
US 20140122407A1 · Duan · 2014 [cited by applicant]
US 20160293159A1 · Belisario · 2016 [cited by examiner]
US 20160307569A1 · Peng · 2016 [cited by examiner]
US 20180174577A1 · Jothilingam · 2018 [cited by examiner]
US 20180196796A1 · Wu · 2018 [cited by examiner]
US 20180218750A1 · Nichkawde · 2018 [cited by examiner]
US 20180268309A1 · Childress · 2018 [cited by examiner]
US 20190189116A1 · Li · 2019 [cited by examiner]
US 20190325864A1 · Anders · 2019 [cited by examiner]
US 20200082806A1 · Kim · 2020 [cited by examiner]
US 20200126536A1 · Farivar · 2020 [cited by examiner]
US 20200251014A1 · Jones · 2020 [cited by examiner]
US 20200286467A1 · Chao · 2020 [cited by examiner]
US 20200286473A1 · Anders · 2020 [cited by examiner]
US 20200380882A1 · Alailima · 2020 [cited by examiner]
US 20200387603A1 · Weldemariam · 2020 [cited by examiner]
US 20210065702A1 · Fink · 2021 [cited by examiner]
US 20210120206A1 · Liu · 2021 [cited by examiner]
US 20210160373A1 · Mcgann · 2021 [cited by examiner]
US 20210295826A1 · Morabia · 2021 [cited by examiner]
US 20220122581A1 · Chen · 2022 [cited by examiner]
US 20220189475A1 · Kwatra · 2022 [cited by examiner]
US 20220284882A1 · Peddinti · 2022 [cited by examiner]
US 20220293124A1 · Weinberg · 2022 [cited by examiner]
Anonymous, “Accelerate change to smarter vehicles of the future with AI and IoT,” IBM.com, [accessed on Jul. 7, 2020], 8 pages, Retrieved from the Internet: <URL: https://www.ibm.com/internet-of-things/explore-iot/vehic… [cited by applicant]
Disclosed Anonymously, “Matching language and accent in virtual assistant responses,” IP.com, Jun. 14, 2014, 5 pages, IP.com No. IPCOM000254265D. [cited by applicant]
Genhart, “Google Home vs. Amazon Echo, round 2: Google strikes back,” CNET.com, May 18, 2017 [accessed on Jul. 7, 2020], 15 pages, Retrieved from the Internet: <URL: https://www.cnet.com/news/google-home-vs-amazon-echo/… [cited by applicant]
Jesdanun, “How Amazon Echo listens and what it stores,” Phys.org, Dec. 29, 2016 [accessed on Jul. 7, 2020], 3 pages, Retrieved from the Internet: <URL: https://phys.org/news/2016-12-amazon-echo.html#jCp>. [cited by applicant]
Mell, et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]