IP Library Granted Patent US 12,198,671
Granted Patent B2
US 12,198,671 · App. 18/309,754 · Granted Jan 14, 2025

Adaptive text-to-speech outputs based on language proficiency

Inventors: Matthew Sharifi (Kilchberg, CH); Jakob Nicolaus Foerster (San Francisco, CA)
Assignee: Google LLC
G10L13/00G06F40/253G06F40/289G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,671
App. No.
18/309,754
Granted
Jan 14, 2025
Kind
B2
Abstract

In some implementations, a language proficiency of a user of a client device is determined by one or more computers. The one or more computers then determines a text segment for output by a text-to-speech module based on the determined language proficiency of the user. After determining the text segment for output, the one or more computers generates audio data including a synthesized utterance of the text segment. The audio data including the synthesized utterance of the text segment is then provided to the client device for output.

Claims (54)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

obtaining previous text queries submitted by a user of a client device;

determining a language proficiency for the user based on the previous text queries;

receiving a query input to the client device by the user;

generating a particular text segment responsive to the query and based on the language proficiency determined for the user, the particular text segment comprising one of:

a first text segment when the language proficiency determined for the user comprises a first level of language proficiency, the first text segment comprising a respective independent clause conveying primary information responsive to the query; or

a second text segment when the language proficiency determined for the user comprises a second level of language proficiency, the second text segment comprising a respective independent clause and one or more subordinate clauses, the one or more subordinate clauses of the second text segment conveying additional information responsive to the query that is not included in the first text segment;

generating audio data comprising a synthesized utterance of the particular text segment responsive to the query; and

providing the audio data for audible output from the client device.

2. The computer-implemented method of claim 1 , wherein the respective independent clause of the second text segment conveys the same primary information responsive to the query as the first text segment.

3. The computer-implemented method of claim 1 , wherein the respective independent clause of the second text segment includes at least one different term than the respective independent clause of the first text segment.

4. The computer-implemented method of claim 1 , wherein the operations further comprise, prior to generating the particular text segment:

identifying multiple candidate text segments that are responsive to the query, each candidate text segment associated with a different level of language complexity; and

selecting, from among the multiple candidate text segments, the particular text segment responsive to the query based on the language proficiency determined for the user.

5. The computer-implemented method of claim 4 , wherein selecting from among the multiple candidate text segments comprises:

determining a language complexity score for each of the multiple candidate text segments; and

selecting the text segment associated with the language complexity score that best matches a reference score that describes the language proficiency determined for the user as the particular text segment.

6. The computer-implemented method of claim 1 , wherein the operations further comprise, prior to generating the particular text segment:

obtaining a baseline text segment responsive to the query; and

generating the particular text segment by increasing a complexity level of the baseline text segment based on the language proficiency designated to the user.

7. The computer-implemented method of claim 1 , wherein the operations further comprise, prior to generating the particular text segment:

obtaining a baseline text segment responsive to the query; and

generating the particular text segment by decreasing a complexity level of the baseline text segment based on the language proficiency designated to the user.

8. The computer-implemented method of claim 1 , wherein:

the second level of language proficiency comprises a higher level of language proficiency than the first level of language proficiency; and

the second text segment is associated with a grammatical structure that is more complex than a grammatical structure associated with the first text segment.

9. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions, that when executed by the data processing hardware, cause the data processing hardware to perform operations comprising:

obtaining previous text queries submitted by a user of a client device;

determining a language proficiency for the user based on the previous text queries;

receiving a query input to the client device by the user;

generating a particular text segment responsive to the query and based on the language proficiency determined for the user, the particular text segment comprising one of:

a first text segment when the language proficiency determined for the user comprises a first level of language proficiency, the first text segment comprising a respective independent clause conveying primary information responsive to the query; or

a second text segment when the language proficiency determined for the user comprises a second level of language proficiency, the second text segment comprising a respective independent clause and one or more subordinate clauses, the one or more subordinate clauses of the second text segment conveying additional information responsive to the query that is not included in the first text segment;

generating audio data comprising a synthesized utterance of the particular text segment responsive to the query; and

providing the audio data for audible output from the client device.

10. The system of claim 9 , wherein the respective independent clause of the second text segment conveys the same primary information responsive to the query as the first text segment.

11. The system of claim 9 , wherein the respective independent clause of the second text segment includes at least one different term than the respective independent clause of the first text segment.

12. The system of claim 9 , wherein the operations further comprise, prior to generating the particular text segment:

identifying multiple candidate text segments that are responsive to the query, each candidate text segment associated with a different level of language complexity; and

selecting, from among the multiple candidate text segments, the particular text segment responsive to the query based on the language proficiency determined for the user.

13. The system of claim 12 , wherein selecting from among the multiple candidate text segments comprises:

determining a language complexity score for each of the multiple candidate text segments; and

selecting the text segment associated with the language complexity score that best matches a reference score that describes the language proficiency determined for the user as the particular text segment.

14. The system of claim 9 , wherein the operations further comprise, prior to generating the particular text segment:

obtaining a baseline text segment responsive to the query; and

generating the particular text segment by increasing a complexity level of the baseline text segment based on the language proficiency designated to the user.

15. The system of claim 9 , wherein the operations further comprise, prior to generating the particular text segment:

obtaining a baseline text segment responsive to the query; and

generating the particular text segment by decreasing a complexity level of the baseline text segment based on the language proficiency designated to the user.

16. The system of claim 9 , wherein:

the second level of language proficiency comprises a higher level of language proficiency than the first level of language proficiency; and

the second text segment is associated with a grammatical structure that is more complex than a grammatical structure associated with the first text segment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2023
From: SHARIFI, MATTHEW; FOERSTER, JAKOB NICOLAUS
To: GOOGLE INC.
Reel/Frame 063486/0437 →
CHANGE OF NAME Recorded Apr 28, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 063499/0994 →
Continuity (7)
Continuation 17153463 · Jan 20, 2021
Continuation 16573492 · Sep 17, 2019
Continuation 16135885 · Sep 19, 2018
Continuation 15653872 · Jul 19, 2017
Continuation 15477360 · Apr 3, 2017
Continuation 15009432 · Jan 28, 2016
Related Publication 20230267911A1 · Aug 24, 2023
References Cited (56)
US 5870709A · Bernstein · 1999 [cited by applicant]
US 6029156A · Lannert · 2000 [cited by examiner]
US 7096183B2 · Junqua · 2006 [cited by applicant]
US 8744855B1 · Rausch · 2014 [cited by applicant]
US 9799324B2 · Sharifi et al. · 2017 [cited by applicant]
US 9886942B2 · Sharifi et al. · 2018 [cited by applicant]
US 10109270B2 · Sharifi et al. · 2018 [cited by applicant]
US 10453441B2 · Sharifi · 2019 [cited by examiner]
US 10923100B2 · Sharifi · 2021 [cited by examiner]
US 11670281B2 · Sharifi · 2023 [cited by examiner]
US 20010049602A1 · Walker et al. · 2001 [cited by applicant]
US 20030084015A1 · Beams · 2003 [cited by examiner]
US 20040117180A1 · Rajput et al. · 2004 [cited by applicant]
US 20040193421A1 · Blass · 2004 [cited by applicant]
US 20050015307A1 · Simpson et al. · 2005 [cited by applicant]
US 20050033582A1 · Gadd et al. · 2005 [cited by applicant]
US 20050175970A1 · Dunlap · 2005 [cited by examiner]
US 20060229873A1 · Eide et al. · 2006 [cited by applicant]
US 20070238076A1 · Burstein et al. · 2007 [cited by applicant]
US 20080162471A1 · Bernard · 2008 [cited by applicant]
US 20100324894A1 · Potkonjak · 2010 [cited by applicant]
US 20110093271A1 · Bernard · 2011 [cited by applicant]
US 20130031476A1 · Coin et al. · 2013 [cited by applicant]
US 20130060763A1 · Chica · 2013 [cited by examiner]
US 20130080173A1 · Talwar et al. · 2013 [cited by applicant]
US 20130238580A1 · D'Orazio Pedro De Matos · 2013 [cited by examiner]
US 20130275138A1 · Gruber et al. · 2013 [cited by applicant]
US 20130325482A1 · Tzirkel-Hancock et al. · 2013 [cited by applicant]
US 20140125558A1 · Miyajima et al. · 2014 [cited by applicant]
US 20140172418A1 · Puppin · 2014 [cited by applicant]
US 20140282098A1 · McConnell · 2014 [cited by examiner]
US 20150332665A1 · Mishra et al. · 2015 [cited by applicant]
JP 03035296A · 1991 [cited by applicant]
JP 2810750B2 · 1998 [cited by applicant]
JP 3225389B2 · 2001 [cited by applicant]
JP 2002171348A · 2002 [cited by applicant]
JP 2002312386A · 2002 [cited by applicant]
JP 2003225389A · 2003 [cited by applicant]
JP 2004193421A · 2004 [cited by applicant]
JP 2006330629A · 2006 [cited by applicant]
JP 2010145873A · 2010 [cited by applicant]
JP 2011100191A · 2011 [cited by applicant]
JP 20140125558A · 2014 [cited by applicant]
JP 2014199323A · 2014 [cited by applicant]
JP 5727810B2 · 2015 [cited by applicant]
KR 1020120120316A · 2012 [cited by applicant]
Invitation to Pay Additional Fees and Where Applicable Protest Fee, with Partial Search Report, May 4, 2017, 8 pages. [cited by applicant]
Janarthanam et al. “Adaptive generation in dialogue systems using dynamic user modeling,” Computational Lirnmistics, MIT Press, vol. 40, No. 4, Dec. 1, 2014, 38 pages. [cited by applicant]
Komatani et al. “Flexible Spoken Dialogue System based on User Models and Dynamic Generation of VoiceXML Scripts,” SIGDIAL, Jan. 1, 2003, 10 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/069182, mailed on Jun. 26, 2017, 21 pages. [cited by applicant]
Japanese Office Action for the related Application No. 2018-539396 dated Aug. 30, 2019. [cited by applicant]
Japanese Office Action for the related Application No. 10-2018-7021923 dated Jul. 27, 2018. [cited by applicant]
Japanese Office Action for the related Application No. 2018-539396 dated Dec. 17, 2019. [cited by applicant]
Korean Office Action for the relatead Application No. 10-2020-7001577 dated Apr. 14, 2020. [cited by applicant]
Japanese Office Action, Application No. 2020-076068, Jan. 18, 2021, 8 pages. [cited by applicant]
USPTO. Office Action relating to U.S. Appl. No. 17/153,463, dated Nov. 2, 2022. [cited by applicant]