IP Library › Granted Patent US 12,205,584
Granted Patent B1
US 12,205,584 · App. 17/532,986 · Granted Jan 21, 2025

Dialog-driven applications supporting alternative vocal input styles

Inventors: John Baker (Bellevue, WA); Anubhav Mishra (Seattle, WA); Bangrui Liu (Seattle, WA); Christopher Michael Hittner (Seattle, WA); Sravan Babu Bodapati (Redmond, WA); Harshal Pimpalkhute (Redmond, WA); Katrin Kirchhoff (Seattle, WA); Anuj Gautam Surana (Mountlake Terrace, WA); Yilai Su (Mercer Island, WA); Brandon Louis Mendez (Seattle, WA); Chengshun Zhang (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G10L15/22G10L13/027G10L15/08G10L2015/221G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,584
App. No.
17/532,986
Granted
Jan 21, 2025
Kind
B1
Abstract

A set of alternative vocal input styles for specifying a parameter of a dialog-driven application is determined. During execution of the application, an audio prompt requesting input in one of the styles is presented. A value of the parameter is determined by applying a collection of analysis tools to vocal input obtained after the prompt is presented. A task of the application is initiated using the value.

Claims (42)

1. A system, comprising:

one or more computing devices;

wherein the one or more computing devices include instructions that upon execution on or across the one or more computing devices:

determine, at a dialog-driven application management service, (a) a first set of alternative vocal input styles for specifying a value of a particular parameter of a dialog-driven application, wherein the first set of alternative vocal input styles includes a word-pronunciation style, a pronounce-each-letter-separately style and a spell-using-example-words style and (b) a default sequence, of at least a subset of alternative vocal input styles of the first set, in which input associated with the particular parameter is to be requested from a client of the dialog-driven application until a value of the particular parameter is determined;

cause to be presented, by the dialog-driven application management service during an execution of the dialog-driven application, in accordance with the default sequence, an audio prompt requesting input in a particular alternative vocal input style of the set of alternative vocal input styles;

obtain, at the dialog-driven application management service subsequent to presentation of at least a portion of the audio prompt, vocal input provided by a client of the dialog-driven application, wherein the vocal input is provided at least in part in the particular alternative vocal input style;

determine, at the dialog-driven application management service based at least in part on applying a collection of analysis tools to the vocal input provided by the client, a value of the particular parameter, wherein an indication of the particular alternative vocal input style is passed to the collection of analysis tools and utilized by the collection of analysis tools to process the vocal input; and

initiate, by the dialog-driven application management service using the value of the particular parameter, a task associated with the dialog-driven application.

2. The system as recited in claim 1 , wherein at least one alternative vocal input style of the set of alternative vocal input styles is determined automatically by the dialog-driven application management service without receiving an indication of the alternative vocal input style from a developer of the dialog-driven application or an owner of the dialog-driven application.

3. The system as recited in claim 1 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices:

determine, at the dialog-driven application management service, a second set of alternative vocal input styles for specifying a value of another parameter of a dialog-driven application, wherein the first set of alternative vocal input styles includes at least one alternative vocal style which is not a member of the second set.

4. The system as recited in claim 1 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices:

cause to be presented, to an end user of the dialog-driven application, an indication of a plurality of alternative vocal input styles for specifying a value of another parameter of the dialog-driven application; and

determine, prior to analyzing vocal input pertaining to the other parameter, that the end user has selected a first alternative vocal input style of the plurality of alternative vocal input styles for specifying the value of the other parameter.

5. The system as recited in claim 1 , wherein the collection of analysis tools comprises at least one automated speech recognition tool and at least one natural language understanding tool.

6. A computer-implemented method, comprising:

determining a set of alternative vocal input styles for specifying a value of a particular parameter of a dialog-driven application, wherein the set of alternative vocal input styles includes a pronounce-each-letter-separately style;

causing to be presented, during an execution of the dialog-driven application, an audio prompt requesting input in a particular alternative vocal input style of the set of alternative vocal input styles;

determining, based at least in part on applying a collection of analysis tools to vocal input obtained subsequent to presentation of the audio prompt, a value of the particular parameter; and

initiating, using the value of the particular parameter, a task of the dialog-driven application.

7. The computer-implemented method as recited in claim 6 , wherein the set of alternative vocal input styles includes one or more of: (a) a word-pronunciation style, (b) a spell-using-example-words style, or (c) a custom style associated with a problem domain of the dialog-driven application.

8. The computer-implemented method as recited in claim 6 , further comprising:

obtaining an indication of a rule to be applied to determine that input in the particular alternative vocal input style is to be requested, wherein the audio prompt is generated in accordance with the rule.

9. The computer-implemented method as recited in claim 8 , wherein according to the rule, the input in the particular alternative vocal input style is to be requested after an attempt to determine the value of the particular parameter based on analysis of vocal input in at least one other alternative vocal input style fails.

10. The computer-implemented method as recited in claim 8 , wherein the indication of the rule is obtained via a programmatic interface from a developer of the dialog-driven application.

11. The computer-implemented method as recited in claim 6 , further comprising:

obtaining an indication, via a programmatic interface from a developer of the dialog-driven application, of at least one alternative vocal input style of the set of alternative vocal input styles.

12. The computer-implemented method as recited in claim 6 , wherein the particular parameter comprises a portion of one or more of: (a) a name of an end user, (b) a street address, (c) an email address, (d) a postal code, or (e) an alphanumeric identifier.

13. The computer-implemented method as recited in claim 6 , wherein the collection of analysis tools comprises one or more of: (a) an automated speech recognition (ASR) tool or (b) a natural language understanding (NLU) tool.

14. The computer-implemented method as recited in claim 6 , wherein the collection of analysis tools comprises one or more of: (a) a finite state transducer (FST), (b) a statistical n-gram model or (c) a neural network-based model.

15. The computer-implemented method as recited in claim 6 , wherein the audio prompt indicates a successfully-interpreted portion of a previous vocal input received prior to presentation of the audio prompt, and wherein a request in the audio prompt for input in the particular alternative vocal input style pertains to a clarification of another portion of the previous vocal input.

16. One or more non-transitory computer-accessible storage media storing program instructions that when executed on or across one or more processors cause the one or more processors to:

determine a set of alternative vocal input styles for specifying a value of a particular parameter of a dialog-driven application, wherein the set of alternative vocal input styles includes a pronounce-each-letter-separately style;

cause to be presented, during an execution of the dialog-driven application, an audio prompt requesting input in a particular alternative vocal input style of the set of alternative vocal input styles;

determine, based at least in part on applying a collection of analysis tools to vocal input obtained subsequent to presentation of the audio prompt, a value of the particular parameter; and

initiate, using the value of the particular parameter, a task of the dialog-driven application.

17. The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the set of alternative vocal input styles includes one or more of: (a) a word-pronunciation style, (b) a spell-using-example-words style, or (c) a custom style associated with a problem domain of the dialog-driven application.

18. The one or more non-transitory computer-accessible storage media as recited in claim 16 , storing further program instructions that when executed on or across one or more processors further cause the one or more processors to:

obtain an indication of a rule to be applied to determine that input in the particular alternative vocal input style is to be requested, wherein the audio prompt is generated in accordance with the rule.

19. The one or more non-transitory computer-accessible storage media as recited in claim 16 , storing further program instructions that when executed on or across one or more processors further cause the one or more processors to:

obtain an indication, via a programmatic interface from a developer of the dialog-driven application, of at least one alternative vocal input style of the set of alternative vocal input styles.

20. The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the vocal input obtained subsequent to presentation of the audio prompt comprises (a) a first portion of input expressed in the particular alternative vocal input style and (b) a second portion of input expressed in a style other than the particular alternative vocal input style.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 23, 2021
From: BAKER, JOHN; MISHRA, ANUBHAV; LIU, BANGRUI; HITTNER, CHRISTOPHER MICHAEL; BODAPATI, SRAVAN BABU; PIMPALKHUTE, HARSHAL; KIRCHHOFF, KATRIN; SURANA, ANUJ GAUTAM; SU, YILAI; MENDEZ, BRANDON LOUIS; ZHANG, CHENGSHUN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 058197/0914 →
References Cited (47)
US 6246981B1 · Papineni et al. · 2001 [cited by applicant]
US 6510411B1 · Norton et al. · 2003 [cited by applicant]
US 7197460B1 · Gupta et al. · 2007 [cited by applicant]
US 8626507B2 · Bangalore et al. · 2014 [cited by applicant]
US 9070367B1 · Hoffmeister et al. · 2015 [cited by applicant]
US 10303773B2 · Curtis et al. · 2019 [cited by applicant]
US 10331791B2 · Anbazhagan et al. · 2019 [cited by applicant]
US 10832008B2 · Banerjee et al. · 2020 [cited by applicant]
US 10848443B2 · Helmy · 2020 [cited by applicant]
US 10891152B2 · Anbazhagan et al. · 2021 [cited by applicant]
US 11568135B1 · Mansour · 2023 [cited by examiner]
US 20070143099A1 · Balchandran et al. · 2007 [cited by applicant]
US 20100298012A1 · Damarla · 2010 [cited by applicant]
US 20150263941A1 · Jung · 2015 [cited by applicant]
US 20160042748A1 · Jain et al. · 2016 [cited by applicant]
US 20160225370A1 · Kannan et al. · 2016 [cited by applicant]
US 20170116982A1 · Gelfenbeyn · 2017 [cited by examiner]
US 20170125008A1 · Maisonnier et al. · 2017 [cited by applicant]
US 20170286916A1 · Skiba · 2017 [cited by examiner]
US 20210132986A1 · Anbazhagan · 2021 [cited by examiner]
EP 2933070 · 2018 [cited by applicant]
JP 2003044988 · 2003 [cited by examiner]
JP 2003044988A · 2003 [cited by examiner]
KR 20090090613 · 2009 [cited by examiner]
KR 20090090613A · 2009 [cited by examiner]
WO 9723088 · 1997 [cited by applicant]
Rouillard, “Web services and speech-based applications,” 2006 ACS/IEEE International Conference on Pervasive Services, Lyon, France, 2006, pp. 341-344, doi: 10.1109/PERSER.2006. 1652258. keywords: {Web services;Speech s… [cited by examiner]
Svetlana Stoyanchev et al “Rapid Prototyping of Form-driven Dialogue Systems Using an Open-Source Framework”, Proceddings of the Sigdial 2016 Conference, pp. 216-219. [cited by applicant]
Claus Brabrand “PowerForms: Declarative client-side form field validation” Brics Report Series, Jan. 1, 2000, pp. 205-214. [cited by applicant]
Robert Jamison, “Announcing a New Tool for Building Interactive Adventure Games on Alexa”, Amazon Mobile App Distribution Blog, Retrieved from URL: https://developer.amazon.com/public/community/post/TxEQV5K754YS77/Annou… [cited by applicant]
“Getting Started with the Alexa Skills Kit”, Amazon Apps & Games Developer Portal, Retrieved from URL: https://developer.amazon.com/pulbic/solutions/slexas/alexa-skills-kit/getting-started-guide on Oct. 30, 2016, pp. 1-… [cited by applicant]
Seth Rosenberg, “How to Build Bots for Messenger”, Facebook for Developers, Retrieved from URL: https://developers.facebook.com/blog/post/2016/04/12/bots-for-messenger on Oct. 30, 2016, pp. 1-5. [cited by applicant]
Ali El-Kahky, et al., “Entending Domain Coverage of Language Understanding Systems Via Intent Transfer Between Domains Using Knowledge Graphs and Search Query Click Logs”, 2014 IEEE International Conference on Acoustic,… [cited by applicant]
Elad Natanson, “Messaging Platforms, Bots and The Future of Mobile”, Retrieved from URL: http://www.forbes.com/sites/eladnatanson/2016/04/08/messaging-platforms-bot-and-the-future-of-mobile/#2d1ab79884af on Oct. 30, 201… [cited by applicant]
“Messenger Platform”, Facebook for Developers, Retrieved from URL: https://developers.facebook.com/doc/messenger-platform on Oct. 30, 2016, pp. 1-3. [cited by applicant]
Collen Estrada, “Microsoft Bot Framework”, Mar. 30, 2016, Retrieved from URL: https://blog.botframework.com/2016/03/30/BotFramework/ on Oct. 30, 2016, pp. 1-7. [cited by applicant]
“Microsoft Cognitive Services—APIs”, Retrieved from URL: https://www.microsoft.com/cognitive-services/en-us/apis on Oct. 30, 2016, pp. 1-8. [cited by applicant]
Himanshu S. Bhatt, et al., “Cross-domain Text Classification with Multiple Domains and Disparate Label Sets”, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Aug. 7-12, 2016, pp.… [cited by applicant]
Amit Fulay, “Say hello to Google Allo: a smarter messaging app”, Retrieved from URL: https://blog.google/products/allo/google-allo-smater-messaging-app on Oct. 30, 2016, pp. 1-14. [cited by applicant]
“Training for the Alexa Skills Kit”, Amazon Apps & Games Developer Portal, Retrieved from URL: https://developer.amazon.com/public/solutions/alexa/alexa-skills-kits/content/alexa-skilss-developer-training on Oct. 30, 20… [cited by applicant]
Wikipedia, “Vorbis”, Retrieved from URL: https://en.wikipedia.org/wiki/Vorbis on Sep. 26, 2016, pp. 1-10. [cited by applicant]
U.S. Appl. No. 17/030,204, filed Sep. 23, 2020, Saab Mansour. [cited by applicant]
U.S. Appl. No. 17/039,900, filed Sep. 30, 2020, Swapandeep Singh et al. [cited by applicant]
U.S. Appl. No. 17/039,920, filed Sep. 30, 2020, Jin Hoon Bang et al. [cited by applicant]
U.S. Appl. No. 17/219,640, filed Mar. 31, 2021, Pokkunuri, et al., Amazon Technologies, Inc. [cited by applicant]
U.S. Appl. No. 17/219,630, filed Mar. 31, 2021, Pushkin, et al., Amazon Technologies, Inc. [cited by applicant]
U.S. Appl. No. 17/039,889, filed Sep. 30, 2020, Swapandeep Singh. [cited by applicant]