IP Library Granted Patent US 12,314,991
Granted Patent B2
US 12,314,991 · App. 17/408,476 · Granted May 27, 2025

Method and system for controlling a graphical user interface by telephone

Inventors: Kamyar Mohajer (Mississauga, CA); Keyvan Mohajer (Los Gatos, CA); James Hom (Palo Alto, CA); Evelyn Jiang (Cupertino, CA)
G06Q30/0619G06Q30/0623G06Q30/0631G06Q30/0633G06Q30/0643
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,991
App. No.
17/408,476
Granted
May 27, 2025
Kind
B2
Abstract

A method and system for controlling a GUI on a user's network-connected device, the control being provided by a telephone call between the user and a speech recognition and speech synthesis system. An example of a restaurant ordering system is provided. The user calls a phone number and is guided through a verbal ordering process that includes one or more of: adding an item, deleting an item, changing quantities, changing sizes, and changing details of an item. The user's choices are added to a display so that a current status of the order is visible to the user. The GUI is updated as changes are made to the order. The GUI can also request additional information, upsell items, and show menus. The GUI aids the user in confirming that the order is correct. The system provides the final order to a restaurant for fulfillment.

Claims (52)

1. A computer-implemented method for an interactive voice and GUI system, comprising:

receiving, at a voice ordering server, a voice telephone call from a mobile device associated with a telephone number;

transmitting the voice telephone call to a middleware reflecting voice from the telephone call;

performing speech recognition on audio information of the voice telephone call;

determining response data in an optimal format based on the audio information of the voice telephone call, wherein the response data comprises interactive audio data and visual data, and wherein the audio data is delivered via synthesized speech on the voice telephone call and the visual data is displayed on the mobile device via a URL;

sending, after confirming via the audio data by the voice ordering server, a text message to the mobile device associated with the telephone number, the text message containing the URL to deliver the visual data displayed on the mobile device;

receiving a request from the mobile device to the URL;

serving code that when rendered by the mobile device, provides a graphical display of a first choice associated with the URL, wherein the first choice is recognized through speech recognition by the voice ordering server;

performing speech recognition on audio received via the mobile device to interpret the interaction of the voice telephone call;

updating the graphical display based on the interaction of the voice telephone call;

serving code that, when rendered by the mobile device, provides an updated graphical display acknowledging the updated first choice; and

terminating the URL upon completion of the voice telephone call.

2. The method of claim 1 , further comprising, during the telephone call:

recognizing a second choice from the interaction of the voice telephone call; and

serving code that changes the graphical display in response to the second choice.

3. The method of claim 2 , wherein the second choice changes a quantity of the first choice.

4. The method of claim 2 , wherein the second choice adds an item to the first choice.

5. The method of claim 2 , wherein the second choice removes an item from the first choice.

6. The method of claim 2 , wherein the second choice changes an item size of the first choice.

7. The method of claim 2 , further comprising:

offering, via synthesized speech during the telephone call, an upsell item; and

sending, during the telephone call, visual information about the upsell item to the graphical user interface;

wherein the user's second choice is the offered upsell item and the graphical display is updated to show the second choice.

8. The method of claim 7 wherein the synthesized voice is sent through the internet for output.

9. The method of claim 7 wherein the synthesized voice is sent via the mobile device.

10. The method of claim 2 , further comprising:

requesting additional information, via synthesized speech during the telephone call, about the user's first choice; and

sending, during the telephone call, visual information about the requested additional information to the graphical display;

wherein the user's second choice is a user's response to the request for additional information and the graphical user interface is updated to show the second choice.

11. The method of claim 1 , further comprising:

sending the user's first choice to a third-party ordering system.

12. The method of claim 1 , further comprising:

sending the user's second choice to a third-party ordering system.

13. The method of claim 1 further comprising serving images in conjunction with the code that provides a graphical display.

14. An interactive voice and GUI ordering system, comprising:

a memory storing software that executed by a processor, causes the system to:

receive, at a voice ordering server, a phone call from a mobile device associated with a telephone number;

transmitting the voice telephone call to a middleware reflecting voice from the telephone call;

perform speech recognition on audio information of the phone;

determining response data in an optimal format based on the audio information of the voice telephone call, wherein the response data comprises interactive audio data and visual data, wherein the audio data is delivered via synthesized speech on the voice telephone call and the visual data is displayed on the mobile device via a URL;

send, after confirming via the audio data by the voice ordering server, a text message to the mobile device associated with the telephone number, the text message containing the URL to deliver the visual data displayed on the mobile device;

receive a request from the mobile device to the URL;

send a dynamic GUI web page reflecting current order information of a voice order from the phone call, wherein the dynamic GUI web page is associated with the URL, and the current order information of the voice order is recognized through speech recognition by the voice ordering server;

receive updated audio information, from the mobile device for the voice order;

perform speech recognition on the received updated audio information to yield recognized speech and to interpret the interaction of the phone call;

update, in accordance with the recognized speech, the dynamic GUI web page reflecting the updated order information; and

terminate the URL upon completion of the voice telephone call.

15. The system of claim 14 , further comprising:

an interface that provides the order information reflecting a state of the voice order to a third party ordering system, wherein the state of the voice order includes items corresponding to the phone call.

16. The voice ordering system of claim 14 , further comprising a memory storing upselling criteria, wherein the software further causes the system to:

determine an upsell by applying the upsell criteria to the state of the voice order; and

send an update to the dynamic GUI web page, the update comprising the upsell.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2021
From: MOHAJER, KAMYAR; MOHAJER, KEYVAN; HOM, JAMES; JIANG, EVELYN
To: SOUNDHOUND, INC.
Reel/Frame 057419/0226 →
Continuity (1)
Related Publication 20230059765A1 · Feb 23, 2023
References Cited (16)
US 8867708B1 · Lavian · 2014 [cited by examiner]
US 9747630B2 · Mierle · 2017 [cited by examiner]
US 20060034270A1 · Haase · 2006 [cited by examiner]
US 20090030708A1 · Glasglow · 2009 [cited by examiner]
US 20100205053A1 · Shuster · 2010 [cited by examiner]
US 20120185240A1 · Goller · 2012 [cited by examiner]
US 20150237189A1 · Schultz · 2015 [cited by examiner]
US 20180053192A1 · Aroori · 2018 [cited by examiner]
US 20180084110A1 · Fraizer · 2018 [cited by examiner]
US 20180293983A1 · Choi · 2018 [cited by examiner]
US 20190325481A1 · Friborg, Jr. · 2019 [cited by examiner]
US 20210090154A1 · Michaelson · 2021 [cited by examiner]
US 20210272183A1 · Turnbull · 2021 [cited by examiner]
Janga, S. “Expediting and Enhancing User Interaction with Interactive Voice Response Systems Utilizing Machine Learning”, Technical Disclosure Commons. (Year: 2019). [cited by examiner]
Min-Jen, T. “VoiceXML dialog system of the multimodal IP—Telephony—The application for voice ordering service”, Expert Systems with Applications, vol. 31, Issue 4, pp. 684-696, https://doi.org/10.1016/j.eswa.2006.01.010… [cited by examiner]
Tsai, M. J. “The VoiceXML dialog system for the e-commerce ordering service,” Proceedings of the Ninth International Conference on Computer Supported Cooperative Work in Design, Coventry, UK, pp. 95-100 vol. 1, doi: 10.… [cited by examiner]