IP Library › Granted Patent US 11,715,469
Granted Patent B2
US 11,715,469 · App. 17/187,163 · Granted Aug 1, 2023

Methods and apparatus for improving search retrieval using inter-utterance context

Inventor: Arpit Sharma (Santa Clara, CA)
Assignee: Walmart Apollo, LLC
G10L15/22G06F40/205G06F40/279G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,715,469
App. No.
17/187,163
Granted
Aug 1, 2023
Kind
B2
Abstract

A system and method of improving the Natural Language Understanding of a voice assistant. A first utterance is converted to text and parsed by a Bi-LSTM neural network to create a vector representing the utterance. A subsequent utterance is similarly converted into a representative vector and the two vector are combined to predict the true intent of a user's subsequent utterance in context with the initial utterance.

Claims (59)

1. An automated assistant comprising a communication system and a computing device operably connected to the communication system, the computing device including at least one memory, and a processor, where the processor is configured to:

receive a first utterance from a user;

convert the first utterance into a first set of text data;

parse the first set of text data to identify a first set of key words;

encode a first intent vector from the first set of key words;

store the first intent vector as a first stored vector;

receive a second utterance from a user;

convert the second utterance into a second set of text data;

parse the second set of text data to identify a second set of key words;

encode a second intent vector from the second set of key words;

combine the second intent vector with the first stored vector forming a combined vector;

forward the combined vector to a dialog manager via a feed forward unit;

store a copy of the combined vector as a second stored vector; and

generate a text response to the user based on an intent of the combined vector.

2. The automated assistant of claim 1 , wherein the processor is further configured to concatenate the first stored vector to the second intent vector to combine the second intent vector and the first stored vector.

3. The automated assistant of claim 1 , wherein the processor is further configured to perform a weighted sum of the second intent vector and the first stored vector to combine the second intent vector and the first stored vector.

4. The automated assistant of claim 3 , wherein the processor is further configured to perform the weighted sum via a gated recurrent unit (GRU).

5. The automated assistant of claim 4 , wherein the processor is further configured to transfer the second intent vector to the GRU via the feed forward unit.

6. The automated assistant of claim 1 , wherein the processor is further configured to form the first intent vector and the second intent vector via a bi-directional long short term memory (Bi-LSTM) network.

7. The automated assistant claim 6 , wherein each keyword is encoded by a separate layer in the Bi-LSTM network.

8. The automated assistant of claim 1 , wherein the processor is further configured to generate the text response via the dialog manager.

9. The automated assistant claim 1 , wherein the processor is further configured to convert the text response to computer-generated speech via a speaker.

10. A computer implemented method of voice recognition and intent classification comprising:

receiving a first utterance from a user;

converting the first utterance into a first set of text data;

parsing the first set of text data to identify a first set of key words;

encoding a first intent vector from the first set of key words;

storing the first intent vector as a first stored vector;

receiving a second utterance from a user;

converting the second utterance into a second set of text data;

parsing the second set of text data to identify a second set of key words;

encoding a second intent vector from the second set of key words;

combining the second intent vector with the first stored vector forming a combined vector;

forwarding the combined vector to a dialog manager via a feed forward unit;

storing a copy of the combined vector as a second stored vector; and

generating a text response to the user based on an intent of the combined vector.

11. The method claim 10 , wherein the step of forming a combined vector further comprises concatenating the first stored vector to the second intent vector.

12. The method of claim 10 wherein the step of forming a combined vector further comprises performing a weighted sum of the second intent vector and the first stored vector.

13. The method of claim 12 , further comprising transferring the second intent vector to a gated recurrent unit (GRU), wherein the step of performing a weighted sum is performed by the GRU.

14. The method of claim 10 , wherein the steps of forming a first intent vector and forming a second intent vector are performed by a bi-directional long short term memory (Bi-LSTM) network.

15. The method of claim 14 , wherein each keyword is encoded by a separate layer in the Bi-LSTM network.

16. The method of claim 10 wherein the step of generating a text response is performed by the dialog manager.

17. The method of claim 10 , further comprising converting the text response to computer-generated speech via a speaker.

18. A non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by at least one processor, cause a device to perform operations comprising:

receiving a first utterance from a user;

converting the first utterance into a first set of text data;

parsing the first set of text data to identify a first set of key words;

encoding a first intent vector from the first set of key words;

storing the first intent vector as a first stored vector;

receiving a second utterance from a user;

converting the second utterance into a second set of text data;

parsing the second set of text data to identify a second set of key words;

encoding a second intent vector from the second set of key words;

combining the second intent vector with the first stored vector forming a combined vector;

forwarding the combined vector to a dialog manager via a feed forward unit;

storing a copy of the combined vector as a second stored vector; and

generating a text response to the user based on am intent of the combined vector.

19. The non-transitory computer readable medium of claim 18 , wherein combining the second intent vector with the first stored vector comprises concatenating the first stored vector to the second intent vector or performing a weighted sum of the second intent vector and the first stored vector.

20. The non-transitory computer readable medium of claim 18 , wherein each keyword is encoded by a separate layer in a bi-directional long short term memory (Bi-LSTM) network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2021
From: SHARMA, ARPIT
To: WALMART APOLLO, LLC
Reel/Frame 055434/0409 →
Continuity (1)
Related Publication 20220277740A1 · Sep 1, 2022
Cited By (1)
US 12,531,060