IP Library Granted Patent US 11,270,084
Granted Patent B2
US 11,270,084 · App. 16/159,199 · Granted Mar 8, 2022

Systems and methods for using trigger words to generate human-like responses in virtual assistants

Inventors: Viswanath Ramamurti (San Leandro, CA); Young M. Lee (Old Westbury, NY)
Assignee: Johnson Controls Tyco IP Holdings LLP
G06F40/56G10L15/16G10L15/1822G10L15/22G06F3/167G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,270,084
App. No.
16/159,199
Granted
Mar 8, 2022
Kind
B2
Abstract

A method for generating a human-like response to a voice or text command includes receiving an input sequence of words and processing the input sequence of words to generate a trigger word that is indicative of a desired nature of the human-like response. The method further includes encoding a neural network using the trigger word and generating the human-like response using an output of the neural network. The method enables implementation of voice command functionality in various types of devices with only a small amount of training data.

Claims (45)

1. A method for generating a human-like response to a voice or text input, the method comprising:

receiving an input sequence of words;

processing, by an extraction module, the input sequence of words to generate a trigger word that is indicative of a desired nature of the human-like response, the extraction module comprising a computer-readable storage medium having instructions stored thereon that are executable by a processor;

encoding, by the extraction module, a state of a neural network using the trigger word by inputting information into the state before an encoder is trained to output the state;

generating the human-like response using an output word sequence of the neural network, the output word sequence generated by the neural network by decoding the output word sequence from the state of the neural network encoded using the trigger word; and

switching from encoding the state with the extraction module to encoding the state of the neural network with the encoder once the encoder is trained to a particular level.

2. The method of claim 1 , further comprising training the neural network with a set of pairs, each pair of the set of pairs comprising one trigger word and one human-like response.

3. The method of claim 2 , further comprising building a log of input sequences and associated human-like responses after deployment of the neural network.

4. The method of claim 3 , further comprising updating the set of pairs using the log and re-training the neural network using the set of pairs.

5. The method of claim 1 , wherein encoding the neural network using the trigger word comprises setting the state of the neural network to a word vector representation of the trigger word, wherein the state is a hidden state.

6. The method of claim 1 , wherein processing the input sequence of words to generate the trigger word comprises performing a part-of-speech tagging process associated with the input sequence of words.

7. The method of claim 1 , wherein processing the input sequence of words to generate the trigger word comprises comparing one or more words of the input sequence of words to previously generated trigger words.

8. The method of claim 1 , wherein the output word sequence of the neural network is a probability of a next word given at least one previously predicted word of the human-like response.

9. The method of claim 8 , wherein generating the human-like response using the output word sequence of the neural network comprises using a beam search.

10. The method of claim 1 , further comprising transmitting the human-like response to a device configured to provide the human-like response to a human.

11. The method claim 1 , wherein the method further comprises:

switching, after an amount of time, to encoding the state of the neural network with the encoder.

12. A system comprising:

a cloud computing platform; and

a device with an integrated virtual assistant, the device in communication with the cloud computing platform, the device comprising a computer-readable storage medium having instructions stored thereon that are executable by a processor of the device and configured to implement an extraction module, the device configured to:

receive an input sequence of words spoken by a human;

transmit the input sequence of words to the cloud computing platform, the cloud computing platform configured to:

extract, by the extraction module, a trigger word from the input sequence of words that is indicative of a desired nature of an output response to be transmitted to the human;

encode, by the extraction module, a state of a neural network using the trigger word by inputting information into the state before an encoder is trained to output the state;

generate the output response using an output word sequence of the neural network, the output word sequence generated by the neural network by decoding the output word sequence from the state of the neural network encoded using the trigger word; and

switch from encoding the state with the extraction module to encoding the state of the neural network with the encoder once the encoder is trained to a particular level;

receive the output response from the cloud computing platform; and

provide the output response to the human.

13. The system of claim 12 , wherein the device is a thermostat, a speaker, a smartphone, a tablet, a watch, a personal computer, or a laptop.

14. The system of claim 12 , wherein the cloud computing platform is further configured to perform a part-of-speech tagging process associated with the input sequence of words.

15. The system of claim 12 , wherein the neural network is trained with a set of pairs, each pair of the set of pairs comprising one trigger word and one output response.

16. The system of claim 15 , wherein the cloud computing platform is configured to build a log of input sequences and associated output responses after deployment of the neural network, update the set of pairs using the log, and re-train the neural network using the set of pairs.

17. A method for providing a response to a speech input, the method comprising:

receiving, by a device with an integrated virtual assistant, the speech input from a human;

transmitting, by the device, the speech input to a cloud computing platform;

extracting, by an extraction module of the cloud computing platform, a trigger word from the speech input that is indicative of a desired nature of the response, the extraction module comprising a computer-readable storage medium having instructions stored thereon that are executable by a processor;

encoding, by the extraction module of the cloud computing platform, a state of a neural network using the trigger word by inputting information into the state before an encoder is trained to output the state;

generating, by the cloud computing platform, a word sequence response using the neural network, the word sequence response generated by the neural network by decoding the word sequence response from the state of the neural network encoded using the trigger word;

receiving, by the device, the response from the cloud computing platform;

providing, by the device, the response to the human; and

switching, by the cloud computing platform, from encoding the state with the extraction module to encoding the state of the neural network with the encoder once the encoder is trained to a particular level.

18. The method of claim 17 , further comprising training the neural network with a set of pairs, each pair of the set of pairs comprising one trigger word and one response.

19. The method of claim 17 , wherein extracting the trigger word from the speech input comprises performing a part-of-speech tagging process associated with the speech input.

20. The method of claim 17 , wherein extracting the trigger word from the speech input comprises comparing one or more words of the speech input to previously generated trigger words.

21. The method of claim 17 , wherein providing the trigger word as an input to the neural network comprises setting the state of the neural network to a word vector representation of the trigger word, wherein the state of the neural network is a hidden state.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2024
From: JOHNSON CONTROLS TYCO IP HOLDINGS LLP
To: TYCO FIRE & SECURITY GMBH
Reel/Frame 067056/0552 →
NUNC PRO TUNC ASSIGNMENT Recorded Feb 4, 2022
From: JOHNSON CONTROLS TECHNOLOGY COMPANY
To: JOHNSON CONTROLS TYCO IP HOLDINGS LLP
Reel/Frame 058959/0764 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2019
From: RAMAMURTI, VISWANATH; LEE, YOUNG M.
To: JOHNSON CONTROLS TECHNOLOGY COMPANY
Reel/Frame 049390/0440 →