IP Library Granted Patent US 11,250,853
Granted Patent B2
US 11,250,853 · App. 16/862,620 · Granted Feb 15, 2022

Sarcasm-sensitive spoken dialog system

Inventors: Zhengyu Zhou (Fremont, CA); In Gyu Choi (Atlanta, GA)
Assignee: ROBERT BOSCH GMBH
G10L15/22G10L15/02G10L15/05G10L15/063G10L15/083G10L15/16G10L2015/227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,250,853
App. No.
16/862,620
Granted
Feb 15, 2022
Kind
B2
Abstract

A dialog system and a method of using the dialog system is disclosed. The method may comprise: receiving audible human speech from a user; determining that the audible human speech comprises sarcasm information; providing an input to a neural network, wherein the input comprises speech data input associated with the audible human speech, an embedding vector associated with the sarcasm information, and a one-hot vector; and based on the input, determining an audible response to the human speech.

Claims (36)

1. A method, comprising:

receiving audible human speech from a user;

determining that the audible human speech comprises sarcasm information;

providing an input to a neural network, wherein the input comprises speech data input associated with the audible human speech, an embedding vector associated with the sarcasm information, and a one-hot vector; and

based on the input, determining an audible response to the human speech.

2. The method of claim 1 , wherein the embedding vector associated with the sarcasm information is appended to the speech data input, wherein the speech data input comprises a word embedding layer which comprises a plurality of word embedding vectors.

3. The method of claim 1 , wherein the one-hot vector comprises a first dimension, wherein the audible response is determined based on a value of the first dimension.

4. The method of claim 3 , wherein the one-hot vector comprises a second dimension, wherein the audible response comprises sarcasm information determined based on a value of the second dimension.

5. The method of claim 3 , wherein the one-hot vector comprises a third dimension, wherein the audible response is determined based on a value of the third dimension, wherein the third dimension is associated with a previous utterance of the user.

6. The method of claim 1 , wherein the input further comprises a dialog history that comprises at least one additional speech data input.

7. The method of claim 1 , wherein the speech input data comprises a word sequence determined using a speech recognition model.

8. The method of claim 1 , wherein a signal knowledge extraction model provides the embedding vector and the one-hot vector based on a determination of sarcasm information in the audible human speech and based on a word sequence received from a speech recognition model.

9. The method of claim 1 , wherein determining that the audible human speech comprises sarcasm information, comprises:

determining textual speech data and signal speech data from the audible human speech;

determining that a text-based sentiment is Positive or Neutral by processing the textual speech data using a text-based sentiment analysis tool;

determining that a signal-based sentiment is Negative by processing the signal speech data using a signal-based sentiment analysis tool; and

detecting sarcasm based on the text-based sentiment being Positive or Neutral while the signal-based sentiment is Negative.

10. The method of claim 1 , wherein a sarcasm-sensitive spoken dialog system comprises the neural network, wherein the sarcasm-sensitive spoken dialog system is embodied in one of: a table-top device, a kiosk, a mobile device, a vehicle, or a robotic machine.

11. A non-transitory computer-readable medium comprising a plurality of computer-executable instructions and memory for maintaining the plurality of computer-executable instructions, the plurality of computer-executable instructions, when executed by one or more processors of a computer, perform the following functions:

receive audible human speech from a user;

determine that the audible human speech comprises sarcasm information;

provide an input to a neural network, wherein the input comprises speech data input associated with the audible human speech, an embedding vector associated with the sarcasm information, and a one-hot vector; and

based on the input, determine an audible response to the human speech.

12. The non-transitory computer-readable medium of claim 11 , wherein the embedding vector associated with the sarcasm information is appended to the speech data input, wherein the speech data input comprises a word embedding layer which comprises a plurality of word embedding vectors.

13. The non-transitory computer-readable medium of claim 11 , wherein the one-hot vector comprises a first dimension, wherein the audible response is determined based on a value of the first dimension.

14. The non-transitory computer-readable medium of claim 11 , wherein a signal knowledge extraction model provides the embedding vector and the one-hot vector based on a determination of sarcasm information in the audible human speech and based on a word sequence received from a speech recognition model.

15. A sarcasm-sensitive spoken dialog system, comprising: one or more processors; and memory coupled to the one or more processors, wherein the memory stores a plurality of instructions executable by the one or more processors, the plurality of instructions comprising, to:

receive audible human speech from a user;

determine that the audible human speech comprises sarcasm information;

provide an input to a neural network, wherein the input comprises speech data input associated with the audible human speech, an embedding vector associated with the sarcasm information, and a one-hot vector; and

based on the input, determine a response to the human speech.

16. The system of claim 15 , wherein the embedding vector associated with the sarcasm information is appended to the speech data input, wherein the speech data input comprises a word embedding layer which comprises a plurality of word embedding vectors.

17. The system of claim 15 , wherein the one-hot vector comprises a first dimension, wherein the audible response is determined based on a value of the first dimension.

18. The system of claim 17 , wherein the one-hot vector comprises a second dimension, wherein the audible response comprises sarcasm information determined based on a value of the second dimension.

19. The system of claim 15 , wherein a signal knowledge extraction model provides the embedding vector and the one-hot vector based on a determination of sarcasm information in the audible human speech and based on a word sequence received from a speech recognition model.

20. A table-top device, a kiosk, a mobile device, a vehicle, or a robotic machine comprising the sarcasm-sensitive spoken dialog system of claim 15 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2020
From: ZHOU, ZHENGYU; CHOI, IN GYU
To: ROBERT BOSCH GMBH
Reel/Frame 052533/0641 →
Continuity (1)
Related Publication 20210343280A1 · Nov 4, 2021