IP Library Granted Patent US 10,565,994
Granted Patent B2
US 10,565,994 · App. 15/827,731 · Granted Feb 18, 2020

Intelligent human-machine conversation framework with speech-to-text and text-to-speech

Inventors: Ching-Ling Huang (San Ramon, CA); Raju Venkataramana (Dublin, CA); Yoshifumi Nishida (San Jose, CA)
Assignee: General Electric Company
G10L15/26G06F16/3329G06F16/60G06N5/04G10L13/033G10L13/0335G10L13/04G10L13/043G10L13/047G10L13/10G10L15/22G10L15/24G10L15/265G10L15/28G10L2015/223G10L2015/225G10L2015/226G10L2015/227G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,565,994
App. No.
15/827,731
Granted
Feb 18, 2020
Kind
B2
Abstract

A method, computer-readable medium, and system including a speech-to-text module to receive an input of speech including one or more words generated by a human and to output data including text, sentiment information, and other parameters corresponding to the speech input; a processing module like Artificial Intelligence to generate a reply to the speech input, the reply including a textual component, sentimental information associated with the textual component, and contextual information associated with the textual component; and a text-to-speech module to receive the textual component, sentimental information, and contextual information and to generate, based on the received textual component and its associated sentimental information and contextual information, a speech output including one or more spoken words, the spoken words to be presented with at least one of a pace, a tone, a volume, and an emphasis representative of the sentimental information and contextual information associated with the textual component.

Claims (24)

1. A system comprising:

a speech-to-text processor to receive an input of speech including one or more words generated by a human and to output data including text, sentiment information, and other parameters information corresponding to the speech input;

a processor to generate a reply to the speech input, the reply including a textual component, sentimental information associated with the textual component, and contextual information associated with the textual component; and

a text-to-speech processor to receive the textual component, sentimental information, and contextual information of the reply and to generate, based on the received textual component and its associated sentimental information and contextual information of the reply, a speech output including one or more spoken words, the spoken words to be presented with at least one of a pace, a tone, a volume, an urgency, a rate, an accent pattern, and an emphasis representative of the sentimental information and contextual information associated with the textual component of the reply, wherein the at least one of the pace, the tone, the volume, the urgency, the rate, the accent pattern, and the emphasis of the speech output is determined on a word by word basis and a sentence by sentence basis for the speech output.

2. The system of claim 1 , wherein the speech-to-text processor is to further receive additional information to aid the speech-to-text processor to accurately output data including the text, the sentiment information, and the other parameters information corresponding to the speech input, the additional information including at least one of an expected human generated response, a keyword, a probability based distribution, a knowledge of prior speeches, a knowledge of an on-going conversation, and combinations thereof.

3. The system of claim 2 , wherein the additional information to aid the speech-to-text processor to accurately output data is received from the processor.

4. The system of claim 1 , wherein the processor receives information used thereby to generate the reply to the speech input from at least one of a database, one or more sensors, one or more controllers, and one or more actuators.

5. The system of claim 1 , wherein the textual component, the sentimental information associated with the textual component, and the contextual information associated with the textual component are synchronized to each other.

6. The system of claim 5 , wherein the processor synchronizes the textual component, the sentimental information, and the contextual information to each other.

7. The system of claim 1 , wherein the processor comprises an Artificial Intelligence processor.

8. The system of claim 1 , wherein the text-to-speech processor can generate speech based on the received textual component in a plurality of different languages.

9. A computer-implemented method, the method comprising:

receiving, by a first processing module, speech input data derived from speech including one or more words generated by a human, the speech input data including text, sentiment information, and other parameters information corresponding to the speech;

generating, by a second processing module, a reply to the speech input data, the reply including a textual component, sentimental information associated with the textual component, and contextual information associated with the textual component; and

transmitting, by a third processing module, the textual component, sentimental information, and contextual information of the reply for the generation of, based on the textual component and its associated sentimental information and contextual information of the reply, a speech output including one or more spoken words, the spoken words to be presented with at least one of a pace, a tone, a volume, an urgency, a rate, an accent pattern, and an emphasis representative of the sentimental information and contextual information associated with the textual component of the reply, wherein the at least one of the pace, the tone, the volume, the urgency, the rate, the accent pattern, and the emphasis of the speech output is determined on a word by word basis and a sentence by sentence basis for the speech output.

10. The method of claim 9 , wherein the second processing module receives information used thereby to generate the reply to the speech input from at least one of a database, one or more sensors, one or more controllers, and one or more actuators.

11. The method of claim 9 , wherein the textual component, the sentimental information associated with the textual component, and the contextual information associated with the textual component are synchronized to each other.

12. The method of claim 11 , wherein the second processing module synchronizes the textual component, the sentimental information, and the contextual information to each other.

13. The method of claim 9 , wherein the second processing module comprises an Artificial Intelligence processor.

14. A non-transitory computer readable medium having processor-executable instructions stored thereon, the medium comprising:

instructions to receive speech input data derived from speech including one or more words generated by a human, the speech input data including text, sentiment information, and other parameters information corresponding to the speech;

instructions to generate a reply to the speech input data, the reply including a textual component, sentimental information associated with the textual component, and contextual information associated with the textual component; and

instructions to transmit the textual component, sentimental information, and contextual information of the reply for the generation of, based on the textual component and its associated sentimental information and contextual information of the reply, a speech output including one or more spoken words, the spoken words to be presented with at least one of a pace, a tone, a volume, an urgency, a rate, an accent pattern, and an emphasis representative of the sentimental information and contextual information associated with the textual component of the reply wherein the at least one of the pace, the tone, the volume, the urgency, the rate, the accent pattern, and the emphasis of the speech output is determined on a word by word basis and a sentence by sentence basis for the speech output.

15. The medium of claim 14 , wherein the textual component, the sentimental information associated with the textual component, and the contextual information associated with the textual component are synchronized to each other.

Assignments (4)
SECURITY INTEREST Recorded Mar 2, 2026
From: INNOVATEPRO MANAGEMENT USA LLC
To: ARES CAPITAL CORPORATION, AS COLLATERAL AGENT
Reel/Frame 073942/0369 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2026
From: GE VERNOVA ELECTRIFICATION SOFTWARE HOLDINGS LLC
To: INNOVATEPRO MANAGEMENT USA LLC
Reel/Frame 073924/0810 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2023
From: GENERAL ELECTRIC COMPANY
To: GE DIGITAL HOLDINGS LLC
Reel/Frame 065612/0085 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2017
From: HUANG, CHING-LING; VENKATARAMANA, RAJU; NISHIDA, YOSHIFUMI
To: GENERAL ELECTRIC COMPANY
Reel/Frame 044265/0395 →
Continuity (1)
Related Publication 20190164554A1 · May 30, 2019