IP Library Granted Patent US 9,514,130
Granted Patent B2
US 9,514,130 · App. 14/985,300 · Granted Dec 6, 2016

Device for extracting information from a dialog

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,514,130
App. No.
14/985,300
Granted
Dec 6, 2016
Kind
B2
Abstract

Computer-implemented systems and methods for extracting information during a human-to-human mono-lingual or multi-lingual dialog between two speakers are disclosed. Information from either the recognized speech (or the translation thereof) by the second speaker and/or the recognized speech by the first speaker (or the translation thereof) is extracted. The extracted information is then entered into an electronic form stored in a data store.

Claims (66)

1. A method comprising:

receiving a first speech input from a first speaker;

determining, by a speech translation system, a first recognized speech result based on the speech input;

determining, by the speech translation system, whether there exists a recognition ambiguity in the first recognized speech result, wherein the recognition ambiguity indicates more than one possible match for the first recognized speech result;

upon a determination that there is recognition ambiguity in the first recognized speech result of the first speaker, determining a confidence score based on the recognition ambiguity; and

responsive to the confidence score being below a threshold, issuing a first disambiguation query to the first speaker via the speech translation system, wherein a response to the first disambiguation query resolves the recognition ambiguity.

2. The method of claim 1 , further comprising:

receiving a second speech input from a second speaker;

determining, by the speech translation system, a second recognized speech result based on the second speech input;

extracting information from the second recognized speech input from the second speaker;

entering the extracted information into an electronic form; and

displaying the electronic form.

3. The method of claim 2 , wherein the determination of whether there exists an ambiguity in the recognized speech result of the first speaker is based one or more of:

an acoustic confidence score in the recognized speech result of the first speaker;

a context of the electronic form; and

a language context given by a translation of one or more utterances from the second speaker from a second language to the first language.

4. The method of claim 2 , further comprising:

determining, by the speech translation system, whether there exists a recognition ambiguity in the second recognized speech result of the second speaker, wherein the recognition ambiguity indicates more than one possible match for the second recognized speech result;

upon a determination that there is recognition ambiguity in the second recognized speech result of the first speaker, determining a confidence score based on the recognition ambiguity; and

responsive to the confidence score being below a threshold, issuing a second disambiguation query to the second speaker via the speech translation system, wherein a response to the second disambiguation query resolves the recognition ambiguity.

5. The method of claim 2 , wherein the second speaker speaks a second language, further comprising:

translating the first recognized speech result into the second language; and

translating the second recognized speech result into the first language,

wherein the information is extracted from the translation of the second recognized speech result.

6. The method of claim 5 , wherein extracting information from the second recognized speech input from the second speaker comprises parsing the translation by a semantic grammar.

7. The method of claim 5 , wherein extracting information from the second recognized speech input from the second speaker comprises detecting one or more keywords in the translation.

8. The method of claim 2 , further comprising:

retrieving one or more documents related to the extracted information from a remote database.

9. The method of claim 2 , further comprising:

soliciting feedback from at least one of the first speaker and the second speaker prior to entering the extracted information into the electronic form.

10. The method of claim 2 , further comprising:

receiving an edit to the extracted information in the electronic form; and

updating the electronic form to reflect the received edit.

11. A non-transitory computer-readable medium comprising instructions that when executed by a processor cause the processor to:

receive a first speech input from a first speaker;

determine, by a speech translation system, a first recognized speech result based on the speech input;

determine, by the speech translation system, whether there exists a recognition ambiguity in the first recognized speech result, wherein the recognition ambiguity indicates more than one possible match for the first recognized speech result;

upon a determination that there is recognition ambiguity in the first recognized speech result of the first speaker, determine a confidence score based on the recognition ambiguity; and

responsive to the confidence score being below a threshold, issue a first disambiguation query to the first speaker via the speech translation system, wherein a response to the first disambiguation query resolves the recognition ambiguity.

12. The non-transitory computer-readable medium of claim 11 , further comprising instructions that cause the processor to:

receive a second speech input from a second speaker;

determine, by the speech translation system, a second recognized speech result based on the second speech input;

extract information from the second recognized speech input from the second speaker;

enter the extracted information into an electronic form; and

display the electronic form.

13. The non-transitory computer-readable medium of claim 12 , wherein the determination of whether there exists an ambiguity in the recognized speech result of the first speaker is based one or more of:

an acoustic confidence score in the recognized speech result of the first speaker;

a context of the electronic form; and

a language context given by a translation of one or more utterances from the second speaker from the second language to the first language.

14. The non-transitory computer-readable medium of claim 12 , further comprising instructions that cause the processor to:

determine, by the speech translation system, whether there exists a recognition ambiguity in the second recognized speech result of the second speaker, wherein the recognition ambiguity indicates more than one possible match for the second recognized speech result;

upon a determination that there is recognition ambiguity in the second recognized speech result of the first speaker, determine a confidence score based on the recognition ambiguity; and

responsive to the confidence score being below a threshold, issue a second disambiguation query to the second speaker via the speech translation system, wherein a response to the second disambiguation query resolves the recognition ambiguity.

15. The non-transitory computer-readable medium of claim 12 , wherein the second speaker speaks the second language, further comprising instructions that cause the processor to:

translate the first recognized speech result into a second language; and

translate the second recognized speech result into the first language,

wherein the information is extracted from the translation of the second recognized speech result.

16. The non-transitory computer-readable medium of claim 15 , wherein the instructions to extract information from the second recognized speech input from the second speaker comprise instructions to parse the translation by a semantic grammar.

17. The non-transitory computer-readable medium of claim 15 , wherein the instructions to extract information from the second recognized speech input from the second speaker comprise instructions to detect one or more keywords in the translation.

18. The non-transitory computer-readable medium of claim 12 , further comprising instructions that cause the processor to:

retrieve one or more documents related to the extracted information from a remote database.

19. The non-transitory computer-readable medium of claim 12 , further comprising instructions that cause the processor to:

solicit feedback from at least one of the first speaker and the second speaker prior to entering the extracted information into the electronic form.

20. The non-transitory computer-readable medium of claim 12 , further comprising instructions that cause the processor to:

receive an edit to the extracted information in the electronic form; and

update the electronic form to reflect the received edit.

Assignments (1)
CHANGE OF NAME Recorded Nov 18, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058897/0824 →