IP Library › Granted Patent US 12,272,353
Granted Patent B2
US 12,272,353 · App. 17/952,503 · Granted Apr 8, 2025

System for enhancing speech understanding with effective and efficient integration of automated speech recognition error correction, out-of-domain detection, and/or domain classification

Inventor: Zhengyu Zhou (Fremont, CA)
Assignee: Robert Bosch GmbH
G10L15/197G06F40/117G06F40/20G06F40/35G10L15/16G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,353
App. No.
17/952,503
Granted
Apr 8, 2025
Kind
B2
Abstract

The systems and methods described herein are related to a new speech understanding system for domain-specific voice interaction. The systems and methods described herein combine the automatic correction of an automatic speech recognition errors with a natural language understanding model in a way that optimizes the recognition and understanding of a received speech input. The systems and methods described herein may further support out-of-domain detection or domain classification by jointly learning and/or performing automatic speech recognition error correction and domain-related classification. Through the joint learning with automatic speech recognition error correction, the out-of-domain detection or domain classification may be conducted based on a plurality of possible speech recognition results with shared feature inputs and shared neural layers. The systems and methods described herein may achieve robust performance with high computational efficiency.

Claims (58)

1. A method for identifying a text string associated with a speech input, the method comprising:

receiving a speech input from a speech-based dialogue interface;

generating, by at least one speech recognition model, first and second text transcription predictions by converting the speech input into respective text strings;

generating a first confidence score for the first text transcription prediction and a second confidence score for the second text transcription prediction;

determining a slot type for respective text strings of the first text transcription prediction and the second text transcription prediction;

applying tags to the respective text strings of the first text transcription prediction and the second text transcription prediction based on the slot type for each respective text string, wherein each tag corresponds to one of the slot types or to an empty tag type that does not include a description;

generating at least one tagged text group by extracting at least one text string of the first text transcription prediction and grouping the at least one text string with at least one text string of the second text transcription prediction based on a tag associated with the text strings;

in response to a determination that the tag of the at least one text string of the first text transcription prediction and the tag of the at least one text string of the second text transcription prediction corresponds to one of the slot types:

determining, using a natural language model, a relevance ranking of the text strings within the at least one tagged text group based on, at least, the slot type of the tag associated with the tagged text group, the first confidence score, and the second confidence score; and

identifying a text string of the at least one tagged text group having a highest relevance ranking, wherein the text string of the at least one tagged text group having the highest relevance ranking is provided to the speech-based dialogue interface; and

in response to a determination that the tag of the at least one text string of the first text transcription prediction and the tag of the at least one text string of the second text transcription prediction corresponds to the empty tag type, discard the at least one tagged text group.

2. The method of claim 1 , wherein the natural language model is one of a class-based natural language understanding model and a word-based natural language understanding model.

3. The method of claim 2 , further comprising:

selecting the class-based natural language understanding model in response to a tag value of the tagged text group indicating an essential tag type; and

selecting the word-based natural language understanding model in response to a tag value of the tagged text group not indicating an essential tag type.

4. The method of claim 1 , further comprising:

applying multiple tags for a respective text string; and

associating a copy of the respective text string to the tagged text group associated with a respective tag.

5. The method of claim 1 , wherein the natural language model includes at least one layer configured to identify text that does not correspond to a tag type.

6. The method of claim 1 , wherein the text strings associated with a tagged text group include a similar linguistic usage.

7. A system for outputting, at a speech-based dialogue interface, a text string associated with a speech input, the system comprising:

a processor; and

a memory including instructions that, when executed by the processor, cause the processor to:

receive a speech input from a speech-based dialogue interface;

generate, by at least one speech recognition model, first and second text transcription predictions by converting the speech input into respective text strings;

generate a first confidence score for the first text transcription prediction and a second confidence score for the second text transcription prediction;

determine a slot type for respective text strings of the first text transcription prediction and the second text transcription prediction;

apply tags to the respective text strings of the first text transcription prediction and the second text transcription prediction based on the slot type for each respective text string, wherein each tag corresponds to one of the slot types or to an empty tag type that does not include a description;

generate at least one tagged text group by extracting at least one text string of the first text transcription prediction and grouping the at least one text string with at least one text string of the second text transcription prediction based on a tag associated with the text strings;

in response to a determination that the tag of the at least one text string of the first text transcription prediction and the tag of the at least one text string of the second text transcription prediction corresponds to one of the slot types:

determine, using a natural language model, a relevance ranking of the text strings within the at least one tagged text group based on, at least, the slot type of the tag associated with the tagged text group, the first confidence score, and the second confidence score;

identify a text string of the at least one tagged text group having a highest relevance ranking; and

output the text string of the at least one tagged text group having the highest relevance ranking to the speech-based dialogue interface; and

in response to a determination that the tag of the at least one text string of the first text transcription prediction and the tag of the at least one text string of the second text transcription prediction corresponds to the empty tag type, discard the at least one tagged text group.

8. The system of claim 7 , wherein the natural language model is one of a class-based natural language understanding model and a word-based natural language understanding model.

9. The system of claim 8 , wherein the instructions further cause the processor to:

select the class-based natural language understanding model in response to a tag value of the tagged text group indicating an essential tag type; and

select the word-based natural language understanding model in response to a tag value of the tagged text group not indicating an essential tag type.

10. The system of claim 7 , wherein at least one tag corresponds to an empty tag type that does not include a description.

11. The system of claim 10 , wherein a tagged text group related to the empty tag type is discarded.

12. The system of claim 7 , wherein the instructions further cause the processor to:

apply multiple tags for a respective text string; and

associate a copy of the respective text string to the tagged text group associated with a respective tag.

13. An apparatus for identifying a text string associated with a speech input, the apparatus comprising:

a processor; and

a memory including instructions that, when executed by the processor, cause the processor to:

receive a speech input from a speech-based dialogue interface captured by at least one input device;

generate, by at least one speech recognition model, first and second text transcription predictions by converting the speech input into respective text strings;

generate a first confidence score for the first text transcription prediction and a second confidence score for the second text transcription prediction;

determine a slot type for respective text strings of the first text transcription prediction and the second text transcription prediction;

apply tags to the respective text strings of the first text transcription prediction and the second text transcription prediction based on the slot type for each respective text string, wherein each tag corresponds to one of the slot types or to an empty tag type that does not include a description;

generate at least one tagged text group by extracting at least one text string of the first text transcription prediction and grouping the at least one text string with at least one text string of the second text transcription prediction based on a tag associated with the text strings;

in response to a determination that the tag of the at least one text string of the first text transcription prediction and the tag of the at least one text string of the second text transcription prediction corresponds to one of the slot types:

determine, using a natural language model, a relevance ranking of the text strings within the at least one tagged text group based on, at least, the slot type of the tag associated with the tagged text group, the first confidence score, and the second confidence score; and

identify a text string of the at least one tagged text group having a highest relevance ranking; and

in response to a determination that the tag of the at least one text string of the first text transcription prediction and the tag of the at least one text string of the second text transcription prediction corresponds to the empty tag type, discard the at least one tagged text group.

14. The apparatus of claim 13 , wherein the natural language model is one of a class-based natural language understanding model and a word-based natural language understanding model.

15. The apparatus of claim 13 , wherein the at least one input device includes at least one microphone, and wherein the at least one input device is associated with at least one of a manufacturing machine, a power tool, an automated personal assistant, a domestic appliance, surveillance system, and a medical imaging system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2022
From: ZHOU, ZHENGYU
To: ROBERT BOSCH GMBH
Reel/Frame 061210/0528 →
Continuity (1)
Related Publication 20240105170A1 · Mar 28, 2024
References Cited (12)
US 9959861B2 · Zhou et al. · 2018 [cited by applicant]
US 10453117B1 · Reavely · 2019 [cited by examiner]
US 11043205B1 · Su · 2021 [cited by examiner]
US 11195522B1 · Makashir · 2021 [cited by examiner]
US 11308281B1 · Craft · 2022 [cited by examiner]
US 20200409943A1 · Diefenbach · 2020 [cited by examiner]
US 20210074279A1 · Rastogi · 2021 [cited by examiner]
US 20220199079A1 · Hanson · 2022 [cited by examiner]
Chen et al., “BERT for Joint Intent Classification and Slot Filling,” arXiv:1902.10909v1, 2019, 6 pages. [cited by applicant]
Liu et al., “Attention-Based Recurrent Neural Network Models for Joint Intent Detection and Slot Filling,” arXiv:1609.01454v1, 2016, 5 pages. [cited by applicant]
Wang et al., “A Bi-model based RNN Semantic Frame Parsing Model for Intent Detection and Slot Filling,” arXiv:1812.10235v1, 2018, 6 pages. [cited by applicant]
Zhou et al., “A Neural Network Based Ranking Framework to Improve ASR With NLU Related Knowledge Deployed,” IEEE ICASSP, 2019, 5 pages. [cited by applicant]
Cited By (1)
US 12,554,940