IP Library Granted Patent US 11,763,936
Granted Patent B2
US 11,763,936 · App. 17/112,279 · Granted Sep 19, 2023

Generating structured text content using speech recognition models

Inventors: Christopher S. Co (Saratoga, CA); Navdeep Jaitly (Mountain View, CA); Lily Hao Yi Peng (Mountain View, CA); Katherine Irene Chou (Palo Alto, CA); Ananth Sankar (Palo Alto, CA)
Assignee: Google LLC
G16H40/20G06F40/47G06F40/58G10L15/063G10L15/142G10L15/16G10L15/183G10L15/1822G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,763,936
App. No.
17/112,279
Granted
Sep 19, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media for speech recognition. One method includes obtaining an input acoustic sequence, the input acoustic sequence representing one or more utterances; processing the input acoustic sequence using a speech recognition model to generate a transcription of the input acoustic sequence, wherein the speech recognition model comprises a domain-specific language model; and providing the generated transcription of the input acoustic sequence as input to a domain-specific predictive model to generate structured text content that is derived from the transcription of the input acoustic sequence.

Claims (29)

1. A computer implemented method comprising:

obtaining an input acoustic sequence, the input acoustic sequence representing one or more utterances;

processing the input acoustic sequence using a speech recognition model to generate a transcription of the input acoustic sequence; and

providing the generated transcription of the input acoustic sequence as input to a domain-specific predictive model to generate structured text content, wherein the domain-specific predictive model comprises a patient instructions predictive model configured to generate patient instructions based on at least a portion of the transcription of the input acoustic sequence that corresponds to utterances spoken by a medical professional, wherein the patient instructions include instructions regarding a future treatment for a patient.

2. The method of claim 1 , wherein the input acoustic sequence includes a digital representation of a conversation between a medical professional and a patient.

3. The method of claim 1 , wherein the patient instruction predictive model is further configured to generate patient instructions that are derived from the transcription of the input acoustic sequence and one or more of (i) the input acoustic sequence, (ii) data associated with the input acoustic sequence, (iii) an acoustic sequence representing a physician dictation, or (iv) data representing a patient's medical record.

4. The method of claim 1 , wherein the speech recognition model comprises a domain-specific language model.

5. The method of claim 4 , wherein the domain-specific language model comprises a medical language model that has been trained using medical-specific training data.

6. The method of claim 1 , wherein the domain-specific predictive model comprises an automated billing predictive model that is configured to generate a bill based on the transcription of the input acoustic sequence.

7. The method of claim 6 , wherein the automated billing predictive model is further configured to generate a bill that is based on the transcription of the input acoustic sequence and one or more of (i) the input acoustic sequence, (ii) data associated with the input acoustic sequence, (iii) an acoustic sequence representing a physician dictation, or (iv) data representing a patient's medical record.

8. The method of claim 1 , further comprising providing the input acoustic sequence as input to a speech prosody detection predictive model configured to process the input acoustic sequence to generate an indication of speech prosody that is derived from the input acoustic sequence.

9. The method of claim 8 , further comprising screening for diseases based on the generated indication of speech prosody.

10. The method of claim 9 , wherein the speech prosody detection predictive model is configured to provide, as output, a document listing results from the screening.

11. The method of claim 1 , wherein the domain-specific predictive model comprises a translation model that is configured to translate the transcription of the input acoustic sequence into a target language.

12. The method of claim 11 , wherein the translation model is further configured to translate the transcription of the input acoustic sequence into a target language using the input acoustic sequence.

13. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining an input acoustic sequence, the input acoustic sequence representing one or more utterances;

processing the input acoustic sequence using a speech recognition model to generate a transcription of the input acoustic sequence; and

providing the generated transcription of the input acoustic sequence as input to a domain-specific predictive model to generate structured text content, wherein the domain-specific predictive model comprises a patient instructions predictive model configured to generate patient instructions based on at least a portion of the transcription of the input acoustic sequence that corresponds to utterances spoken by a medical professional, wherein the patient instructions include instructions regarding a future treatment for a patient.

14. The system of claim 13 , wherein the patient instruction predictive model is further configured to generate patient instructions that are derived from the transcription of the input acoustic sequence and one or more of (i) the input acoustic sequence, (ii) data associated with the input acoustic sequence, (iii) an acoustic sequence representing a physician dictation, or (iv) data representing a patient's medical record.

15. The system of claim 13 , wherein the speech recognition model comprises a domain specific language model.

16. The system of claim 13 , wherein the speech recognition model comprises a hybrid deep neural network—hidden Markov model automatic speech recognition model.

17. The system of claim 13 , wherein the speech recognition model comprises an end-to-end speech recognition model with attention.

18. One or more non-transitory computer-readable storage media comprising instructions stored thereon that are executable by one or more processing devices and upon such execution cause the one or more processing devices to perform operations comprising:

obtaining an input acoustic sequence, the input acoustic sequence representing one or more utterances;

processing the input acoustic sequence using a speech recognition model to generate a transcription of the input acoustic sequence; and

providing the generated transcription of the input acoustic sequence as input to a domain-specific predictive model to generate structured text content, wherein the domain-specific predictive model comprises a patient instructions predictive model configured to generate patient instructions based on at least a portion of the transcription of the input acoustic sequence that corresponds to utterances spoken by a medical professional, wherein the patient instructions include instructions regarding a future treatment for a patient.

19. The one or more non-transitory computer-readable storage media of claim 18 , wherein the input acoustic sequence includes a digital representation of a conversation between a medical professional and a patient.

20. The one or more non-transitory computer-readable storage media of claim 18 , wherein the domain-specific predictive model comprises a translation model that is configured to translate the transcription of the input acoustic sequence into a target language.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2020
From: CO, CHRISTOPHER S.; JAITLY, NAVDEEP; PENG, LILY HAO YI; CHOU, KATHERINE IRENE; SANKAR, ANANTH
To: GOOGLE INC.
Reel/Frame 054550/0553 →
ENTITY CONVERSION Recorded Dec 4, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 054591/0993 →
Continuity (2)
Continuation 15362643 · Nov 28, 2016
Related Publication 20210090724A1 · Mar 25, 2021