IP Library Granted Patent US 12,315,624
Granted Patent B2
US 12,315,624 · App. 18/234,350 · Granted May 27, 2025

Generating structured text content using speech recognition models

Inventors: Christopher S. Co (Saratoga, CA); Navdeep Jaitly (Mountain View, CA); Lily Hao Yi Peng (Mountain View, CA); Katherine Irene Chou (Palo Alto, CA); Ananth Sankar (Palo Alto, CA)
Assignee: Google LLC
G16H40/20G06F40/47G06F40/58G10L15/063G10L15/142G10L15/16G10L15/1822G10L15/183G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,624
App. No.
18/234,350
Granted
May 27, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media for speech recognition. One method includes obtaining an input acoustic sequence, the input acoustic sequence representing one or more utterances; processing the input acoustic sequence using a speech recognition model to generate a transcription of the input acoustic sequence, wherein the speech recognition model comprises a domain-specific language model; and providing the generated transcription of the input acoustic sequence as input to a domain-specific predictive model to generate structured text content that is derived from the transcription of the input acoustic sequence.

Claims (27)

1. A computer implemented method comprising:

obtaining an input acoustic sequence that includes a digital representation of a conversation between a medical professional and a patient;

processing the input acoustic sequence using a speech recognition model to generate a transcription of the input acoustic sequence; and

providing the generated transcription of the input acoustic sequence as input to a domain-specific predictive model to generate structured text content, wherein the domain-specific predictive model comprises an automated billing predictive model that is configured to generate billing information based on the transcription of the input acoustic sequence, wherein the billing information comprises data indicating a cost associated with an interaction between the medical professional and the patient.

2. The method of claim 1 , wherein the automated billing predictive model is configured to generate bill information based on the transcription of the input acoustic sequence and one or more of (i) the input acoustic sequence, (ii) data associated with the input acoustic sequence, (iii) an acoustic sequence representing a physician dictation, or (iv) data representing a patient's medical record.

3. The method of claim 1 , wherein the billing information includes a billing code.

4. The method of claim 1 , wherein automated billing predictive model is configured to generate a bill from the billing information.

5. The method of claim 4 , wherein the automated billing predictive model is configured to populate a section of the bill with a summary of an interaction between a medical professional and a patient.

6. The method of claim 4 , wherein the generated bill comprises a formatted document that is organized into one or more sections or fields.

7. The method of claim 1 , wherein the speech recognition model comprises a domain-specific language model.

8. The method of claim 7 , wherein the domain-specific language model comprises a medical language model that has been trained using medical- specific training data.

9. The method of claim 1 , further comprising providing the input acoustic sequence as input to a speech prosody detection predictive model configured to process the input acoustic sequence to generate an indication of speech prosody that is derived from the input acoustic sequence.

10. The method of claim 9 , further comprising screening for diseases based on the generated indication of speech prosody.

11. The method of claim 9 , wherein the speech prosody detection predictive model is configured to provide, as output, a document listing results from the screening.

12. The method of claim 1 , wherein the domain-specific predictive model comprises a translation model that is configured to translate the transcription of the input acoustic sequence into a target language.

13. The method of claim 12 , wherein the translation model is further configured to translate the transcription of the input acoustic sequence into a target language using the input acoustic sequence.

14. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining an input acoustic sequence that includes a digital representation of a conversation between a medical professional and a patient;

processing the input acoustic sequence using a speech recognition model to generate a transcription of the input acoustic sequence; and

providing the generated transcription of the input acoustic sequence as input to a domain-specific predictive model to generate structured text content, wherein the domain-specific predictive model comprises an automated billing predictive model that is configured to generate billing information based on the transcription of the input acoustic sequence, wherein the billing information comprises data indicating a cost associated with an interaction between the medical professional and the patient.

15. The system of claim 14 , wherein the automated billing predictive model is configured to generate bill information based on the transcription of the input acoustic sequence and one or more of (i) the input acoustic sequence, (ii) data associated with the input acoustic sequence, (iii) an acoustic sequence representing a physician dictation, or (iv) data representing a patient's medical record.

16. The system of claim 14 , wherein the billing information includes a billing code.

17. The system of claim 14 , wherein the automated billing predictive model is configured to generate a bill from the billing information.

18. One or more non-transitory computer-readable storage media comprising instructions stored thereon that are executable by one or more processing devices and upon such execution cause the one or more processing devices to perform operations comprising:

obtaining an input acoustic sequence that includes a digital representation of a conversation between a medical professional and a patient;

processing the input acoustic sequence using a speech recognition model to generate a transcription of the input acoustic sequence; and

providing the generated transcription of the input acoustic sequence as input to a domain-specific predictive model to generate structured text content, wherein the domain-specific predictive model comprises an automated billing predictive model that is configured to generate billing information based on the transcription of the input acoustic sequence, wherein the billing information comprises data indicating a cost associated with an interaction between the medical professional and the patient.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: CO, CHRISTOPHER S.; JAITLY, NAVDEEP; PENG, LILY HAO YI; CHOU, KATHERINE IRENE; SANKAR, ANANTH
To: GOOGLE, INC.
Reel/Frame 065505/0165 →
CHANGE OF NAME Recorded Nov 9, 2023
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 065530/0754 →
Continuity (3)
Continuation 17112279 · Dec 4, 2020
Continuation 15362643 · Nov 28, 2016
Related Publication 20230386652A1 · Nov 30, 2023
References Cited (49)
US 7668718B2 · Kahn · 2010 [cited by applicant]
US 9190050B2 · Mathias · 2015 [cited by examiner]
US 9569593B2 · Casella dos Santos · 2017 [cited by examiner]
US 9785753B2 · Casella dos Santos · 2017 [cited by applicant]
US 9966066B1 · Corfield · 2018 [cited by examiner]
US 10199124B2 · Casella dos Santos · 2019 [cited by applicant]
US 10283110B2 · Bellegarda · 2019 [cited by examiner]
US 10296160B2 · Shah · 2019 [cited by examiner]
US 10339937B2 · Koll et al. · 2019 [cited by applicant]
US 10860685B2 · Co · 2020 [cited by examiner]
US 20020072896A1 · Roberge et al. · 2002 [cited by applicant]
US 20020087357A1 · Singer · 2002 [cited by applicant]
US 20030182117A1 · Monchi et al. · 2003 [cited by applicant]
US 20040019482A1 · Holub · 2004 [cited by applicant]
US 20060041428A1 · Fritsch et al. · 2006 [cited by applicant]
US 20060149558A1 · Kahn · 2006 [cited by examiner]
US 20090326937A1 · Chitsaz et al. · 2009 [cited by applicant]
US 20100088095A1 · John · 2010 [cited by applicant]
US 20120323574A1 · Wang · 2012 [cited by examiner]
US 20130238312A1 · Waibel · 2013 [cited by examiner]
US 20130238330A1 · Casella dos Santos · 2013 [cited by examiner]
US 20140019128A1 · Riskin · 2014 [cited by examiner]
US 20140249818A1 · Yegnanarayanan · 2014 [cited by examiner]
US 20140324477A1 · Oez · 2014 [cited by applicant]
US 20160078661A1 · Iurascu · 2016 [cited by applicant]
US 20160196821A1 · Yegnanarayanan · 2016 [cited by applicant]
US 20170116392A1 · Casella dos Santos · 2017 [cited by applicant]
US 20180150605A1 · Co · 2018 [cited by examiner]
US 20200134571A1 · Demick · 2020 [cited by examiner]
US 20200143946A1 · Lewis · 2020 [cited by examiner]
US 20230386652A1 · Co · 2023 [cited by examiner]
CN 104380375A · 2015 [cited by applicant]
EP 1787288A2 · 2007 [cited by applicant]
JP 2008136646A · 2008 [cited by applicant]
JP 2009541800A · 2009 [cited by applicant]
WO WO2006023622A2 · 2006 [cited by applicant]
Azzini et al. “Application of Spoken Dialogue Technology in a Medical Domain,” Text, Speech and Dialogue, Springer Berlin Heidelberg, vol. 2448, Jan. 1, 2006, 4 pages. [cited by applicant]
Chan et al., “Listen, Attend and Spell,” (Aug. 20, 2015) [online] (retrieved from https://arxiv.org/pdf/1508.01211v2.pdf), 16 pages. [cited by applicant]
EP Office Action in European Application No. 17811802, dated Jul. 8, 2020, 6 pages. [cited by applicant]
Graves et al., Hybrid Speech Recognition with Bidirectional LSTM (2013) [online] (retrieved from http://www.cs.toronto.edu/˜graves/asru_2013_poster.pdf) 1 page. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2017/063301, dated Jun. 6, 2019, 15 pages. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US2017/063301, mailed on Feb. 8, 2018, 15 pages. [cited by applicant]
Nguyen et al., “Deepr: A Convolutional Net for Medical Records,” (Jul. 26, 2016) [online] (retrieved from https://arxiv.org/pdf/1607.07519v1.pdf), 9 pages. [cited by applicant]
Office Action in Chinese Appln. No. 201780073503.3, dated Dec. 2, 2022, 33 pages (with English Translation). [cited by applicant]
Shang et al., “Neural Responding Machine for Short-text Conversation,” Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on National Langu… [cited by applicant]
Shriberg and Stolcke, “Prosody Modeling for Automatic Speech Recognition and Understanding,” (Nov. 28, 2016) [online] (retrieved from http://www.icsi.berkeley.edu/pubs/speech/prosodymodeling.pdf), 10 pages. [cited by applicant]
Sordoni et al., “A Neural Network Approach to Context-sensitive Generation of Conversational Responses,” Human Language Technologies: The 2015 Annual Conference of the North American Chapter of the ACL, pp. 196-205, Den… [cited by applicant]
Vinyals and Le, “A Neural Conversational Model,” (Jul. 22, 2015) [online] (retrieved from https://arxiv.org/pdf/1506.05869.pdf), 8 pages. [cited by applicant]
Wolk and Marasek, “Neural-based Machine Translation for Medical Text Domain. Based on European Medicines Agency Leaflet Texts,” (Nov. 28, 2016) [online] (retrieved from https://arxiv.org/ftp/arxiv/papers/1509/1509.08644… [cited by applicant]
Cited By (2)
US 12,620,008 US 12,711,418