IP Library Patent Application 18423490
Patent Application
App. No. 18/423,490

AUTOMATIC SPEECH RECOGNITION SYSTEM CONTEXTUALLY BIASED FOR MEDICAL SPEECH

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/423,490
Abstract

Methods and systems of generating text representation of spoken medical speech are presented herein. Some methods may include the steps of providing a pre-trained automatic speech recognition (ASR) system stored in memory and executed on a processor; receiving, by the pre-trained ASR system, spoken medical speech; and generating text of the spoken medical speech by biasing the pre-trained ASR system using a contextual language model, where the contextual language model may include medical terminology that is not included in a vocabulary used to train the pre-trained ASR system.

Claims (35)

1 . A method of generating text of medical speech, the method comprising:

providing a pre-trained automatic speech recognition (ASR) system stored in memory and executed on a processor;

receiving, by the pre-trained ASR system, spoken medical speech; and

generating text of the spoken medical speech by biasing the pre-trained ASR system using a contextual language model, wherein the contextual language model comprises medical terminology that is not included in a vocabulary used to train the pre-trained ASR system.

2 . The method of claim 1 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been jointly trained using the vocabulary.

3 . The method of claim 1 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been separately trained, and wherein the language model is trained using the vocabulary.

4 . The method of claim 1 , wherein the biased ASR system is a shallow fusion model.

5 . The method of claim 1 , wherein the medical terminology comprises a plurality of medical terms.

6 . The method of claim 1 , wherein the contextual language model is a contextual n-gram language model.

7 . The method of claim 6 , wherein the step of generating text of the spoken medical speech comprises determining an n-gram score based on an overall model score generated by the pre-trained ASR system and a bias score generated by the contextual language model to generate a textual representation of a medical term that is not included in the vocabulary used to train the pre-trained ASR system.

8 . The method of claim 1 , wherein the language model biases the ASR system during beam searching.

9 . The method of claim 1 , wherein the language model biases the ASR system before beam searching.

10 . A method of generating a medical report, comprising:

the method of claim 1 ; and,

writing a report based on the text of the medical speech.

11 . A system for generating text of spoken medical speech comprising:

an input interface configured to receive spoken medical speech;

a memory configured to store a plurality of processor-executable instruction, the memory including:

a pre-trained ASR system; and,

a contextual language model, wherein the contextual language model receives a plurality of medical terms; and,

a processor configured to execute the plurality of processor-executable instructions to perform operations including:

biasing the pre-trained ASR system using the contextual language model; and,

generating text of the spoken medical speech using the biased pre-trained ASR system, wherein at least one of the plurality of medical terms is not included in a vocabulary used to train the pre-trained ASR system.

12 . The system of claim 11 , wherein the biased pretrained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been jointly trained using the vocabulary.

13 . The system of claim 11 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been separately trained, and wherein the language model is trained using the vocabulary.

14 . The system of claim 12 , wherein generating text of the spoken medical speech comprises determining an n-gram score based on an overall model score generated by the pre-trained ASR system and a bias score generated by the contextual language model.

15 . The system of claim 11 , wherein the contextual language model biases the pre-trained ASR system during beam search decoding.

16 . A non-transitory processor-readable storage medium storing a plurality of processor-executable instructions for generating text of spoken medical speech, the instructions being executed by a processor to perform operations comprising:

providing a pre-trained automatic speech recognition (ASR) model;

biasing the pre-trained ASR model using a contextual language model, wherein the contextual language model comprises medical terminology; and,

generating text of the spoken medical speech by biasing the pre-trained ASR system using a contextual language model, wherein the contextual language model comprises medical terminology that is not included in a vocabulary used to train the pre-trained ASR system

17 . The non-transitory processor-readable storage medium of claim 16 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been jointly trained using the vocabulary.

18 . The non-transitory processor-readable storage medium of claim 16 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been separately trained, and wherein the language model is trained using the vocabulary.

19 . The non-transitory processor-readable storage medium of claim 16 , wherein the contextual language model is a contextual n-gram language model.

20 . The non-transitory processor-readable storage medium of claim 19 , wherein generating text of the spoken medical speech comprises determining an n-gram score based on an overall model score generated by the pre-trained ASR system and a bias score generated by the contextual language model.

Assignments (2)
CHANGE OF NAME Recorded May 4, 2026
From: VERILY LIFE SCIENCES LLC
To: VERILY HEALTH INC.
Reel/Frame 075501/0627 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2024
From: SHOR, JOEL
To: VERILY LIFE SCIENCES LLC
Reel/Frame 066458/0077 →