IP Library › Granted Patent US 10,990,266
Granted Patent B2
US 10,990,266 · App. 16/909,115 · Granted Apr 27, 2021

Method and system for generating transcripts of patient-healthcare provider conversations

Inventors: Melissa Strader (San Jose, CA); William Ito (Mountain View, CA); Christopher Co (Saratoga, CA); Katherine Chou (Palo Alto, CA); Alvin Rajkomar (Mountain View, CA); Rebecca Rolfe (Menlo Park, CA)
Assignee: GOOGLE LLC
G06F3/04855G06F40/289G06F40/30G10L15/26G16H10/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,990,266
App. No.
16/909,115
Granted
Apr 27, 2021
Kind
B2
Abstract

A method and workstation for generating a transcript of a conversation between a patient and a healthcare practitioner is disclosed. A workstation is provided with a tool for rendering of an audio recording of the conversation and generating a display of a transcript of the audio recording using a speech-to-text engine, thereby enabling inspection of the accuracy of conversion of speech to text. A tool is provided for scrolling through the transcript and rendering the portion of the audio according to the position of the scrolling. There is a highlighting in the transcript of words or phrases spoken by the patient relating to symptoms, medications or other medically relevant concepts. Additionally, there is provided a set of transcript supplement tools enabling editing of specific portions of the transcript based on the content of the corresponding portion of audio recording.

Claims (65)

1. A method for generating a transcript of a conversation between a patient and a healthcare practitioner, comprising:

providing a rendering of an audio recording of the conversation and generating a display of a transcript of the audio recording using a speech-to-text engine in substantial real time with the rendering of the audio recording, thereby enabling inspection of the accuracy of conversion of speech to text;

providing a user interface element enabling scrolling through the transcript and rendering a portion of the audio according to a position of the scrolling;

receiving a prediction from a trained machine learning model automatically recognizing in the transcript or in the recording words or phrases spoken by the patient relating to symptoms, medications or other medically relevant concepts;

highlighting such recognized words or phrases spoken by the patient in the transcript; and

providing a set of transcript supplement tools enabling editing of specific portions of the transcript based on a content of a corresponding portion of the audio recording.

2. The method of claim 1 , wherein the transcript supplement tools include at least one of the following:

a) a display of smart suggestions for words or phrases and a tool for editing, approving, rejecting or providing feedback on the suggestions;

b) a display of suggested corrected medical terminology;

c) a display of an indication of confidence level in suggested words or phrases.

3. The method of claim 2 , wherein the transcript supplement tools include each of the displays a), b) and c) recited in claim 2 .

4. A computer-implemented system for displaying a transcript of a conversation between a patient and a healthcare practitioner, comprising:

a ocessor: and

memory storing computer-readable instructions that, when executed by the processor, cause:

a user interface element to trigger a rendering of an audio recording of the conversation and generating a display of the transcript of the audio recording using a speech-to-text engine in substantial real time with the rendering of the audio recording, thereby enabling inspection of the accuracy of conversion of speech to text;

a user interface element to enable scrolling through the transcript and rendering the portion of the audio according to the position of the scrolling;

a machine learning model to automatically recognize in the recording or in the transcript words or phrases spoken by the patient relating to symptoms, medications or other medically relevant concepts and wherein the display of the transcript includes a highlighting of such recognized words or phrases spoken by the patient; and

a set of transcript supplement tools to enable editing of specific portions of the transcript based on the content of the corresponding portion of audio recording.

5. The system of claim 4 , wherein the transcript supplement tools include at least one of the following:

a) a display of smart suggestions for words or phrases and a tool for editing, approving, rejecting or providing feedback on the suggestions;

b) a display of suggested corrected medical terminology;

c) a display of an indication of confidence level in suggested words or phrases.

6. The system of claim 5 , wherein the transcript supplement tools includes each of the displays a), b), and c) recited in claim 5 .

7. The system of claim 5 , wherein the instructions, when executed by the processor, cause:

display of a note simultaneously with the display of the transcript and populating the note with the highlighted words or phrases in substantial real time with the rendering of the audio; and providing supplementary information for symptom s including labels for phrases required for billing.

8. The system of claim 7 , wherein the instructions, when executed by the processor, cause:

linking of words or phrases in the note to relevant parts of the transcript from which the words or phrases in the note originated.

9. A server for displaying a transcript of a conversation between a patient and a healthcare practitioner, comprising:

one or more processors; and

memory storing computer-executable instructions that, when executed by the one or more processors, cause the server to:

receive an audio recording of the conversation;

apply a machine learning model to generate a transcript of the audio recording using a speech-to-text engine in substantial real time with a rendering of the audio recording, thereby enabling inspection of an accuracy of conversion of speech to text, wherein the machine learning model is trained to automatically recognize in the recording or in the transcript words or phrases spoken by the patient relating to symptoms, medications or other medically relevant concepts; and

transmit the transcript by an application programming interface to a workstation, and wherein the instructions cause the workstation to:

generate a display of the transcript of the audio recording;

provide a user interface element enabling scrolling through the transcript and rendering a portion of the audio according to a position of the scrolling;

highlight the recognized words or phrases spoken by the patient in the transcript; and

provide a set of transcript supplement tools enabling editing of specific portions of the transcript based on a content of a corresponding portion of the audio recording.

10. The server of claim 9 , wherein the transcript supplement tools include at least one of the following:

a) a display of smart suggestions for words or phrases and a tool for editing, approving, rejecting or providing feedback on the suggestions;

b) a display of suggested corrected medical terminology;

c) a display of an indication of confidence level in suggested words or phrases.

11. The server of claim 10 , wherein the transcript supplement tools include each of the displays a), b) and c) recited in claim 10 .

12. The server of claim 9 , wherein the instructions further cause theworkstation to:

display a note simultaneously with the display of the transcript and populating the note with the highlighted words or phrases in substantial real time with the rendering of the audio; and

provide supplementary information for symptoms including labels for phrases required for billing.

13. The server of claim 12 , wherein the instructions further cause the workstation to link words or phrases in the note to relevant parts of the - transcript from which the words or phrases in the note originated.

14. An article of manufacture comprising one or more computer eadable media having computer-readable instructions stored thereon that, when executed by one or more processors of a computing device, cause the computing device to:

provide a rendering of an audio recording of the conversation and generating a display of a transcript of the audio recording using a speech-to-text engine in substantial real time with the rendering of the audio recording, thereby enabling inspection of the accuracy of conversion of speech to text;

provide a user interface element enabling scrolling through the transcript and rendering a portion of the audio according to a position of the scrolling;

receive a prediction from a trained machine learning model automatically recognizing in the transcript, or in the recording words or phrases spoken by the patient relating to sympto s, medications or other medically relevant concepts;

highlight such recognized words or phrases spoken by the patient in the transcript; and

provide a set of transcript supplement tools enabling editing of specific portions of the transcript based on a content of a corresponding portion of the audio recording.

15. The article of manufacture of claim 14 , wherein the transcript supplement tools include at least one of the following:

a) a display of smart suggestions for words or phrases and a tool for editing, approving, rejecting or providing feedback on the suggestions;

b) a display of suggested corrected medical terminology;

c) a display of an indication of confidence level in suggested words or phrases.

16. The article of manufacture of claim 15 , wherein the transcript supplement tools include each of the displays a), b) and c) recited in claim 15 .

17. The article of manufacture of claim 14 , wherein the instructions, when executed by the one or more processors, cause the computing device to:

display a note simultaneously with the display of the transcript and populating the note with the highlighted words or phrases in substantial real time with the rendering of the audio; and providing supplementary information for symptoms including labels for phrases required for billing.

18. The article of manufacture of claim 17 . wherein the instructions, when executed by the one or more processors, cause the computing device to:

link words or phrases in the note to relevant parts of the transcript from which the words or phrases in the note originated.

19. The method of claim 1 , further comprising:

displaying a note simultaneously with the display of the transcript and populating the note with the highlighted words or phrases in substantial real time with the rendering of the audio; and providing supplementary information for symptoms including labels for phrases required for billing.

20. The method of claim 19 , further comprising:

linking words or phrases in the note to relevant parts of the transcript from which the words or phrases in the note originated.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 12, 2021
From: STRADER, MELISSA; ITO, WILLIAM; CO, CHRISTOPHER; CHOU, KATHERINE; RAJKOMAR, ALVIN; ROLFE, REBECCA
To: GOOGLE LLC
Reel/Frame 055248/0569 →
Continuity (3)
Continuation 15988657 · May 24, 2018
Provisional Application 62575725 · Oct 23, 2017
Related Publication 20200319787A1 · Oct 8, 2020