IP Library › Granted Patent US 11,650,732
Granted Patent B2
US 11,650,732 · App. 17/819,698 · Granted May 16, 2023

Method and system for generating transcripts of patient-healthcare provider conversations

Inventors: Melissa Strader (San Jose, CA); William Ito (Mountain View, CA); Christopher Co (Saratoga, CA); Katherine Chou (Palo Alto, CA); Alvin Rajkomar (Mountain View, CA); Rebecca Rolfe (Menlo Park, CA)
Assignee: Google LLC
G06F3/04855G06F40/289G06F40/30G10L15/26G16H10/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,650,732
App. No.
17/819,698
Granted
May 16, 2023
Kind
B2
Abstract

A method and workstation for generating a transcript of a conversation between a patient and a healthcare practitioner is disclosed. A workstation is provided with a tool for rendering of an audio recording of the conversation and generating a display of a transcript of the audio recording using a speech-to-text engine, thereby enabling inspection of the accuracy of conversion of speech to text. A tool is provided for scrolling through the transcript and rendering the portion of the audio according to the position of the scrolling. There is a highlighting in the transcript of words or phrases spoken by the patient relating to symptoms, medications or other medically relevant concepts. Additionally, there is provided a set of transcript supplement tools enabling editing of specific portions of the transcript based on the content of the corresponding portion of audio recording.

Claims (62)

1. A method for generating a user interface displaying a plurality of portions associated with a conversation between a patient and a healthcare practitioner, comprising:

providing, in a first portion of the plurality of portions, a rendering of an audio recording of the conversation;

displaying, in a second portion of the plurality of portions, a transcript of the audio recording generated using a speech-to-text engine;

displaying, in a third portion of the plurality of portions, a note summarizing the conversation, wherein the displaying of the note comprises:

automatically recognizing, in the transcript, a text segment spoken by the patient relating to one or more of symptoms, medications or other medically relevant concepts, and

populating the note with the recognized text segment; and

linking the plurality of portions of the user interface, wherein the linking comprises at least a linking of the recognized text segment in the note to a relevant part of the transcript from which the recognized text segment in the note originated.

2. The method of claim 1 , further comprising:

enabling editing of the relevant part of the transcript based on a content of a corresponding portion of the audio recording comprising an audio version of the recognized text segment; and

subsequent to a user selection of the linked recognized text segment in the note, enabling playback of the corresponding portion of the audio recording by providing one or more audio playback options.

3. The method of claim 2 , wherein the enabling of the one or more audio playback options further comprises:

speeding up and skipping of at least another portion of the audio recording that does not comprise the recognized text segment.

4. The method of claim 2 , wherein the one or more audio playback options comprise pause, rewind, play, or fast forward.

5. The method of claim 1 , wherein the displaying of the note further comprises:

determining that the recognized text segment is a common term corresponding to a medically preferred term; and

replacing, in the note, the common term with the medically preferred term.

6. The method of claim 1 , further comprising:

providing a set of note supplement tools enabling editing of a particular portion of the note based on a content of a corresponding portion of the transcript or the audio recording.

7. The method of claim 6 , wherein the set of note supplement tools include at least one of:

a) a display of smart suggestions for words or phrases and a tool for editing, approving, rejecting or providing feedback on the smart suggestions;

b) a display of suggested corrected medical terminology; and

c) a display of an indication of confidence level in suggested words or phrases.

8. The method of claim 7 , wherein the set of note supplement tools include each of the displays a), b) and c) recited in claim 7 .

9. The method of claim 7 , wherein the display of the indication of the confidence level comprises display upon a determination that the confidence level exceeds a threshold confidence level.

10. The method of claim 1 , further comprising:

minimizing display of the transcript and enabling viewing the note only, and wherein the note is generated in substantial real time with the rendering of the audio recording.

11. The method of claim 1 , wherein the recognized text segment is placed into an appropriate category in the note.

12. The method of claim 1 , further comprising:

providing, in the note, supplementary information for one or more symptoms, wherein the supplementary information comprises one or more labels indicating phrases associated with medical billing terminology.

13. The method of claim 1 , further comprising:

a set of transcript supplement tools to enable editing of a particular portion of the transcript based on a content of a corresponding portion of the audio recording.

14. The method of claim 1 , wherein the automatically recognizing of the text segment is performed by using a trained machine learning model.

15. A server for displaying a transcript of a conversation between a patient and a healthcare practitioner, comprising:

one or more processors; and

memory storing computer-executable instructions that, when executed by the one or more processors, cause the server to perform operations comprising:

applying a machine learning model to generate the transcript of an audio recording of the conversation using a speech-to-text engine in substantial real time with a rendering of the audio recording, wherein the machine learning model is trained to automatically recognize, in the audio recording or in the transcript, a text segment spoken by the patient relating to one or more of symptoms, medications or other medically relevant concepts; and

transmitting, by an application programming interface to a workstation, the transcript and a link to the audio recording, and wherein the instructions cause the workstation to perform operations comprising:

providing, in a first portion of the plurality of portions, a rendering of an audio recording of the conversation;

displaying, in a second portion of the plurality of portions, a transcript of the audio recording generated using a speech-to-text engine;

displaying, in a third portion of the plurality of portions, a note summarizing the conversation, wherein the displaying of the note comprises:

automatically recognizing, in the transcript, a text segment spoken by the patient relating to one or more of symptoms, medications or other medically relevant concepts, and

populating the note with the recognized text segment; and

linking the plurality of portions of the user interface, wherein the linking of the plurality of portions comprises at least a linking of the recognized text segment in the note to a relevant part of the transcript from which the recognized text segment in the note originated.

16. The server of claim 15 , wherein the instructions cause the server to perform operations further comprising:

enabling editing of the relevant part of the transcript based on a content of a corresponding portion of the audio recording comprising an audio version of the recognized text segment; and

subsequent to a user selection of the linked recognized text segment in the note, enabling playback of the corresponding portion of the audio recording by providing one or more audio playback options.

17. The server of claim 16 , wherein the instructions cause the server to perform operations further comprising:

speeding up and skipping of at least another portion of the audio recording that does not comprise the recognized text segment.

18. The server of claim 16 , wherein the one or more audio playback options comprise pause, rewind, play, or fast forward.

19. The server of claim 16 , wherein the instructions for the displaying of the note further comprising:

determining that the recognized text segment is a common term corresponding to a medically preferred term; and

replacing, in the note, the common term with the medically preferred term.

20. The server of claim 15 , wherein the instructions cause the server to:

apply the machine learning model to generate the note in substantial real time with the rendering of the audio recording; and

transmit the note by the application programming interface to the workstation.

21. An article manufacture comprising one or more computer readable media having computer-readable instructions stored thereon that, when executed by one or more processors of a computing device, cause the computing device to perform operations comprising:

providing, in a first portion of a plurality of portions of a user interface associated with a conversation between a patient and a healthcare practitioner, a rendering of an audio recording of the conversation;

displaying, in a second portion of the plurality of portions, a transcript of the audio recording generated using a speech-to-text engine;

displaying, in a third portion of the plurality of portions, a note summarizing the conversation, wherein the displaying of the note comprises:

automatically recognizing, in the transcript, a text segment spoken by the patient relating to one or more of symptoms, medications or other medically relevant concepts, and

populating the note with the recognized text segment; and

linking the plurality of portions of the user interface, wherein the linking comprises at least a linking of the recognized text segment in the note to a relevant part of the transcript from which the recognized text segment in the note originated.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2022
From: STRADER, MELISSA; ITO, WILLIAM; CO, CHRISTOPHER; CHOU, KATHERINE; RAJKOMAR, ALVIN; ROLFE, REBECCA
To: GOOGLE LLC
Reel/Frame 060805/0604 →
Continuity (5)
Continuation 17215512 · Mar 29, 2021
Continuation 16909115 · Jun 23, 2020
Continuation 15988657 · May 24, 2018
Provisional Application 62575725 · Oct 23, 2017
Related Publication 20220391083A1 · Dec 8, 2022