IP Library › Granted Patent US 12,597,512
Granted Patent B2
US 12,597,512 · App. 19/302,810 · Granted Apr 7, 2026

Real-time use of multiple parallel automatic speech recognition (ASR) modules in a conversational artificial intelligence (AI) architecture

Inventors: Subhabrata Mukherjee (Palo Alto, CA); Shanil Puri (Palo Alto, CA); Tanmay Laud (Palo Alto, CA); Woojeong Jin (Palo Alto, CA); Jan Schellenberger (Palo Alto, CA)
Assignee: HealthGPT, Inc.
G16H40/20G10L15/183G10L15/22G10L25/66G10L15/1822G10L15/19
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,512
App. No.
19/302,810
Filed
Aug 18, 2025
Granted
Apr 7, 2026
Kind
B2
Examiner
HANG, VU B
Art Unit
2654
USPC
704/257
Abstract

As an example, a conversational artificial intelligence (AI) that is in a conversation receives a human response from a human, augments the human response to create an augmented response, determines a context of the conversation, and provides the context and the augmented response to a plurality of automatic speech recognition (ASR) modules that individually process the augmented response in parallel. The conversational AI receives a plurality of intermediate text outputs from the plurality of ASR modules, wherein individual intermediate text outputs of the plurality of intermediate text outputs are received from individual ones of the ASR modules. A reconciliation AI performs a contextual reconciliation of the plurality of intermediate text outputs based at least in part on the context of the conversation to create a final text output. The conversational AI provides, in real-time, an artificial intelligence response to the human based on the final text output.

Claims (104)

1 . A method, comprising:

initiating, by a conversational artificial intelligence executed by one or more processors, a conversation with a human;

receiving, by the conversational artificial intelligence, a human response from the human;

augmenting, by the one or more processors, the human response to create an augmented response;

determining, by the one or more processors, a context of the conversation;

providing, by the one or more processors, the context of the conversation and the augmented response to a plurality of automatic speech recognition (ASR) modules that individually process the augmented response in parallel;

receiving, by the one or more processors, a plurality of intermediate text outputs from the plurality of ASR modules, wherein individual intermediate text outputs of the plurality of intermediate text outputs are received from individual modules of the plurality of ASR modules;

performing in real-time, by a reconciliation artificial intelligence, a contextual reconciliation of the plurality of intermediate text outputs based at least in part on the context of the conversation and on a medical history of the human to create a final text output; and

generating, by the conversational artificial intelligence, an artificial intelligence response to the human based at least in part on the final text output.

2 . The method of claim 1 , wherein augmenting the human response to create the augmented response comprises at least one of:

performing an intonation analysis of the human response;

performing noise cancellation by reducing an amount of background noise present in the human response; or

any combination thereof.

3 . The method of claim 1 , wherein determining the context of the conversation comprises:

accessing electronic medical records (EMR) associated with the human; and

determining a conversation history of the conversation between the conversational artificial intelligence and the human.

4 . The method of claim 1 , wherein the plurality of ASR modules comprise at least:

a first ASR module implemented using a first artificial intelligence algorithm;

a second ASR module implemented using a second artificial intelligence algorithm; and

a third ASR module implemented using a third artificial intelligence algorithm;

wherein the first artificial intelligence algorithm, the second artificial intelligence algorithm, and the third artificial intelligence algorithm are different from each other.

5 . The method of claim 1 , wherein the plurality of ASR modules comprise at least:

a first ASR module trained using a first corpus of speech of people having a first type of accent when speaking a particular language;

a second ASR module trained using a second corpus of speech of people having a second type of accent when speaking the particular language; and

a third ASR module trained using a third corpus of speech of people having a third type of accent when speaking the particular language;

wherein the first type of accent, the second type of accent, and the third type of accent are different from each other.

6 . The method of claim 1 , wherein the conversational artificial intelligence is trained using multi-turn reinforcement learning through human feedback (RLHF).

7 . The method of claim 1 , wherein the conversational artificial intelligence, during the conversation with the human:

identifies a turn-yielding cue;

performs interruption detection to detect when the human is attempting to interrupt the conversationa; artificial intelligence;

identifies a non-verbal cue associated with the human; or

any combination thereof.

8 . A server comprising:

one or more processors; and

one or more computer-readable storage media to store instructions executable by the one or more processors to perform operations comprising:

initiating, by a conversational artificial intelligence, a conversation with a human;

receiving, by the conversational artificial intelligence, a human response from the human;

augmenting the human response to create an augmented response;

determining a context of the conversation;

providing the context of the conversation and the augmented response to a plurality of automatic speech recognition (ASR) modules that individually process the augmented response in parallel;

receiving a plurality of intermediate text outputs from the plurality of ASR modules, wherein individual intermediate text outputs of the plurality of intermediate text outputs are received from individual modules of the plurality of ASR modules;

performing in real-time, by a reconciliation artificial intelligence, a contextual reconciliation of the plurality of intermediate text outputs based at least in part on the context of the conversation and on a medical history of the human to create a final text output; and

generating, by the conversational artificial intelligence, an artificial intelligence response to the human based at least in part on the final text output.

9 . The server of claim 8 , wherein augmenting the human response to create the augmented response comprises at least one of:

performing an intonation analysis of the human response including determining a volume, a pitch, and a rhythm of individual words in the human response;

performing noise cancellation by identifying speech content spoken by the human and reducing the volume of other content in the human response including other human speech; or

any combination thereof.

10 . The server of claim 8 , wherein determining the context of the conversation comprises:

determining electronic medical records (EMR) associated with the human; and

determining a conversation history of the conversation between the conversational artificial intelligence and the human.

11 . The server of claim 10 , wherein the conversation history of the conversation between the conversational artificial intelligence and the human is stored in a cache memory.

12 . The server of claim 8 , wherein the plurality of ASR modules comprise at least:

a first ASR module implemented using a first artificial intelligence algorithm;

a second ASR module implemented using a second artificial intelligence algorithm; and

a third ASR module implemented using a third artificial intelligence algorithm;

wherein the first artificial intelligence algorithm, the second artificial intelligence algorithm, and the third artificial intelligence algorithm are different from each other.

13 . The server of claim 8 , wherein the plurality of ASR modules comprise at least:

a first ASR module trained using a first corpus of speech of people having a first type of accent when speaking a particular language;

a second ASR module trained using a second corpus of speech of people having a second type of accent when speaking the particular language; and

a third ASR module trained using a third corpus of speech of people having a third type of accent when speaking the particular language;

wherein the first type of accent, the second type of accent, and the third type of accent are different from each other.

14 . The server of claim 8 , wherein the conversational artificial intelligence is configured to perform specialized healthcare-related functions comprising one or more of:

gathering data related to performing Healthcare Effectiveness Data and Information Set (HEDIS) calculations;

performing a Health Records Assessment (HRA);

determining a Risk Adjustment Factor (RAF);

reviewing a pre-op checklist;

reviewing a discharge checklist;

reviewing a chronic care checklist;

determining social determinants of health (SDOH); or

any combination thereof.

15 . A non-transitory memory device to store instructions executable by one or more processors to perform operations comprising:

initiating, by a conversational artificial intelligence, a conversation with a human;

receiving, by the conversational artificial intelligence, a human response from the human;

augmenting the human response to create an augmented response;

determining a context of the conversation;

providing the context of the conversation and the augmented response to a plurality of automatic speech recognition (ASR) modules that individually process the augmented response in parallel;

receiving a plurality of intermediate text outputs from the plurality of ASR modules, wherein individual intermediate text outputs of the plurality of intermediate text outputs are received from individual modules of the plurality of ASR modules;

performing in real-time, by a reconciliation artificial intelligence, a contextual reconciliation of the plurality of intermediate text outputs based at least in part on the context of the conversation and on a medical history of the patient to create a final text output; and

generating, by the conversational artificial intelligence, an artificial intelligence response to the human based at least in part on the final text output.

16 . The non-transitory memory device of claim 15 , wherein augmenting the human response to create the augmented response comprises at least one of:

performing an intonation analysis of the human response including determining a volume, a pitch, and a rhythm of individual words in the human response;

performing noise cancellation by identifying speech content spoken by the human and reducing the volume of other content in the human response including other human speech; or

any combination thereof.

17 . The non-transitory memory device of claim 15 , wherein determining the context of the conversation comprises:

determining electronic medical records (EMR) associated with the human; and

determining a conversation history of the conversation between the conversational artificial intelligence and the human.

18 . The non-transitory memory device of claim 15 , wherein the plurality of ASR modules comprise at least:

a first ASR module implemented using a first artificial intelligence algorithm;

a second ASR module implemented using a second artificial intelligence algorithm; and

a third ASR module implemented using a third artificial intelligence algorithm;

wherein the first artificial intelligence algorithm, the second artificial intelligence algorithm, and the third artificial intelligence algorithm are different from each other.

19 . The non-transitory memory device of claim 15 , wherein the plurality of ASR modules comprise at least:

a first ASR module trained using a first corpus of speech of people having a first type of accent when speaking a particular language;

a second ASR module trained using a second corpus of speech of people having a second type of accent when speaking the particular language; and

a third ASR module trained using a third corpus of speech of people having a third type of accent when speaking the particular language;

wherein the first type of accent, the type of second accent, and the third type of accent are different from each other.

20 . The non-transitory memory device of claim 15 , wherein the conversational artificial intelligence is engaged in a task that includes one or more of:

performing a preventative screening;

an intake-related task;

a scheduling-related task;

a pre-op related task;

a discharge-related task;

a chronic care related task; or

any combination thereof.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2025
From: MUKHERJEE, SUBHABRATA; PURI, SHANIL; LAUD, TANMAY; JIN, WOOJEONG; SCHELLENBERGER, JAN
To: HEALTHGPT, INC. DBA HIPPOCRATIC AI
Reel/Frame 072944/0854 →
Continuity (5)
Continuation In Part 18900289 · Sep 27, 2024
Continuation 18592441 · Feb 29, 2024
Provisional Application 63611762 · Dec 18, 2023
Provisional Application 63466712 · May 15, 2023
Related Publication 20250384995A1 · Dec 18, 2025
References Cited (31)
US 8548937B2 · Saigal et al. · 2013 [cited by applicant]
US 9824188B2 · Brown et al. · 2017 [cited by applicant]
US 10282512B2 · Bennett et al. · 2019 [cited by applicant]
US 10452816B2 · Kidd et al. · 2019 [cited by applicant]
US 10504379B2 · Vinkers · 2019 [cited by examiner]
US 10748644B2 · Shriberg et al. · 2020 [cited by applicant]
US 11329933B1 · Kiyanda et al. · 2022 [cited by applicant]
US 11348694B2 · Kim et al. · 2022 [cited by applicant]
US 11532393B2 · Arkoff et al. · 2022 [cited by applicant]
US 11693990B1 · Arkoff et al. · 2023 [cited by applicant]
US 11694807B2 · Golan et al. · 2023 [cited by applicant]
US 11843565B2 · Lee et al. · 2023 [cited by applicant]
US 11977854B2 · Tunstall-Pedoe et al. · 2024 [cited by applicant]
US 20140365885A1 · Carson · 2014 [cited by examiner]
US 20190341052A1 · Allibhai · 2019 [cited by examiner]
US 20230245651A1 · Wang · 2023 [cited by applicant]
US 20240185968A1 · Doerr et al. · 2024 [cited by applicant]
Anonymous, Error Handling and Fallback Mechanisms in AI Assistants: A Comprehensive Guide, Nexus Flow Innovations, published on Nov. 11, 2024, 4 pages. Retrieved from the internet on Aug. 8, 2025. URL: https://www.nexus… [cited by applicant]
Biao et al, Root Mean Square Layer Normalization, School of Informatics, University of Edinburgh, Retrieved on Jun. 6, 2024, 12 pages. [NeurIPS 2019]. [cited by applicant]
Dosovitskiy et al., An Image Is Worth 16×16 Words: Transformers For Image Recognition at Scale, ICLR 2021, dated Jun. 3, 2021, 22 pages. [cited by applicant]
Gao et al. Retrieval-Augmented Generation for Large Language Models: A Survey, Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University, Mar. 27, 2024, 21 pages. [Url: arXiv:2312.10997]. [cited by applicant]
Karan et al, Large Language Models Encode Clinical Knowledge, Google Research, Dec. 26, 2022, 44 pages. [arXiv:2212.13138]. [cited by applicant]
McGreevy et al., Clinical, Legal, and Ethical Aspects of Artificial Intelligence-Assisted Conversational Agents in Health Care, JAMA, vol. 324, No. 6, dated Aug. 11, 2020, 2 pages. [cited by applicant]
Peter et al, Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. The New England Journal of Medicine, Mar. 30, 2023, 7 pages [N Engl Med 388;13]. [cited by applicant]
PGR2025-00075—Mandatory Notices, filed on Aug. 25, 2025, 4 pages. [cited by applicant]
PGR2025-00075—Petition for Post-Grant Review Updated, filed on Aug. 12, 2025, 83 pages. [cited by applicant]
PGR2025-00075—Petition for Post-Grant Review Updated, filed on Aug. 27, 2025, 4 pages. [cited by applicant]
PGR2025-00075—Petition for Post-Grant Review Updated, filed on Aug. 27, 2025, 83 pages. [cited by applicant]
Surani et al., Understanding Privacy and Security Postures of Healthcare Chatbots, CHI-22, Apr. 30-May 6, 2022, 7 pages. [cited by applicant]
Tao et al, Towards Conversational Diagnostic AI, Google Research, Jan. 11, 2024, 46 pages. [arXiv:2401.05654]. [cited by applicant]
Vaswani et al., Attention Is All You Need, Advances in Neural Information Processing Systems 2017, dated Dec. 6, 2017, 15 pages. [cited by applicant]