IP Library Granted Patent US 11,232,794
Granted Patent B2
US 11,232,794 · App. 17/315,857 · Granted Jan 25, 2022

System and method for multi-microphone automated clinical documentation

Inventors: Dushyant Sharma (Woburn, MA); Patrick A. Naylor (Reading, GB)
Assignee: NUANCE COMMUNICATIONS, INC.
G10L15/22G06F16/65G06F16/686G06N20/00G10L15/20G10L15/32G10L17/06G10L21/028G10L25/78G10L25/84G16H15/00H04R1/406H04R3/005H04R3/04H04R5/04H04R29/005H04S7/307G10L15/26G10L21/0216G10L21/0272G10L2021/02166G16H10/60G16H40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,232,794
App. No.
17/315,857
Granted
Jan 25, 2022
Kind
B2
Abstract

A method, computer program product, and computing system for receiving audio encounter information from a microphone array. Speech activity within one or more portions of the audio encounter information may be identified based upon, at least in part, a correlation among the audio encounter information received from the microphone array. Location information for the one or more portions of the audio encounter information may be determined based upon, at least in part, the correlation among the signals received by each microphone of the microphone array. The one or more portions of the audio encounter information may be labeled with the speech activity and the location information.

Claims (33)

1. A computer-implemented method, executed on a computing device, comprising:

receiving audio encounter information from a microphone array;

identifying speech activity within each of one or more portions of the audio encounter information based upon, at least in part, a correlation among the audio encounter information received from the microphone array;

determining location information for the one or more portions of the audio encounter information based upon, at least in part, the correlation among the audio encounter information received from the microphone array; and

labeling, for each of the one or more portions of the audio encounter information, the portion with the speech activity and the location information.

2. The computer-implemented method of claim 1 , wherein determining the location information for the one or more portions of the audio encounter information includes determining a time difference of arrival between each pair of microphones of the microphone array for the one or more portions of the audio encounter information.

3. The computer-implemented method of claim 1 , wherein identifying the speech activity and determining the location information are performed jointly for the one or more portions of the audio encounter information.

4. The computer-implemented method of claim 3 , wherein identifying the speech activity and determining the location information are performed jointly for the one or more portions of the audio encounter information using a machine learning model.

5. The computer-implemented method of claim 3 , further comprising:

receiving information associated with an acoustic environment.

6. The computer-implemented method of claim 5 , wherein identifying the speech activity within each of the one or more portions of the audio encounter information is based upon, at least in part, the location information for the one or more portions of the audio encounter information and the information associated with the acoustic environment.

7. The computer-implemented method of claim 1 , wherein labeling, for each of the one or more portions of the audio encounter information, the portion with the speech activity and the location information includes generating acoustic metadata for the one or more portions of the audio encounter information.

8. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

receiving audio encounter information from a microphone array;

identifying speech activity within each of one or more portions of the audio encounter information based upon, at least in part, a correlation among the audio encounter information received from the microphone array;

determining location information for the one or more portions of the audio encounter information based upon, at least in part, the correlation among the audio encounter information received from the microphone array; and

labeling for each of the one or more portions of the audio encounter information, the portion with the speech activity and the location information.

9. The computer program product of claim 8 , wherein determining the location information for the one or more portions of the audio encounter information includes determining a time difference of arrival between each pair of microphones of the microphone array for the one or more portions of the audio encounter information.

10. The computer program product of claim 8 , wherein identifying the speech activity and determining the location information are performed jointly for the one or more portions of the audio encounter information.

11. The computer program product of claim 10 , wherein identifying the speech activity and determining the location information are performed jointly for the one or more portions of the audio encounter information using a machine learning model.

12. The computer program product of claim 10 , wherein the operations further comprise:

receiving information associated with an acoustic environment.

13. The computer program product of claim 12 , wherein identifying the speech activity within each of the one or more portions of the audio encounter information is based upon, at least in part, the location information for the one or more portions of the audio encounter information and the information associated with the acoustic environment.

14. The computer program product of claim 8 , wherein labeling, for each of the one or more portions of the audio encounter information, the portion with the speech activity and the location information includes generating acoustic metadata for the one or more portions of the audio encounter information.

15. A computing system comprising:

a memory; and

a processor configured to receive audio encounter information from a microphone array, wherein the processor is further configured to identify speech activity within each of one or more portions of the audio encounter information based upon, at least in part, a correlation among the audio encounter information received from the microphone array, wherein the processor is further configured to location information for the one or more portions of the audio encounter information based upon, at least in part, the correlation among the audio encounter information received from the microphone array, and wherein the processor is further configured to label, for each of the one or more portions of the audio encounter information, the portion with the speech activity and the location information.

16. The computing system of claim 15 , wherein determining the location information for the one or more portions of the audio encounter information includes determining a time difference of arrival between each pair of microphones of the microphone array for the one or more portions of the audio encounter information.

17. The computing system of claim 15 , wherein identifying the speech activity and determining the location information are performed jointly for the one or more portions of the audio encounter information.

18. The computing system of claim 17 , wherein identifying the speech activity and determining the location information are performed jointly for the one or more portions of the audio encounter information using a machine learning model.

19. The computing system of claim 17 , wherein the processor is further configured to:

receive information associated with an acoustic environment.

20. The computing system of claim 19 , wherein identifying the speech activity within each of the one or more portions of the audio encounter information is based upon, at least in part, the location information for the one or more portions of the audio encounter information and the information associated with the acoustic environment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2021
From: SHARMA, DUSHYANT; NAYLOR, PATRICK A.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 056187/0325 →