IP Library Granted Patent US 10,839,807
Granted Patent B2
US 10,839,807 · App. 16/732,291 · Granted Nov 17, 2020

Systems and methods for voice identification and analysis

Inventors: Timothy Degraye (Geneva, CH); Liliane Huguet (Geneva, CH)
Assignee: HED Technologies Sarl
G10L15/265G10L15/285G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,839,807
App. No.
16/732,291
Granted
Nov 17, 2020
Kind
B2
Abstract

Obtaining configuration audio data including voice information for a plurality of meeting participants. Generating localization information indicating a respective location for each meeting participant. Generating a respective voiceprint for each meeting participant. Obtaining meeting audio data. Identifying a first meeting participant and a second meeting participant. Linking a first meeting participant identifier of the first meeting participant with a first segment of the meeting audio data. Linking a second meeting participant identifier of the second meeting participant with a second segment of the meeting audio data. Generating a GUI indicating the respective locations of the first and second meeting participants, and the GUI indicating a first transcription of the first segment and a second transcription of the second segment. The first transcription is associated with the first meeting participant in the GUI, and the second transcription is associated with the second meeting participant in the GUI.

Claims (64)

1. A system comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the system to perform:

obtaining configuration audio data including voice information from a plurality of meeting participants in a meeting room, the voice information captured by one or more microphones of a plurality of microphones in the meeting room, each of the plurality of microphones having a respective position relative to each other, the voice information from each participant of the plurality of participants including a respective participant identifier;

generating localization information based on the configuration audio data and on the respective positions of the plurality of microphones, the localization information indicating a respective location of each participant of the plurality of meeting participants in the meeting room, the respective location of each participant being associated with the respective participant identifier;

generating, based on the configuration audio data and the localization information, a respective voiceprint for each of the plurality of meeting participants, the respective voiceprint being associated with the respective participant identifier;

at least during a first time period:

obtaining first meeting audio data;

identifying, based on the localization information and the respective voiceprints, a first segment of the first meeting audio data as associated with a first meeting participant of the plurality of meeting participants and a second segment of the first meeting audio data as associated with a second meeting participant of the plurality of meeting participants;

linking a first meeting participant identifier of the first meeting participant with the first segment of the first meeting audio data; and

linking a second meeting participant identifier of the second meeting participant with the second segment of the first meeting audio data;

updating the respective voiceprint of the first meeting participant based on the first segment of the first meeting audio data;

updating the respective voiceprint of the second meeting participant based on the second segment of the first meeting audio data; and

generating a holistic graphical user interface (GUI), the holistic GUI illustrating the respective locations of the first and second meeting participants of the plurality of meeting participants in the meeting room, and the holistic GUI indicating a first transcription of the first segment of the first meeting audio data and a second transcription of the second segment of the first meeting audio data, the first transcription being associated with the first meeting participant in the holistic GUI, and the second transcription being associated with the second meeting participant in the holistic GUI.

2. The system of claim 1 , wherein the instructions further cause the system to perform:

at least during a second time period subsequent to the first time period:

obtaining additional meeting audio data;

identifying, based on the respective voiceprints and on a reduced weight of localization information, at least the first meeting participant and the second meeting participant of the plurality of meeting participants;

linking, based on the identification of the first meeting participant, the first meeting participant identifier of the first meeting participant with a first segment of the additional meeting audio data; and

linking, based on the identification of the second meeting participant, the second meeting participant identifier of the second meeting participant with a second segment of the additional meeting audio data.

3. The system of claim 1 , wherein the instructions further cause the system to perform:

receiving user feedback associated with the linking the first meeting participant identifier of the first meeting participant with the first segment of the additional meeting audio data;

unlinking, based on the user feedback, the first meeting participant identifier of the first meeting participant with the first segment of the additional meeting audio data; and

updating, based on the unlinking, the respective voiceprint of the first meeting participant.

4. The system of claim 3 , wherein the instructions further cause the system to perform:

linking, based on additional user feedback, a third meeting participant identifier of a third meeting participant of the plurality of meeting participants with the first segment of the first meeting audio data; and

updating, based on the linking of the third meeting participant identifier with the first segment of the first meeting audio data, the respective voiceprint of the third meeting participant.

5. The system of claim 1 , wherein the holistic GUI indicates a third transcription of the first segment of the first meeting audio data and a fourth transcription of the second segment of the first meeting audio data, the third transcription being associated with the first meeting participant in the holistic GUI, and the fourth transcription being associated with the second meeting participant in the holistic GUI.

6. The system of claim 1 , wherein the holistic GUI indicates a first voice recording of the first segment of the first meeting audio data and a second voice recording of the second segment of the first meeting audio data, the first voice recording being associated with the first meeting participant in the holistic GUI, and the second voice recording being associated with the second meeting participant in the holistic GUI.

7. The system of claim 6 , wherein each of the first and second voice recordings may be played back within the holistic GUI responsive to user input.

8. The system of claim 1 , wherein a first set of microphones of the plurality of microphones is disposed in a first directional audio recording device, and a second set of microphones of the plurality of microphones is disposed in a second directional audio recording device distinct and remote from the first directional audio recording device.

9. The system of claim 8 , wherein a first segment of the first meeting audio data is captured by the first directional audio recording device, and the second segment of the first meeting audio data is captured by the second directional audio recording device.

10. The system of claim 1 , wherein the voice information includes voice audio data and signal strength data associated with the voice audio data.

11. A method being implemented by a computing system including one or more physical processors and storage media storing machine-readable instructions, the method comprising:

obtaining configuration audio data including voice information from a plurality of meeting participants in a meeting room, the voice information captured by one or more microphones of a plurality of microphones in the meeting room, each of the plurality of microphones having a respective position relative to each other, the voice information from each participant of the plurality of participants including a respective participant identifier;

generating localization information based on the configuration audio data and on the respective positions of the plurality of microphones, the localization information indicating a respective location of each participant of the plurality of meeting participants in the meeting room, the respective location of each participant being associated with the respective participant identifier;

generating, based on the configuration audio data and the localization information, a respective voiceprint for each of the plurality of meeting participants, the respective voiceprint being associated with the respective participant identifier;

at least during a first time period:

obtaining first meeting audio data;

identifying, based on the localization information and the respective voiceprints, a first segment of the first meeting audio data as associated with a first meeting participant of the plurality of meeting participants and a second segment of the first meeting audio data as associated with a second meeting participant of the plurality of meeting participants;

linking a first meeting participant identifier of the first meeting participant with the first segment of the first meeting audio data; and

linking a second meeting participant identifier of the second meeting participant with the second segment of the first meeting audio data;

updating the respective voiceprint of the first meeting participant based on the first segment of the first meeting audio data;

updating the respective voiceprint of the second meeting participant based on the second segment of the first meeting audio data; and

generating a holistic graphical user interface (GUI), the holistic GUI illustrating the respective locations of the first and second meeting participants of the plurality of meeting participants in the meeting room, and the holistic GUI indicating a first transcription of the first segment of the first meeting audio data and a second transcription of the second segment of the first meeting audio data, the first transcription being associated with the first meeting participant in the holistic GUI, and the second transcription being associated with the second meeting participant in the holistic GUI.

12. The method of claim 11 , further comprising:

at least during a second time period subsequent to the first time period:

obtaining additional meeting audio data;

identifying, based on the respective voiceprints and on a reduced weight of localization information, at least the first meeting participant and the second meeting participant of the plurality of meeting participants;

linking, based on the identification of the first meeting participant, the first meeting participant identifier of the first meeting participant with a first segment of the additional meeting audio data; and

linking, based on the identification of the second meeting participant, the second meeting participant identifier of the second meeting participant with a second segment of the additional meeting audio data.

13. The method of claim 11 , further comprising:

receiving user feedback associated with the linking the first meeting participant identifier of the first meeting participant with the first segment of the additional meeting audio data;

unlinking, based on the user feedback, the first meeting participant identifier of the first meeting participant with the first segment of the additional meeting audio data; and

updating, based on the unlinking, the respective voiceprint of the first meeting participant.

14. The method of claim 13 , further comprising:

linking, based on additional user feedback, a third meeting participant identifier of a third meeting participant of the plurality of meeting participants with the first segment of the first meeting audio data; and

updating, based on the linking of the third meeting participant identifier with the first segment of the first meeting audio data, the respective voiceprint of the third meeting participant.

15. The method of claim 11 , wherein the holistic GUI indicates a third transcription of the first segment of the first meeting audio data and a fourth transcription of the second segment of the first meeting audio data, the third transcription being associated with the first meeting participant in the holistic GUI, and the fourth transcription being associated with the second meeting participant in the holistic GUI.

16. The method of claim 11 , wherein the holistic GUI indicates a first voice recording of the first segment of the first meeting audio data and a second voice recording of the second segment of the first meeting audio data, the first voice recording being associated with the first meeting participant in the holistic GUI, and the second voice recording being associated with the second meeting participant in the holistic GUI.

17. The method of claim 16 , wherein each of the first and second voice recordings may be played back within the holistic GUI responsive to user input.

18. The method of claim 11 , wherein a first set of microphones of the plurality of microphones is disposed in a first directional audio recording device, and a second set of microphones of the plurality of microphones is disposed in a second directional audio recording device distinct and remote from the first directional audio recording device.

19. The method of claim 18 , wherein a first segment of the first meeting audio data is captured by the first directional audio recording device, and the second segment of the first meeting audio data is captured by the second directional audio recording device.

20. The method of claim 11 , wherein the voice information includes voice audio data and signal strength data associated with the voice audio data.

Assignments (3)
SECURITY INTEREST Recorded Jan 31, 2023
From: HED TECHNOLOGIES SÀRL
To: CELLO HOLDINGS LIMITED
Reel/Frame 062542/0920 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CORRECT SECOND ASSIGNOR NAME PREVIOUSLY RECORDED AT REEL: 053958 FRAME: 0948. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 5, 2020
From: DEGRAYE, TIMOTHY; HUGUET, LILIANE
To: HED TECHNOLOGIES SARL
Reel/Frame 053981/0158 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2020
From: DEGRAYE, TIMOTHY; HUGUET, LILIAN
To: HED TECHNOLOGIES SARL
Reel/Frame 053958/0948 →
Continuity (2)
Provisional Application 62786915 · Dec 31, 2018
Related Publication 20200211561A1 · Jul 2, 2020