IP Library › Granted Patent US 11,978,456
Granted Patent B2
US 11,978,456 · App. 17/651,208 · Granted May 7, 2024

System, method and programmed product for uniquely identifying participants in a recorded streaming teleconference

Inventors: Shlomi Medalion (Ramat Gan, IL); Omri Allouche (Tel Aviv, IL); Maxim Bulanov (Ramat Gan, IL)
Assignee: GONG.IO LTD
G10L17/06G06V10/82G06V20/40G06V40/171G10L17/02G10L17/18G10L21/028G10L25/57H04L65/403G06Q30/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,978,456
App. No.
17/651,208
Granted
May 7, 2024
Kind
B2
Abstract

Systems, methods and programmed products for using visual information in a video stream of a recording streaming teleconference among a plurality of participants to diarize speech, involving obtaining respective components of the teleconference including a respective audio component, a respective video component, respective teleconference metadata, and transcription data, parsing components into speech segments, tagging speech segments with source feeds, and diarizing the teleconference so as to label the speech segments based on neural network or heuristic analysis of visual information.

Claims (42)

1. A method for using visual information in a video stream of a first recorded teleconference among a plurality of participants to diarize speech, the method comprising:

(a) obtaining, by a computer system, components of the first recorded teleconference among the plurality of participants conducted over a network, wherein the respective components include:

(i) an audio component including utterances of respective participants that spoke during the first recorded teleconference;

(ii) a video component including a video feed as to respective participants that spoke during the first recorded teleconference;

(iii) teleconference metadata associated with the first recorded teleconference and including a first plurality of timestamp information and respective speaker identification information associated with each respective timestamp information; and

(iv) transcription data associated with the first recorded teleconference, wherein said transcription data is indexed by timestamps;

(b) parsing, by the computer system, the audio component into a plurality of speech segments in which one or more participants were speaking during the first recorded teleconference, wherein each respective speech segment is associated with a respective time segment including a start timestamp indicating a first time in the first recorded teleconference when the respective speech segment begins, and a stop timestamp associated with a second time in the first recorded teleconference when the respective speech segment ends;

(c) tagging, by the computer system, each respective speech segment with the respective speaker identification information based on the teleconference metadata associated with the respective time segment; and

(d) diarizing, by the computer system, the first recorded teleconference in a process comprising:

(i) indexing the transcription data in accordance with respective speech segments and the respective speaker identification information to generate a segmented transcription data set for the first recorded teleconference;

(ii) identifying respective speaker information associated with respective speech segments using a neural network with at least a portion of the video feed corresponding in time to at least a portion of the segmented transcription data set determined according to the indexing as an input, and providing source indication information for each respective speech segment as an output and using a training set including visual content tagged with prior source indication information, wherein the portion of the video feed includes a first artificial visual representation not including a face generated by telephone conferencing software in the visual content associated with a first participant that spoke during a first speech segment of the first recorded teleconference, and the portion of the video feed does not include any artificial visual representation associated with a second participant that did not speak during the first speech segment of the recorded teleconference, and the source indication information is based at least on presence of the first artificial visual representation; and

(iii) labeling each respective speech segment based on the identified respective speaker information associated with the respective speech segment;

wherein the identified respective speaker information is based on the source indication information.

2. The method of claim 1 , wherein at least some of the visual content shows lips in the process of speaking and at least some other of the visual content shows lips not in the process of speaking and the source indication information includes an indication of whether lips are moving.

3. The method of claim 1 , wherein the artificial visual representation is a colored shape appearing around a designated portion of a screen.

4. The method of claim 1 , wherein the artificial visual representation is predesignated text.

5. The method of claim 1 , wherein the step of identifying respective speaker information is further based on a look-up, by the computer system, of the output source indication information in a database containing speaker identification information associated with a plurality of potential speakers.

6. The method of claim 5 , wherein the look-up, by the computer system, is performed using a customer relationship management system.

7. The method of claim 1 , wherein the step of identifying respective speaker information further uses a second neural network with the at least a portion of the video feed corresponding in time to the at least a portion of the segmented transcription data set determined according to the indexing as an input, and second source indication information as an output and a second training set including second visual content tagged with prior source indication information.

8. The method of claim 7 , wherein at least some of the visual content shows lips and at least some other of the visual content shows an absence of lips and the source indication information includes an indication of whether lips are present.

9. The method of claim 8 , wherein at least some of the second visual content shows lips in the process of speaking and at least some other of the second visual content shows lips not in the process of speaking and the second source indication information includes an indication of whether lips are speaking, and wherein the identifying of respective speaker information selectively occurs accordingly to whether both the source indication information as outputted indicates lips being present and the second source indication information as outputted indicates lips are speaking.

10. The method of claim 1 , wherein at least some of the visual content shows lips in the process of pronouncing a first sound and at least some other of the visual content shows lips in the process of pronouncing a second sound and the source indication information includes an indication of a particular sound being pronounced.

11. The method of claim 1 , wherein the teleconference metadata is generated by the computer system.

12. The method of claim 1 , wherein the transcription data is generated by the computer system.

13. The method of claim 1 , wherein the respective speaker identification information associated with at least one of the respective timestamp information identifies multiple speakers among the plurality of participants.

14. The method of claim 13 , wherein the neural network is selectively used for the identifying the respective speaker information associated with respective speech segments according to whether respective speaker identification information of the teleconference metadata identifies multiple speakers among the plurality of participants.

15. The method of claim 1 , wherein identifying, by the computer system, the respective speaker information associated with respective speech segments further comprises performing optical character recognition on at least a portion of the video feed corresponding in time to at least a portion of the segmented transcription data set determined according to the indexing, so as to determine source-indicative characters or text.

16. The method of claim 1 , wherein identifying, by the computer system, the respective speaker information associated with respective speech segments further comprises performing symbol recognition on at least a portion of the video feed corresponding in time to at least a portion of the segmented transcription data set determined according to the indexing, so as to determine whether a source-indicative colored shape appears around a designated portion of a display associated with the video feed.

17. The method of claim 1 , further comprising performing an analysis, by the computer system, of the diarization of the first recorded teleconference, and providing, by the computer system, results of such analysis to a user.

18. The method of claim 17 , wherein the analysis, by the computer system, of the diarization of the first recorded teleconference, comprises a determination of conversation participant talk times, a determination of conversation participant talk ratios, a determination of conversation participant longest monologues, a determination of conversation participant longest uninterrupted speech segments, a determination of conversation participant interactivity, a determination of conversation participant patience, a determination of conversation participant question rates, or a determination of a topic duration.

19. A method for using video content of a video stream of a first recorded teleconference among a plurality of participants to diarize speech, the method comprising:

(a) obtaining, by a computer system, components of the first recorded teleconference among the plurality of participants conducted over a network, wherein the components include:

(i) an audio component including utterances of respective participants that spoke during the first recorded teleconference;

(ii) a video component including a video feed comprising video of respective participants that spoke during the first recorded teleconference;

(iii) teleconference metadata associated with the first recorded teleconference and including a first plurality of timestamp information and respective speaker identification information associated with each respective timestamp information; and

(iv) transcription data associated with the first recorded teleconference, wherein said transcription data is indexed by timestamps;

(b) parsing, by the computer system, the audio component into a plurality of speech segments in which one or more participants were speaking during the first recorded teleconference, wherein each respective speech segment is associated with a respective time segment including a start timestamp indicating a first time in the first recorded teleconference when the respective speech segment begins, and a stop timestamp associated with a second time in the first recorded teleconference when the respective speech segment ends;

(c) tagging, by the computer system, each respective speech segment with the respective speaker identification information based on the teleconference metadata associated with the respective time segment; and

(d) diarizing, by the computer system, the first recorded teleconference in a process comprising:

(i) indexing the transcription data in accordance with respective speech segments and the respective speaker identification information to generate a segmented transcription data set for the first recorded teleconference;

(ii) identifying respective spoken dialogue information associated with respective speech segments using a neural network with at least a portion of the video feed comprising video of at least one participant among the respective participants corresponding in time to at least a portion of the segmented transcription data set determined according to the indexing as an input, and spoken dialogue indication information as an output and a training set including a plurality of videos of persons tagged with indications of what spoken dialogue the respective persons are speaking, wherein the portion of the video feed includes a first artificial visual representation not including a face generated by telephone conferencing software in the visual content associated with a first participant that spoke during a first speech segment of the first recorded teleconference, and the portion of the video feed does not include any artificial visual representation associated with a second participant that did not speak during the first speech segment of the recorded teleconference, and the speaker identification information is based at least on presence of the first artificial visual representation; and,

(iii) updating the transcription data based on the identified respective spoken dialogue information associated with the respective speech segment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2024
From: MEDALION, SHLOMI; ALLOUCHE, OMRI; BULANOV, MAXIM
To: GONG.IO LTD
Reel/Frame 066464/0790 →
Continuity (1)
Related Publication 20230260519A1 · Aug 17, 2023
Cited By (1)
US 12,518,742