IP Library Granted Patent US 12701023
Granted Patent B1
US 12701023 · App. 18/461,172 · Granted Aug 4, 2026

Annotating a conference video stream with participant names using speech recognition

Inventors: Zhenghang Gu (San Jose, CA); Zhaofeng Jia (Saratoga, CA); Robert Aaron Klegon (Chicago, IL); Tiffany Hui Lai (Seattle, WA); Cynthia Eshiuan Lee (Austin, TX); Bo Ling (Saratoga, CA); Ka Ki Ng (San Mateo, CA); Jing-An Tzeng (San Jose, CA); Zhenyi Ye (Aliso Viejo, CA); Huixi Zhao (San Jose, CA)
Assignee: Zoom Communications, Inc.
H04L12/1822G06F21/32G10L17/04G06F2221/2117G06F2221/2143
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12701023
App. No.
18/461,172
Granted
Aug 4, 2026
Kind
B1
Abstract

A conference video stream is annotated with participant names determined using speech recognition. A computing device used with a video conference obtains voiceprints representing speech of individual conference participants invited to the video conference. The computing device obtains, as output of a speech recognition process performed based on the voiceprints, name information of a conference participant whose speech is represented within an audio segment captured during the video conference. The computing device outputs an annotation representing the name information to configure one or more participant devices connected to the video conference to display the name information proximate to a depiction of the conference participant within a video stream depicting the conference participant. The voiceprints representing the speech of the individual conference participants are uploaded to a data store prior to the video conference as part of an enrollment and authorization process performed by the individual conference participants.

Claims (55)

1 . A method, comprising:

obtaining, by a computing device used with a video conference, voiceprints representing speech of individual conference participants invited to the video conference, wherein the voiceprints are stored within a local memory of the computing device for use with the video conference;

obtaining, by the computing device and as output of a speech recognition process performed based on the voiceprints, name information of a conference participant whose speech is represented within an audio segment captured during the video conference;

outputting, by the computing device, an annotation representing the name information to configure one or more participant devices connected to the video conference to display, throughout the video conference independent of whether the conference participant is an active speaker, the name information proximate to a depiction of the conference participant within a video stream depicting the conference participant;

deleting, by the computing device based on an indication to end the video conference, the voiceprints from the local memory; and

ending, by the computing device, the video conference after the deletion of the voiceprints.

2 . The method of claim 1 , wherein outputting the annotation comprises:

transmitting, by the computing device, the annotation and information indicating a location within the video stream at which to display the annotation.

3 . The method of claim 1 , wherein outputting the annotation comprises:

including, by the computing device, the annotation within the video stream at a location proximate to the depiction of the conference participant; and

transmitting, by the computing device, the video stream including the annotation.

4 . The method of claim 1 , wherein obtaining the voiceprints comprises:

downloading, by the computing device from a data store, a voiceprint representing speech of the conference participant based on one or more of an identification of the conference participant within a participant list of the video conference or a pairing of a personal computing device of the conference participant to the computing device.

5 . The method of claim 1 , wherein obtaining the voiceprints comprises:

obtaining, by the computing device, a voiceprint representing speech of the conference participant from a personal computing device of the conference participant based on one or more of an identification of the conference participant within a participant list of the video conference or a pairing of the personal computing device to the computing device.

6 . The method of claim 1 , wherein obtaining the voiceprints comprises:

validating a quality of a voiceprint of the voiceprints based on one or more audio signal components of the voiceprint.

7 . The method of claim 1 , wherein obtaining the name information of the conference participant comprises:

performing the speech recognition process including determining a match between high dimensional features of a voiceprint representing speech of the conference participant and high dimensional features of the audio segment; and

obtaining the name information from a data store based on the match.

8 . The method of claim 1 , comprising:

storing, by the computing device, data associated with the voiceprints within the local memory; and

deleting, by the computing device, the data associated with the voiceprints from the local memory based on the ending of the video conference.

9 . The method of claim 1 , wherein outputting the annotation comprises:

outputting the annotation to indicate the conference participant as an active speaker during the video conference.

10 . The method of claim 1 , wherein outputting the annotation comprises:

removing the annotation when the conference participant is other than an active speaker during the video conference.

11 . The method of claim 1 , comprising:

recording the name information of the conference participant in connection with speech of the conference participant within a transcript of the video conference.

12 . The method of claim 1 , wherein the audio segment is captured using a microphone located within a physical space and the computing device is located within the physical space.

13 . A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:

obtaining, by a computing device used with a video conference, voiceprints representing speech of individual conference participants invited to the video conference, wherein the voiceprints are stored within a local memory of the computing device for use with the video conference;

obtaining, by the computing device and as output of a speech recognition process performed based on the voiceprints, name information of a conference participant whose speech is represented within an audio segment captured during the video conference;

outputting, by the computing device, an annotation representing the name information to configure one or more participant devices connected to the video conference to display, throughout the video conference independent of whether the conference participant is an active speaker, the name information proximate to a depiction of the conference participant within a conference video stream depicting the conference participant;

deleting, by the computing device based on an indication to end the video conference, the voiceprints from the local memory; and

ending, by the computing device, the video conference after the deletion of the voiceprints.

14 . The non-transitory computer readable medium of claim 13 , the operations comprising:

performing the speech recognition process by comparing a first set of high dimensional features determined for a voiceprint representing speech of the conference participant and a second set of high dimensional features determined for the audio segment.

15 . The non-transitory computer readable medium of claim 13 , the operations comprising:

prompting for input to confirm the name information prior to outputting the annotation.

16 . The non-transitory computer readable medium of claim 13 , wherein other information associated with the voiceprints is stored within the local memory of the computing device during the video conference and deleted from the local memory based on the ending of the video conference.

17 . A system, comprising:

a memory subsystem storing instructions; and

processing circuitry configured to execute the instructions to:

obtain voiceprints representing speech of individual conference participants invited to a video conference, wherein the voiceprints are stored within a local memory of the memory subsystem for use with the video conference;

obtain, as output of a speech recognition process performed based on the voiceprints, name information of a conference participant whose speech is represented within an audio segment captured during the video conference;

output an annotation representing the name information to configure one or more participant devices connected to the video conference to display, throughout the video conference independent of whether the conference participant is an active speaker, the name information proximate to a depiction of the conference participant within a conference video stream depicting the conference participant;

delete, based on an indication to end the video conference, the voiceprints from the local memory; and

end the video conference after the deletion of the voiceprints.

18 . The system of claim 17 , wherein the processing circuitry is configured to execute the instructions to:

determine location information describing a location within the video stream at which to cause the display of the annotation based on a location of the depiction of the conference participant within the video stream.

19 . The system of claim 17 , wherein, to obtain the voiceprints, the processing circuitry is configured to execute the instructions to:

determine that personal computing devices for the individual conference participants are paired to a computing device within a physical space within which the individual conference participants are located; and

determine that the individual conference participants are included in a participant invite list for the video conference.

20 . The system of claim 17 , wherein the voiceprints are uploaded for use during the video conference as part of an enrollment and authorization process performed by each of the individual conference participants.