IP Library Granted Patent US 12671603
Granted Patent B1
US 12671603 · App. 18/202,544 · Granted Jun 30, 2026

Generating summaries of videoconferences

Inventors: Ronald Hutto (Willis, TX); Gavin Paxton (Magnolia, TX)
Assignee: Zoom Communications, Inc.
H04L12/1831G06V20/47H04L12/1822
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12671603
App. No.
18/202,544
Granted
Jun 30, 2026
Kind
B1
Abstract

Techniques for automatically generating meeting summaries of video conferences and virtual meetings are disclosed. In an example, a method involves receiving, from a virtual meeting, one or more items of content. The method further involves providing the one or more items of content to a machine-learning model. The method further involves receiving a summary of the virtual meeting from the machine-learning model. The method further involves transmitting the summary to one or more client devices.

Claims (56)

1 . A method comprising:

receiving, from a virtual meeting, one or more items of meeting content;

accessing a data store having a plurality of available multimedia content items;

identifying, based on metadata for the virtual meeting, one or more additional items of multimedia content relevant to the virtual meeting from the plurality of available multimedia content items;

providing the one or more items of meeting content and the one or more additional items of multimedia meeting content to a machine-learning model;

receiving, from the machine-learning model, a summary of the virtual meeting; and

transmitting the summary to one or more client devices of a plurality of client devices.

2 . The method of claim 1 , further comprising establishing the virtual meeting and joining a plurality of client devices to the virtual meeting, each client device associated with a corresponding participant.

3 . The method of claim 1 , wherein the meeting content includes audio, the method further comprising:

receiving, from a client device of the plurality of client devices, an audio stream;

extracting, from the audio stream, one or more keywords; and

providing the keywords to the machine-learning model with the meeting content.

4 . The method of claim 1 , wherein the meeting content includes one or more documents, the method further comprising:

identifying, in the one or more documents, textual content;

providing the textual content to the machine-learning model with the meeting content.

5 . The method of claim 1 , wherein the meeting content includes one or more messages shared between participants, one or more content items shared between participants, or one or more segments of audio shared between participants.

6 . The method of claim 1 , wherein the machine-learning model identifies one or more topics discussed in the virtual meeting, and wherein the summary comprises the one or more topics.

7 . The method of claim 1 , further comprising:

composing a message comprising the summary; and

providing the summary via a message to one or more client devices of the plurality of client devices.

8 . A system comprising:

a non-transitory computer-readable medium storing processor-executable program instructions; and

a processor communicatively coupled to the non-transitory computer-readable medium for executing the processor-executable program instructions, wherein executing the processor-executable program instructions configures the processor to:

receive, from a virtual meeting, one or more items of meeting content;

access a data store having a plurality of available multimedia content items;

identify, based on metadata for the virtual meeting, one or more additional items of multimedia content relevant to the virtual meeting from the plurality of available multimedia content items;

provide the one or more items of meeting content and the one or more additional items of multimedia meeting content to a machine-learning model;

receive, from the machine-learning model, a summary of the virtual meeting; and

transmit the summary to one or more client devices of a plurality of client devices.

9 . The system of claim 8 , wherein executing the processor-executable program instructions configure the processor to: establish the virtual meeting and joining a plurality of client devices to the virtual meeting, each client device associated with a corresponding participant.

10 . The system of claim 8 , wherein the meeting content includes audio, wherein executing the processor-executable program instructions configures the processor to:

receive, from a client device of the plurality of client devices, an audio stream;

extract, from the audio stream, one or more keywords; and

provide the keywords to the machine-learning model with the meeting content.

11 . The system of claim 8 , wherein the meeting content includes one or more documents, wherein executing the processor-executable program instructions configures the processor to:

identify, in the one or more documents, textual content; and

provide the textual content to the machine-learning model with the meeting content.

12 . The system of claim 8 , wherein the meeting content includes one or more messages shared between participants, one or more content items shared between participants, or one or more segments of audio shared between participants.

13 . The system of claim 8 , wherein the machine-learning model identifies one or more topics discussed in the virtual meeting, and wherein the summary comprises the one or more topics.

14 . The system of claim 8 , wherein executing the processor-executable program instructions configures the processor to:

compose a message comprising the summary; and

provide the summary via a message to one or more client devices of the plurality of client devices.

15 . A non-transitory computer-readable medium comprising processor-executable instructions, wherein when executed by a processing device, the processor-executable program instructions cause the processing device to:

receive, from a virtual meeting, one or more items of meeting content;

access a data store having a plurality of available multimedia content items;

identify, based on metadata for the virtual meeting, one or more additional items of multimedia content relevant to the virtual meeting from the plurality of available multimedia content items;

provide the one or more items of meeting content and the one or more additional items of multimedia meeting content to a machine-learning model;

receive, from the machine-learning model, a summary of the virtual meeting; and

transmit the summary to one or more client devices of a plurality of client devices.

16 . The non-transitory computer-readable medium of claim 15 , further comprising processor-executable program instructions configured to cause the processing device to establish the virtual meeting and joining a plurality of client devices to the virtual meeting, each client device associated with a corresponding participant.

17 . The non-transitory computer-readable medium of claim 15 , further comprising processor-executable program instructions configured to cause the processing device to end the virtual meeting prior to transmitting the summary.

18 . The non-transitory computer-readable medium of claim 15 , wherein the content includes audio, further comprising processor-executable program instructions configured to cause the processing device to:

receive, from a client device of the plurality of client devices, an audio stream;

extract, from the audio stream, one or more keywords; and

provide the keywords to the machine-learning model with the content.

19 . The method of claim 1 , wherein the one or more additional items of multimedia content relevant to the virtual meeting comprise content shared before or after the virtual meeting.