IP Library Granted Patent US 12,198,433
Granted Patent B2
US 12,198,433 · App. 18/104,138 · Granted Jan 14, 2025

Searching within segmented communication session content

Inventors: Andrew Miller-Smith (Chicago, IL); Renjie Tao (Sunnyvale, CA); Ling Tsou (Los Angeles, CA)
Assignee: Zoom Video Communications, Inc.
G06V20/41G06V10/762G06V20/49G06V20/70G06V30/19
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,433
App. No.
18/104,138
Granted
Jan 14, 2025
Kind
B2
Abstract

Methods and systems provide for search results within segmented communication session content. In one embodiment, the system receives a transcript and video content of a communication session between participants, the transcript including timestamps for a number of utterances associated with speaking participants; processes the video content to extract textual content visible within the frames of the video content; segments frames of the video content into a number of contiguous topic segments; determines a title for each topic segment; assigns a category label for each topic segment; receives a request from a user to search for specified text within the video content; determines one or more titles or category labels for which a prediction of relatedness with the specified text is present; and presents content from at least one topic segment associated with the one or more titles or category labels for which a prediction of relatedness is present.

Claims (37)

1. A method, comprising:

receiving video content of a conversation between participants produced during a communication session;

performing video-based segmentation on the video content to classify a category label from a list of category labels for each video frame of the video content, the video content comprising a plurality of topic segments each associated with a different title and category label;

receiving, from a client device associated with a user, a request to search for specified text within the video content;

in response to receiving the request, determining one or more of the titles or category labels for which a prediction of relatedness with the specified text is present; and

presenting, to the client device, content from at least one topic segment associated with the one or more titles or category labels for which a prediction of relatedness is present.

2. The method of claim 1 , wherein each topic segment is associated with a starting timestamp, and wherein the content comprises playback of video of the communication session at a starting timestamp associated with the at least one topic segment associated with the titles or category labels for which a prediction of relatedness is present.

3. The method of claim 1 , wherein the title associated with each topic segment was determined via performing optical character recognition (OCR) on one or more frames within the topic segment.

4. The method of claim 1 , wherein the title associated with a topic segment is an empty or null title.

5. The method of claim 1 , wherein presenting the content from at least one topic segment comprises:

presenting one or more frames from the at least one topic segment associated with each title or category label for which a prediction of relatedness is present.

6. The method of claim 5 , wherein related titles and category labels are visually highlighted within any presented frame from the topic segment in which they appear.

7. The method of claim 6 , wherein determining that a prediction of relatedness with a specified text is present is based on one or more of: entity extraction techniques, relationship embedding techniques, and matching synonyms.

8. The method of claim 1 , wherein determining one or more titles or category labels for which a prediction of relatedness with the specified text is present comprises determining one or more exact matches between a title or category label and the specified text.

9. The method of claim 1 , wherein determining that a prediction of relatedness with a specified text is present comprises determining one or more exact matches between titles or category labels and a spell-corrected version of the specified text.

10. The method of claim 1 , wherein determining that a prediction of relatedness with a specified text is present comprises determining one or more non-exact matches with the specified text.

11. The method of claim 1 , wherein the content from the topic segments associated with titles or category labels for which a prediction of relatedness is present are presented to the client device in chronological order based on associated timestamps.

12. A communication system comprising one or more processors configured to perform operations of:

receiving video content of a conversation between participants produced during a communication session;

performing video-based segmentation on the video content to classify a category label from a list of category labels for each video frame of the video content, the video content comprising a plurality of topic segments each associated with a different title and category label;

receiving, from a client device associated with a user, a request to search for specified text within the video content;

in response to receiving the request, determining one or more of the titles or category labels for which a prediction of relatedness with the specified text is present; and

presenting, to the client device, content from at least one topic segment associated with the one or more titles or category labels for which a prediction of relatedness is present.

13. The communication system of claim 12 , wherein each topic segment is associated with a starting timestamp, and wherein the content comprises playback of video of the communication session at a starting timestamp associated with the at least one topic segment associated with the titles or category labels for which a prediction of relatedness is present.

14. The communication system of claim 12 , wherein the title associated with each topic segment was determined via performing optical character recognition (OCR) on one or more frames within the topic segment.

15. The communication system of claim 12 , wherein the title associated with a topic segment is an empty or null title.

16. The communication system of claim 12 , wherein presenting the content from at least one topic segment comprises:

presenting one or more frames from the at least one topic segment associated with each title or category label for which a prediction of relatedness is present.

17. The communication system of claim 16 , wherein related titles and category labels are visually highlighted within any presented frame from the topic segment in which they appear.

18. The communication system of claim 12 , wherein determining one or more titles or category labels for which a prediction of relatedness with the specified text is present comprises determining one or more exact matches between a title or category label and the specified text.

19. The communication system of claim 12 , wherein determining that a prediction of relatedness with a specified text is present comprises determining one or more exact matches between titles or category labels and a spell-corrected version of the specified text.

20. A non-transitory computer-readable medium containing instructions comprising:

instructions for receiving video content of a conversation between participants produced during a communication session;

performing video-based segmentation on the video content to classify a category label from a list of category labels for each video frame of the video content, the video content comprising a plurality of topic segments each associated with a different title and category label;

instructions for receiving, from a client device associated with a user, a request to search for specified text within the video content;

in response to receiving the request, instructions for determining one or more of the titles or category labels for which a prediction of relatedness with the specified text is present; and

instructions for presenting, to the client device, content from at least one topic segment associated with the one or more titles or category labels for which a prediction of relatedness is present.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2023
From: MILLER-SMITH, ANDREW; TAO, RENJIE; TSOU, LING
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 063480/0735 →