IP Library Patent Application 19222552
Patent Application
App. No. 19/222,552

METHODS AND SYSTEMS FOR SEGMENTING VIDEO CONTENT BASED ON SPEECH DATA AND FOR RETREIVING VIDEO SEGMENTS TO GENERATE VIDEOS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/222,552
Abstract

A method includes receiving a series of video segments and providing the series of video segments as input to a first machine learning model to produce text data. The text data is provided as input to a second machine learning model to produce categorized text data that includes a classification indication. The classification indication is added to metadata of the video segment, and the categorized text data is provided as input to a third machine learning model to produce a semantic vector. The method also includes causing the video segment and the metadata that includes the classification indication to be stored at a location of a database based on the semantic vector, the database being configured to be searched based on a search query associated with the semantic vector.

Claims (67)

1 . A non-transitory, machine-readable medium storing instructions that, when executed by a processor, cause the processor to:

receive input data;

search, based on the input data, a plurality of semantic vectors associated with a plurality of video segments;

in response to determining an association between the input data and at least one semantic vector from the plurality of semantic vectors, select a semantic vector that is (1) from the at least one semantic vector and (2) associated with a video segment from the plurality of video segments, based on a comparison between the input data and metadata that is associated with the video segment and includes a classification indication; and

add the video segment to a series of video segments based on the classification indication.

2 . The non-transitory, machine-readable medium of claim 1 , wherein the plurality of semantic vectors is a first plurality of semantic vectors, the plurality of video segments is a plurality of verbal video segments, and the video segment is a verbal video segment, the non-transitory, machine-readable medium further storing instructions to cause the processor to:

in response to determining an absence of an association between the input data and the first plurality of semantic vectors, search a second plurality of semantic vectors based on the input data to identify a nonverbal video segment from a plurality of nonverbal video segments, the second plurality of semantic vectors being associated with the plurality of nonverbal video segments;

provide the nonverbal video segment as input to a first machine learning model to produce storyline data;

provide the nonverbal video segment as input to a second machine learning model to produce audio data;

include the nonverbal video segment in the series of video segments that includes the verbal video segment to produce an updated series of video segments; and

generate video data based on the updated series of video segments, the storyline data, and the audio data.

3 . The non-transitory, machine-readable medium of claim 1 , wherein the instructions to cause the processor to select the semantic vector from the at least one semantic vector include instructions to cause the processor to provide the at least one semantic vector and the input data as input to a machine learning model to select the semantic vector.

4 . The non-transitory, machine-readable medium of claim 1 , further storing instructions to cause the processor to:

receive at least one of a text prompt or an image prompt; and

provide the at least one of the text prompt or the image prompt as input to at least one machine learning model to produce the input data.

5 . The non-transitory, machine-readable medium of claim 1 , wherein the instructions cause the processor to search the plurality of semantic vectors include instructions to cause the processor to determine at least one cosine similarity value based on the input data and the plurality of semantic vectors.

6 . The non-transitory, machine-readable medium of claim 1 , wherein:

the metadata further includes at least one of an orientation indication, a resolution indication, a video segment length indication, or a frame rate indication; and

the instructions to cause the processor to select the semantic vector include instructions to cause the processor to select the semantic vector based on a comparison between the input data and the at least one of the orientation indication, the resolution indication, the video segment length indication, or the frame rate indication.

7 . The non-transitory, machine-readable medium of claim 1 , wherein the video segment is a first video segment, the non-transitory, machine-readable medium further storing instructions to cause the processor to:

update the input data based on the video segment to produce updated input data; and

search, based on the updated input data, the plurality of semantic vectors to select a second video segment.

8 . The non-transitory, machine-readable medium of claim 1 , wherein the video segment is a first video segment, the non-transitory, machine-readable medium further storing instructions to cause the processor to:

cause display of the series of video segments via a graphical user interface (GUI) of a user compute device;

receive an indication of a second video segment from the user compute device in response to causing the display of the series of video segments; and

include the second video segment in the series of video segments.

9 . A method, comprising:

receiving, at a processor, input data;

searching, via the processor and based on the input data, a plurality of semantic vectors associated with a plurality of video segments;

in response to determining an association between the input data and at least one semantic vector from the plurality of semantic vectors, selecting, via the processor, a semantic vector that is (1) from the at least one semantic vector and (2) associated with a video segment from the plurality of video segments, based on a comparison between the input data and metadata that is associated with the video segment and includes a classification indication; and

adding, via the processor, the video segment to a series of video segments based on the classification indication.

10 . The method of claim 9 , wherein the plurality of semantic vectors is a first plurality of semantic vectors, the plurality of video segments is a plurality of verbal video segments, and the video segment is a verbal video segment, the method further comprising:

in response to determining an absence of an association between the input data and the first plurality of semantic vectors, searching, via the processor, a second plurality of semantic vectors based on the input data to identify a nonverbal video segment from a plurality of nonverbal video segments, the second plurality of semantic vectors being associated with the plurality of nonverbal video segments;

providing, via the processor, the nonverbal video segment as input to a first machine learning model to produce storyline data;

providing, via the processor, the nonverbal video segment as input to a second machine learning model to produce audio data;

including, via the processor, the nonverbal video segment in the series of video segments that includes the verbal video segment to produce an updated series of video segments; and

generating, via the processor, video data based on the updated series of video segments, the storyline data, and the audio data.

11 . The method of claim 9 , wherein the selecting the semantic vector from the at least one semantic vector includes providing, via the processor, the at least one semantic vector and the input data as input to a machine learning model to select the semantic vector.

12 . The method of claim 9 , further comprising:

receiving, at the processor, at least one of a text prompt or an image prompt; and

providing, via the processor, the at least one of the text prompt or the image prompt as input to at least one machine learning model to produce the input data.

13 . The method of claim 9 , wherein the searching the plurality of semantic vectors includes determining, via the processor, at least one cosine similarity value based on the input data and the plurality of semantic vectors.

14 . The method of claim 9 , wherein:

the metadata further includes at least one of an orientation indication, a resolution indication, a video segment length indication, or a frame rate indication; and

the selecting the semantic vector includes selecting, via the processor, the semantic vector based on a comparison between the input data and the at least one of the orientation indication, the resolution indication, the video segment length indication, or the frame rate indication.

15 . The method of claim 9 , wherein the video segment is a first video segment, the method further comprising:

updating, via the processor, the input data based on the video segment to produce updated input data; and

searching, via the processor and based on the updated input data, the plurality of semantic vectors to select a second video segment.

16 . The method of claim 9 , wherein the video segment is a first video segment, the method further comprising:

causing, via the processor, display of the series of video segments via a graphical user interface (GUI) of a user compute device;

receiving, at the processor, an indication of a second video segment from the user compute device in response to causing the display of the series of video segments; and

including, via the processor, the second video segment in the series of video segments.

17 . An apparatus, comprising:

a memory; and

a processor operatively coupled to the memory, the processor configured to:

receive input data;

search, based on the input data, a plurality of semantic vectors associated with a plurality of video segments;

in response to determining an association between the input data and at least one semantic vector from the plurality of semantic vectors, select a semantic vector that is (1) from the at least one semantic vector and (2) associated with a video segment from the plurality of video segments, based on a comparison between the input data and metadata that is associated with the video segment and includes a classification indication; and

add the video segment to a series of video segments based on the classification indication.

18 . The apparatus of claim 17 , wherein the plurality of semantic vectors is a first plurality of semantic vectors, the plurality of video segments is a plurality of verbal video segments, and the video segment is a verbal video segment, the processor further configured to:

in response to determining an absence of an association between the input data and the first plurality of semantic vectors, search a second plurality of semantic vectors based on the input data to identify a nonverbal video segment from a plurality of nonverbal video segments, the second plurality of semantic vectors being associated with the plurality of nonverbal video segments;

provide the nonverbal video segment as input to a first machine learning model to produce storyline data;

provide the nonverbal video segment as input to a second machine learning model to produce audio data;

include the nonverbal video segment in the series of video segments that includes the verbal video segment to produce an updated series of video segments; and

generate video data based on the updated series of video segments, the storyline data, and the audio data.

19 . The apparatus of claim 17 , wherein the metadata further includes at least one of an orientation indication, a resolution indication, a video segment length indication, or a frame rate indication, the processor being configured to select the semantic vector based on a comparison between the input data and the at least one of the orientation indication, the resolution indication, the video segment length indication, or the frame rate indication.

20 . The apparatus of claim 17 , wherein the processor is configured to provide the at least one semantic vector and the input data as input to a machine learning model to select the semantic vector.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2026
From: SANGHAVI, SUNDEEP; YU, JIN; EDWARDS, GROWSON; SHAH, HARSHIL
To: CIPIO INC.
Reel/Frame 076007/0788 →
CHANGE OF NAME Recorded May 20, 2026
From: CIPIO, INC.
To: VIDEOFORCEAI, INC.
Reel/Frame 075583/0687 →