IP Library Granted Patent US 11,403,676
Granted Patent B2
US 11,403,676 · App. 16/895,005 · Granted Aug 2, 2022

Interleaving video content in a multi-media document using keywords extracted from accompanying audio

Inventors: Jason S. Bayer (Mountain View, CA); Ronojoy Chakrabarti (Santa Clara, CA); Keval Desai (San Francisco, CA); Manish P Gupta (Santa Clara, CA); Jill A Huchital (Saratogo, CA); Willard V T Rusch, II (Woodside, CA)
Assignee: GOOGLE LLC
G06Q30/0277G06Q30/02G06Q30/0242H04N7/17318H04N21/23106H04N21/23424H04N21/252H04N21/25891H04N21/2665H04N21/2668H04N21/26208H04N21/26603H04N21/435H04N21/4331H04N21/44016H04N21/466H04N21/4667H04N21/812
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,403,676
App. No.
16/895,005
Granted
Aug 2, 2022
Kind
B2
Abstract

Provided herein are systems and methods of classifying video content. At least one server can identify a video content item identifying a plurality of segments to play primary video content. The at least one server can identify a set of words from a segment of the plurality of segments by using at least one of a transcription corresponding to the segment or speech recognition on audio content corresponding to the segment. The at least one server can determine a classification for the segment based on the set of words from the segment. The at least one server can store, in one or more data structures, an association between the video content item and the classification to categorize the segment of the video content item.

Claims (55)

1. A method, comprising:

identifying, by at least one server, a video content item identifying a plurality of segments to play primary video content;

identifying, by the at least one server, a set of words from a segment of the plurality of segments by using at least one of a transcript corresponding to the segment or speech recognition on audio content corresponding to the segment;

determining, by the at least one server, a classification for the segment based on the set of words from the segment;

storing, by the at least one server, in one or more data structures, an association between the video content item and the classification to categorize the segment of the video content item, wherein the classification indicates a topic or concept associated with the segment;

receiving, by the at least one server from a client device, a request for video content to insert into a content spot in the video content item, the content spot temporally adjacent to at least one of the plurality of segments;

responsive to the request, selecting, by the at least one server, from a plurality of supplemental video content items, a supplemental video content item based on the classification for the segment; and

providing, by the least one server to the client device, the supplemental video content item to play in the content spot in the video content item.

2. The method of claim 1 , further comprising:

identifying, by the at least one server, a second set of words across the plurality of segments identified by the video content item by using at least one of a second transcript corresponding to the video content item or speech recognition on audio content of the video content item;

determining, by the at least one server, a second classification for an entirety of the video content item based on the second set of words; and

storing, by the at least one server, in the one or more data structures, a second association between the video content item and the second classification to classify the entirety of the video content item.

3. The method of claim 1 , further comprising associating, by the at least one server, the classification with a timestamp defining the segment within the video content item; and

wherein storing the association between the video content item and the classification further comprises storing a second association between the classification and the timestamp.

4. The method of claim 1 , wherein determining the classification further comprises determining a plurality of topical categories based on the set of words from the segment; and

wherein storing the association between the video content item and the classification further comprises storing a second association between the video content item and the plurality of topical categories.

5. The method of claim 1 , wherein storing the association between the video content item and the classification further comprises storing the classification to be used to select one of a plurality of supplemental video content items to play in a content spot in the video content item, the content spot temporally adjacent to at least one of the plurality of supplemental video content items.

6. The method of claim 1 , further comprising:

identifying, by the at least one server, a temporal difference between the segment and a content spot identified by the video content item in which to play supplemental video content;

determining, by the at least one server, a weight for the classification to the content spot based on the temporal difference; and

wherein selecting the supplemental video content item further comprises selecting the supplemental video content item based on the weight.

7. The method of claim 1 , wherein determining the weight further comprises determining the weight different from a second weight for the classification to a second content spot based on a second temporal difference, the second temporal difference different from the temporal difference.

8. The method of claim 1 , further comprising filtering, by the at least one server, the plurality of supplemental video content items to identify a subset of supplemental video content items in accordance with a filtering policy for the content spot; and

wherein selecting the supplemental video content item further comprises selecting the supplemental video content item from the subset of supplemental video content item.

9. The method of claim 1 , wherein receiving the request further comprises receiving the request for video content including content selection parameters; and

wherein selecting the supplemental video content item further comprises selecting the supplemental video content item based on the content selection parameters.

10. A system for extracting topical categories from video content, comprising:

at least one server having one or more processors coupled with memory, configured to:

identify a video content item identifying a plurality of segments to play primary video content;

identify a set of words from a segment of the plurality of segments by using at least one of a transcript corresponding to the segment or speech recognition on audio content corresponding to the segment;

determine a classification for the segment based on the set of words from the segment;

store, in one or more data structures, an association between the video content item and the classification to categorize the segment of the video content item, wherein the classification indicates a topic or concept associated with the segment;

receive, from a client device, a request for video content to insert into a content spot in the video content item, the content spot temporally adjacent to at least one of the plurality of segments;

responsive to the request, select, from a plurality of supplemental video content items, a supplemental video content item based on the classification for the segment; and

provide, to the client device, the supplemental video content item to play in the content spot in the video content item.

11. The system of claim 10 , wherein the at least one server is further configured to:

identify a second set of words across the plurality of segments identified by the video content item by using at least of a second transcript corresponding to the video content item or speech recognition on audio content of the video content item;

determine a second classification for an entirety of the video content item based on the second set of words; and

store, in the one or more data structures, a second association between the video content item and the second classification to categorize the entirety of the video content item.

12. The system of claim 10 , wherein the at least one server is further configured to:

associate the classification with a timestamp defining the segment within the video content item; and

store, in the one or more data structures, a second association of the classification with the timestamp.

13. The system of claim 10 , wherein the at least one server is further configured to determine a plurality of topical categories based on the set of words from the segment and to store a second association between the video content item and the plurality of topical categories.

14. The system of claim 10 , wherein the at least one server is further configured to store the association between the video content item and the classification to be used to select one of a plurality of supplemental video content items to play in a content spot in the video content item, the content spot temporally adjacent to at least one of the plurality of supplemental video content items.

15. The system of claim 10 , wherein the at least one server is further configured to:

identify a temporal difference between the segment and a content spot identified by the video content item in which to play supplemental video content;

determine a weight for the classification to the content spot based on the temporal difference; and

select the supplemental video content item based on the weight.

16. The system of claim 10 , wherein the at least one server is further configured to determine the weight different from a second weight for the classification to a second content spot based on a second temporal difference, the second temporal difference different from the temporal difference.

17. The system of claim 10 , wherein the at least one server is further configured to:

filter the plurality of supplemental video content items to identify a subset of supplemental video content items in accordance with a filtering policy for the content spot; and

select the supplemental video content item from the subset of supplemental video content item.

18. The system of claim 10 , wherein the at least one server is further configured to:

receive the request for video content including content selection parameters; and

select the supplemental video content item based on the content selection parameters.

Assignments (3)
CHANGE OF NAME Recorded Oct 19, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 058538/0920 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2021
From: BALUJA, SHUMEET; ROWLEY, HENRY A.
To: GOOGLE INC.
Reel/Frame 057337/0293 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2021
From: BAYER, JASON; CHAKRABARTI, RONOJOY; DESAI, KEVAL; GUPTA, MANISH; HUCHITAL, JILL A.; RUSCH, WILLARD
To: GOOGLE INC.
Reel/Frame 057337/0728 →
Continuity (6)
Continuation 16140213 · Sep 24, 2018
Continuation 14251824 · Apr 14, 2014
Continuation 14246826 · Apr 7, 2014
Continuation 14193695 · Feb 28, 2014
Continuation 11323327 · Dec 30, 2005
Related Publication 20200302491A1 · Sep 24, 2020
Cited By (4)
US 12,189,637 US 12,361,380 US 12,639,318 US 12,664,170