IP Library Granted Patent US 10,679,261
Granted Patent B2
US 10,679,261 · App. 16/140,213 · Granted Jun 9, 2020

Interleaving video content in a multi-media document using keywords extracted from accompanying audio

Inventors: Jason S. Bayer (Mountain View, CA); Ronojoy Chakrabarti (Santa Clara, CA); Keval Desai (San Francisco, CA); Manish P Gupta (Santa Clara, CA); Jill A Huchital (Saratogo, CA); Willard V T Rusch, II (Woodside, CA)
Assignee: Google LLC
G06Q30/0277G06Q30/02G06Q30/0242H04N7/17318H04N21/23106H04N21/23424H04N21/252H04N21/25891H04N21/2665H04N21/2668H04N21/26208H04N21/26603H04N21/435H04N21/4331H04N21/44016H04N21/466H04N21/4667H04N21/812
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,679,261
App. No.
16/140,213
Granted
Jun 9, 2020
Kind
B2
Abstract

Provided herein are systems and methods of inserting content into videos based on associated text. A video content server can receive a request for video content into a video content slot of a video item played on the client device. The request can be generated responsive to execution of an encoding embedded in the video item. The video content server can identify words derived from a segment of the video item playable prior to the video content slot. The video content server can determine a topical category for the segment slot based on the identified words. The video content server can select a secondary video content item based on the topical category of the segment of the video item. The video content server can provide the secondary video content item to the client device to insert into the video content slot during the video item played on the client device.

Claims (53)

1. A method of inserting content into videos based on associated text, comprising:

receiving, by a video content server having one or more processors, from a client device, a request for video content to insert into a video content slot of a plurality of video content slots included in a video item playing on the client device, the request generated responsive to execution of an encoding embedded in the video item for the video content slot;

identifying, by the video content server, a set of words derived from a segment of a plurality of segments included in the video item, the segment playable prior to the video content slot included in the video item playing on the client device;

determining, by the video content server, a topical category for the segment playable prior to the video content slot based on the set of words derived from the segment of the video item playable prior to the video content slot;

selecting, by the video content server, from a plurality of candidate secondary video content items, responsive to receiving the request for video content, a secondary video content item based on the topical category of the segment of the video item; and

providing, by the video content server, the secondary video content item to the client device to insert into the video content slot playable subsequent to the segment of the plurality of segments included in the video item playing on the client device.

2. The method of claim 1 , further comprising:

determining, by the video content server, a second topical category for a second set of words derived from a second segment of the video item in a second timeframe prior to the video content slot, the second timeframe of the second segment temporally further than a timeframe of the segment relative to the video content slot;

determining, by the video content server, a first relevancy weight for the topical category based on a first temporal length between a first timeframe of the segment and the video content slot and a second relevancy weight for the second topical category based on a second temporal length between the second timeframe of the second segment in the video item and the video content slot, the first relevancy weight for the topical category greater than the second relevancy weight for the second topical category based on the first temporal length being less than the second temporal length; and

wherein selecting the secondary video content item further comprises selecting the secondary video content item based on the first relevancy weight for the topical category and the second relevancy weight for the second topical category.

3. The method of claim 1 , further comprising:

determining, by the video content server, a second topical category for an entire set of words derived from an entirety of the video item including the segment based on the entire set of words, the entire set of words including the set of words;

determining, by the video content server, a first relevancy weight for the topical category based on determining the topical category for the set of words derived from the segment and a second relevancy weight for the second topical category based on determining the second topical category from the entire set of words of the entirety of the video item, the first relevancy weight greater than the second relevancy weight; and

wherein selecting the secondary video content item further comprises selecting the secondary video content item based on the first relevancy weight for the topical category and the second relevancy weight for the second topical category.

4. The method of claim 1 , further comprising

determining, by the video content server, a content selection parameter for the video content slot in a timeframe relative to the content slot in the video item based on topical category; and

wherein selecting the secondary video content item further comprises selecting the secondary video content item from the plurality of candidate secondary video content items based on the content selection parameter.

5. The method of claim 1 , wherein identifying the set of words derived from the video item further comprises performing audio analysis on the segment of the video item in a timeframe prior to the video content slot to generate audio information, the audio information including at least one of a pitch and tone; and

wherein selecting the secondary video content item further comprises selecting the secondary video content item from the plurality of candidate secondary video content items based on the audio information.

6. The method of claim 1 , wherein identifying the set of words derived from the segment of the video item further comprises performing speech recognition onto the segment of the video item to generate the set of words.

7. The method of claim 1 , wherein identifying the set of words derived from the segment of the video item further comprises retrieving a transcript for the video item including the set of words.

8. The method of claim 1 , wherein receiving the request for video content further comprises receiving the request for video content including the set of words of the segment of the video item; and

wherein identifying the set of words derived from the video item further comprises identifying the set of words from the request for video content received from the client device.

9. The method of claim 1 , wherein receiving the request for video content further comprises receiving the request for video content to insert into the video content slot of the video item played on the client, the video item streamed by a primary video content server different from the video content server.

10. The method of claim 1 , wherein providing the secondary video content item to the client device further comprises providing the second video content item to the client device to insert the secondary video content item into the video content slot at one of a prior to, during, or subsequent to a second segment of the plurality of segments included in the video item.

11. A system for inserting content into videos based on associated text, comprising:

a video content server having one or more processors and memory, configured to:

receive, from a client device, a request for video content to insert into a video content slot of a plurality of video content slots included in a video item playing on the client device, the request generated responsive to execution of an encoding embedded in the video item for the video content slot;

identify a set of words derived from a segment of a plurality of segments included in the video item, the segment playable prior to the video content slot included in the video item playing on the client device;

determine a topical category for the segment playable prior to the video content slot based on the set of words derived from the segment of the video item playable prior to the video content slot;

select, from a plurality of candidate secondary video content items, responsive to receiving the request for video content, a secondary video content item based on the topical category of the segment of the video item; and

provide the secondary video content item to the client device to insert into the video content slot playable subsequent to the segment of the plurality of segments included in the video item playing on the client device.

12. The system of claim 11 , wherein the video content server is further configured to:

determine a second topical category for a second set of words derived from a second segment of the video item in a second timeframe prior to the video content slot, the second timeframe of the second segment temporally further than a timeframe of the segment relative to the video content slot;

determine a first relevancy weight for the topical category based on a first temporal length between the timeframe of the segment and the video content slot and a second relevancy weight for the second topical category based on a second temporal length between the second timeframe of the second segment in the video item and the video content slot, the first relevancy weight for the topical category greater than the second relevancy weight for the second topical category based on the first temporal length being less than the second temporal length; and

select the secondary video content item based on the first relevancy weight for the topical category and the second relevancy weight for the second topical category.

13. The system of claim 11 , wherein the video content server is further configured to:

determine a second topical category for an entire set of words derived from an entirety of the video item including the segment based on the entire set of words, the entire set of words including the set of words;

determine a first relevancy weight for the topical category based on determining the topical category for the set of words derived from the segment and a second relevancy weight for the second topical category based on determining the second topical category from the entire set of words of the entirety of the video item, the first relevancy weight greater than the second relevancy weight; and

select the secondary video content item based on the first relevancy weight for the topical category and the second relevancy weight for the second topical category.

14. The system of claim 11 , wherein the video content server is further configured to:

determine a content selection parameter for the video content slot in a timeframe relative to the content slot in the video item based on topical category; and

select the secondary video content item from the plurality of candidate secondary video content items based on the content selection parameter.

15. The system of claim 11 , wherein the video content server is further configured to:

determine a content selection parameter for the video content slot in a timeframe relative to the set of words in the video item based on topical category; and

select the secondary video content item from the plurality of candidate secondary video content items based on the content selection parameter.

16. The system of claim 11 , wherein the video content server is further configured to perform speech recognition onto the segment of the video item to generate the set of words.

17. The system of claim 11 , wherein the video content server is further configured to retrieve a transcript for the video item including the set of words.

18. The system of claim 11 , wherein the video content server is further configured to:

receive the request for video content including the set of words of the segment of the video item; and

identify the set of words from the request for video item content received from the client device.

19. The system of claim 11 , wherein the video content server is further configured to receive the request for video content to insert into the video content slot of the video item played on the client, the video item streamed by a primary video content server different from the video content server.

20. The system of claim 11 , wherein the video content server is further configured to provide the second video content item to the client device to insert the secondary video content item into the video content slot at one of a prior to, during, or subsequent to a second segment of the plurality of segments included in the video item.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2020
From: BAYER, JASON; CHAKRABARTI, RONOJOY; DESAI, KEVAL; GUPTA, MANISH; HUCHITAL, JILL A.; RUSCH, WILLARD
To: GOOGLE LLC
Reel/Frame 052548/0478 →
Continuity (5)
Continuation 14251824 · Apr 14, 2014
Continuation 14246826 · Apr 7, 2014
Continuation 14193695 · Feb 28, 2014
Continuation 11323327 · Dec 30, 2005
Related Publication 20190026790A1 · Jan 24, 2019