IP Library › Granted Patent US 11,720,793
Granted Patent B2
US 11,720,793 · App. 17/069,638 · Granted Aug 8, 2023

Video anchors

Inventors: Gabe Culbertson (Palo Alto, CA); Wei Peng (Fremont, CA); Nicolas Crowell (San Francisco, CA)
Assignee: GOOGLE LLC
G06N3/08G06F18/22G06F18/23G06N5/02G06V10/761G06V10/762G06V20/41G06V20/47G06V20/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,720,793
App. No.
17/069,638
Granted
Aug 8, 2023
Kind
B2
Abstract

In one aspect, a method includes obtaining videos and for each video: obtaining a set of anchors for the video, each anchor beginning at the playback time and including anchor text; identifying, from text generated from audio of the video, a set of entities specified in the text, wherein each entity in the set of entities is associated with a times stamp at which the entity is mentioned; determining, by a language model and from the text generated from the audio of the video, an importance value for each entity; for a subset of the videos, receiving rater data that describes, for each anchor, the accuracy of the anchor text in describing subject matter of the video; and training, using the human rater data, the importance values, the text, and the set of entities, an anchor model that predicts an entity label for an anchor for a video.

Claims (56)

1. A computer-implemented method, comprising:

obtaining a plurality of videos, and for each video of the plurality of videos:

obtaining a set of anchors for the video, each anchor in the set of anchors for the video beginning at the playback time specified by a respective time index value of a time in the video, and each anchor in the set of anchors including anchor text;

identifying, from text generated from audio of the video, a set of entities specified in the text, wherein each entity in the set of entities is an entity specified in an entity corpus that defines a list of entities and is associated with a times stamp that indicates a time in the video at which the entity is mentioned;

determining, by a language model and from the text generated from the audio of the video, an importance value for each entity in the set of entities, each importance value indicating an importance of the entity for a context defined by the text generated from the audio of the video;

for a proper subset of the videos, receiving, for each video in the proper subset of videos, human rater data that describes, for each anchor for the video, the accuracy of the anchor text of the anchor in describing subject matter of the video beginning at the time index value specified by the respective time index value of the anchor; and

training, using the human rater data, the importance values, the text generated from the audio of the videos, the set of entities, an anchor model that predicts an entity label for an anchor for a video at a particular time in the video.

2. The computer-implemented method of claim 1 , further comprising, for each video, determining salient terms for the video, where each salient term is a term that is descriptive of the video; and

wherein training the anchor model further comprises training the anchor model using the salient terms.

3. The computer-implemented method of claim 2 , wherein identifying, from text generated from audio of the video, a set of entities specified in the text comprises:

determining hypernyms for each entity;

clustering the entities into entity clusters based on a similarity of the hypernyms; and

filtering entity clusters that are determined to not meet filtering criteria.

4. The computer-implemented method of claim 3 , wherein the filtering criteria includes one or more of: broadness of the entities in an entity cluster, a minimum number of entities in the entity cluster, and a similarity threshold of the hypernyms of entities that belong to the entity cluster and the salient terms determined for the video.

5. The computer-implemented method of claim 1 , wherein obtaining the plurality of videos comprises, for each video of the plurality of videos, obtaining the video only if the video includes a minimum plurality of anchors in the set of anchors.

6. The computer-implemented method of claim 1 , wherein identifying a set of entities specified in the text comprises identifying an entity only when the entity has a unique entry in a knowledge graph.

7. The computer-implemented method of claim 1 , further comprising:

providing, after training the anchor label model and as input to the anchor model, a video; and

receiving, as output from the anchor model, a set of anchors for the video, each anchor in the set of anchors for the video beginning at the playback time specified by a respective time index value of a time in the video, and each anchor in the set of anchors including anchor text that is predicted to be descriptive of subject matter in the video beginning at the time index value.

8. A system, comprising:

a data processing apparatus; and

a non-transitory computer readable medium in data communication with the data processing apparatus and storing instructions executable by the data processing apparatus and that upon such execution cause the data processing apparatus to perform operations comprising:

obtaining a plurality of videos and for each video of the plurality of videos:

obtaining a set of anchors for the video, each anchor in the set of anchors for the video beginning at the playback time specified by a respective time index value of a time in the video, and each anchor in the set of anchors including anchor text;

identifying, from text generated from audio of the video, a set of entities specified in the text, wherein each entity in the set of entities is an entity specified in an entity corpus that defines a list of entities and is associated with a times stamp that indicates a time in the video at which the entity is mentioned;

determining, by a language model and from the text generated from the audio of the video, an importance value for each entity in the set of entities, each importance value indicating an importance of the entity for a context defined by the text generated from the audio of the video;

for a proper subset of the videos, receiving, for each video in the proper subset of videos, human rater data that describes, for each anchor for the video, the accuracy of the anchor text of the anchor in describing subject matter of the video beginning at the time index value specified by the respective time index value of the anchor; and

training, using the human rater data, the importance values, the text generated from the audio of the videos, the set of entities, an anchor model that predicts an entity label for an anchor for a video at a particular time in the video.

9. The system of claim 8 , the operations further comprising, for each video, determining salient terms for the video, where each salient term is a term that is descriptive of the video; and

wherein training the anchor model further comprises training the anchor model using the salient terms.

10. The system of claim 9 , wherein identifying, from text generated from audio of the video, a set of entities specified in the text comprises:

determining hypernyms for each entity;

clustering the entities into entity clusters based on a similarity of the hypernyms; and

filtering entity clusters that are determined to not meet filtering criteria.

11. The system of claim 10 , wherein the filtering criteria includes one or more of: broadness of the entities in an entity cluster, a minimum number of entities in the entity cluster, and a similarity threshold of the hypernyms of entities that belong to the entity cluster and the salient terms determined for the video.

12. The system of claim 8 , wherein obtaining the plurality of videos comprises, for each video of the plurality of videos, obtaining the video only if the video includes a minimum plurality of anchors in the set of anchors.

13. The system of claim 8 , wherein identifying a set of entities specified in the text comprises identifying an entity only when the entity has a unique entry in a knowledge graph.

14. The system of claim 8 , further comprising:

providing, after training the anchor label model and as input to the anchor model, a video; and

receiving, as output from the anchor model, a set of anchors for the video, each anchor in the set of anchors for the video beginning at the playback time specified by a respective time index value of a time in the video, and each anchor in the set of anchors including anchor text that is predicted to be descriptive of subject matter in the video beginning at the time index value.

15. A non-transitory computer readable medium storing instructions executable by the data processing apparatus and that upon such execution cause the data processing apparatus to perform operations comprising;

obtaining a plurality of videos and for each video of the plurality of videos:

obtaining a set of anchors for the video, each anchor in the set of anchors for the video beginning at the playback time specified by a respective time index value of a time in the video, and each anchor in the set of anchors including anchor text;

identifying, from text generated from audio of the video, a set of entities specified in the text, wherein each entity in the set of entities is an entity specified in an entity corpus that defines a list of entities and is associated with a times stamp that indicates a time in the video at which the entity is mentioned;

determining, by a language model and from the text generated from the audio of the video, an importance value for each entity in the set of entities, each importance value indicating an importance of the entity for a context defined by the text generated from the audio of the video;

for a proper subset of the videos, receiving, for each video in the proper subset of videos, human rater data that describes, for each anchor for the video, the accuracy of the anchor text of the anchor in describing subject matter of the video beginning at the time index value specified by the respective time index value of the anchor; and

training, using the human rater data, the importance values, the text generated from the audio of the videos, the set of entities, an anchor model that predicts an entity label for an anchor for a video at a particular time in the video.

16. The non-transitory computer readable medium of claim 15 , further comprising, for each video, determining salient terms for the video, where each salient term is a term that is descriptive of the video; and

wherein training the anchor model further comprises training the anchor model using the salient terms.

17. The non-transitory computer readable medium of claim 16 , wherein identifying, from text generated from audio of the video, a set of entities specified in the text comprises:

determining hypernyms for each entity;

clustering the entities into entity clusters based on a similarity of the hypernyms; and

filtering entity clusters that are determined to not meet filtering criteria.

18. The non-transitory computer readable medium of claim 17 , wherein the filtering criteria includes one or more of: broadness of the entities in an entity cluster, a minimum number of entities in the entity cluster, and a similarity threshold of the hypernyms of entities that belong to the entity cluster and the salient terms determined for the video.

19. The non-transitory computer readable medium of claim 15 , wherein obtaining the plurality of videos comprises, for each video of the plurality of videos, obtaining the video only if the video includes a minimum plurality of anchors in the set of anchors.

20. The non-transitory computer readable medium of claim 16 , wherein each video is included in a resource page that also includes text, and the salient terms are determined from the text included in the resource page.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2020
From: CULBERTSON, GABE; PENG, WEI; CROWELL, NICOLAS
To: GOOGLE LLC
Reel/Frame 054425/0237 →
Continuity (2)
Provisional Application 62914684 · Oct 14, 2019
Related Publication 20210110163A1 · Apr 15, 2021