IP Library Granted Patent US 10,869,105
Granted Patent B2
US 10,869,105 · App. 15/912,790 · Granted Dec 15, 2020

Voice-driven metadata media content tagging

Inventor: Jason Henderson (Littleton, CO)
Assignee: DISH Network L.L.C.
H04N21/8405H04N21/2353H04N21/23109G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,869,105
App. No.
15/912,790
Granted
Dec 15, 2020
Kind
B2
Abstract

Various arrangements for voice-based metadata tagging of video content are presented. A request to add a spoken metadata tag to be linked with a video content instance may be received. A voice clip that includes audio spoken by a user may be received. Speech-to-text conversion of the voice clip to produce a proposed spoken metadata tag may be performed. A metadata integration database to link the spoken metadata tag with the video content instance may be updated.

Claims (112)

1. A method for voice-based metadata tagging of video content, the method comprising:

receiving, by a television receiver, via an electronic programming guide (EPG), a request to add a spoken metadata tag to be linked with a video content instance;

receiving, by the television receiver via a microphone integrated as part of a remote control unit, a voice clip, wherein the voice clip comprises audio spoken by a user;

transmitting, by the television receiver, the voice clip to a metadata integration server system via the Internet, wherein:

the metadata integration server system maintains a crowdsourced metadata integration database that is updated based on spoken metadata tags submitted by a plurality of content viewers via a plurality of television receivers, the plurality of television receivers comprising the television receiver;

performing, by the metadata integration server system, speech-to-text conversion of the voice clip to produce a proposed spoken metadata tag;

transmitting, by the metadata integration server system, the proposed spoken metadata tag to the television receiver;

outputting, by the television receiver, the proposed spoken metadata tag for presentation;

receiving, by the television receiver, from the remote control unit, confirmation of the proposed spoken metadata tag to be the spoken metadata tag;

in response to the confirmation, updating, by the metadata integration server system, the crowdsourced metadata integration database to link the spoken metadata tag with the video content instance, wherein updating the crowdsourced metadata integration database comprises:

determining a number of times that the spoken metadata tag has been submitted for the video content instance;

determining that the number of times exceeds a minimum tag threshold; and

linking the spoken metadata tag with the video content instance in response to the number of times being determined to exceed the minimum tag threshold;

determining that the number of times exceeds a presentation threshold; and

in response to determining that the number of times exceeds the presentation threshold, updating an EPG entry for the video content instance such that the spoken metadata tag is visually presented as part of the EPG entry;

receiving, by the metadata integration server system, a content search;

transmitting, by the metadata integration server system, content search results that are indicative of the video content instance, wherein the content search results are based at least in part on the spoken metadata tag being linked with the video content instance in the metadata integration database; and

outputting, by the television receiver, for presentation the content search results.

2. The method for voice-based metadata tagging and searching of video content of claim 1 , further comprising:

receiving, by the television receiver, selection of the video content instance from the content search results; and

in response to the selection of the video content instance from the content search results, outputting, by the television receiver, for presentation the EPG entry for the video content instance such that the spoken metadata tag is visually presented as part of the EPG entry.

3. The method for voice-based metadata tagging and searching of video content of claim 1 , wherein updating the crowdsourced metadata integration database further comprises:

determining that the number of times does not exceed a presentation threshold; and

in response to the number of times not exceeding the presentation threshold but exceeding the minimum tag threshold, causing the content search results to include the video content instance, but not visually presenting the spoken metadata tag as part of an EPG entry.

4. The method for voice-based metadata tagging and searching of video content of claim 1 , further comprising:

accessing, by the metadata integration server system, a third-party database that maintains metadata for a plurality of video content instances; and

updating, by the metadata integration server system, the crowdsourced metadata integration database based on metadata from the third-party database.

5. The method for voice-based metadata tagging and searching of video content of claim 1 , wherein performing the speech-to-text conversion of the voice clip to produce the proposed spoken metadata tag comprises:

accessing a third-party database that maintains metadata for a plurality of video content instances, wherein the plurality of video content instances comprises the video content instance; and

determining a spelling of the proposed spoken metadata tag at least partially based on metadata linked with the video content instance in the third-party database.

6. The method for voice-based metadata tagging and searching of video content of claim 1 , wherein performing the speech-to-text conversion of the voice clip to produce the proposed spoken metadata tag comprises:

extracting only nouns from the voice clip to produce the proposed spoken metadata tag.

7. A method for voice-based metadata tagging of video content, the method comprising:

receiving, by a television receiver, via an electronic programming guide (EPG), a request to add a spoken metadata tag to be linked with a video content instance;

receiving, by the television receiver via a microphone integrated as part of a remote control unit, a voice clip, wherein the voice clip comprises audio spoken by a user;

transmitting, by the television receiver, the voice clip to a metadata integration server system via the Internet, wherein:

the metadata integration server system maintains a crowdsourced metadata integration database that is updated based on spoken metadata tags submitted by a plurality of content viewers via a plurality of television receivers, the plurality of television receivers comprising the television receiver;

performing, by the metadata integration server system, speech-to-text conversion of the voice clip to produce a proposed spoken metadata tag;

transmitting, by the metadata integration server system, the proposed spoken metadata tag to the television receiver;

outputting, by the television receiver, the proposed spoken metadata tag for presentation;

receiving, by the television receiver, from the remote control unit, confirmation of the proposed spoken metadata tag to be the spoken metadata tag;

in response to the confirmation, updating, by the metadata integration server system, the crowdsourced metadata integration database to link the spoken metadata tag with the video content instance;

receiving, by the metadata integration server system, a content search;

transmitting, by the metadata integration server system, content search results that are indicative of the video content instance, wherein the content search results are based at least in part on the spoken metadata tag being linked with the video content instance in the metadata integration database;

outputting, by the television receiver, for presentation the content search results;

receiving, by the television receiver, via the EPG, a second request to add a second spoken metadata tag to be linked with a second video content instance;

receiving, by the television receiver via the microphone integrated as part of the remote control unit, a second voice clip, wherein the second voice clip comprises audio spoken by a user;

transmitting, by the television receiver, the second voice clip to the metadata integration server system via the Internet;

performing, by the metadata integration server system, a second speech-to-text conversion of the second voice clip to produce a second proposed spoken metadata tag;

transmitting, by the metadata integration server system, the second proposed spoken metadata tag to the television receiver;

outputting, by the television receiver, the second proposed spoken metadata tag for presentation;

receiving, by the television receiver, from the remote control unit, cancellation of the second proposed spoken metadata tag; and

in response to the cancellation, updating, by the metadata integration server system, the crowdsourced metadata integration database to link the second proposed spoken metadata tag with the second video content instance, wherein the second proposed spoken metadata tag is assigned a lower weight due to the cancellation than a higher weight assigned the spoken metadata tag.

8. A system for voice-based metadata tagging of video content, the system comprising:

a remote control comprising an integrated microphone to capture spoken audio clips;

a television receiver, configured to:

receive, via an electronic programming guide (EPG) interface, a request to add a spoken metadata tag to be linked with a video content instance;

receive, from the remote control, a voice clip, wherein the voice clip comprises audio spoken by a user; and

transmit the voice clip to a metadata integration server system via the Internet; and

the metadata integration server system that maintains a crowdsourced metadata integration database updated based on spoken metadata tags submitted by a plurality of content viewers via a plurality of television receivers, the plurality of television receivers comprising the television receiver, the metadata integration server system configured to:

perform speech-to-text conversion of the voice clip to produce a proposed spoken metadata tag;

transmit the proposed spoken metadata tag to the television receiver;

in response to a received confirmation, update a metadata integration database to increase a weight of the spoken metadata tag for the video content instance, wherein:

updating the metadata integration database comprises:

determine a number of times that the spoken metadata tag has been submitted for the video content instance;

determine that the number of times exceeds a minimum tag threshold;

link the spoken metadata tag with the video content instance in response to the number of times being determined to exceed the minimum tag threshold;

determine that the number of times does not exceed a presentation threshold; and

in response to the number of times not exceeding the presentation threshold but exceeding a minimum tag threshold, cause the content search results to include the video content instance, but not visually presenting the spoken metadata tag as part of an EPG entry; and

the spoken metadata tag has been received previously for the video content instance from another television receiver of the plurality of television receivers;

receive a content search; and

transmit content search results that are indicative of the video content instance, wherein the content search results are based at least in part on the spoken metadata tag being linked with the video content instance in the crowdsourced metadata integration database, wherein the television receiver is further configured to output for presentation the content search results.

9. The system for voice-based metadata tagging and searching of video content of claim 8 , wherein the metadata integration server system being configured to update the metadata integration database comprises the metadata integration server system being configured to:

determine that the number of times exceeds a presentation threshold; and

in response to determining that the number of times exceeds the presentation threshold, update an EPG entry for the video content instance such that the spoken metadata tag is visually presented as part of the EPG entry.

10. The system for voice-based metadata tagging and searching of video content of claim 9 , wherein the television receiver is further configured to:

receive selection of the video content instance from the content search results; and

in response to the selection of the video content instance from the content search results, output for presentation the EPG entry for the video content instance such that the spoken metadata tag is visually presented as part of the EPG entry.

11. The system for voice-based metadata tagging and searching of video content of claim 8 , wherein the metadata integration server system is further configured to:

access a third-party database that maintains metadata for a plurality of video content instances; and

update the metadata integration database based on metadata from the third-party database.

12. The system for voice-based metadata tagging and searching of video content of claim 8 , wherein the metadata integration server system being configured to perform the speech-to-text conversion of the voice clip to produce the proposed spoken metadata tag comprises the metadata integration server system being configured to:

access a third-party database that maintains metadata for a plurality of video content instances, wherein the plurality of video content instances comprises the video content instance; and

determine a spelling of the proposed spoken metadata tag at least partially based on metadata linked with the video content instance in the third-party database.

13. The system for voice-based metadata tagging and searching of video content of claim 8 , wherein the metadata integration server system being configured to perform the speech-to-text conversion of the voice clip to produce the proposed spoken metadata tag comprises the metadata integration server system being configured to:

extract only nouns from the voice clip to produce the proposed spoken metadata tag.

14. The system for voice-based metadata tagging and searching of video content of claim 8 , wherein the television receiver is further configured to:

receive, via the EPG, a second request to add a second spoken metadata tag to be linked with a second video content instance;

receive, from the remote control, a second voice clip, wherein the second voice clip comprises audio spoken by a user;

transmit the second voice clip to the metadata integration server system via the Internet; and

wherein the metadata integration server system is further configured to:

perform a second speech-to-text conversion of the second voice clip to produce a second proposed spoken metadata tag;

transmit the second proposed spoken metadata tag to the television receiver for presentation; and

in response to a cancellation received from the television receiver, update the metadata integration database to link the second proposed spoken metadata tag with the second video content instance, wherein the second proposed spoken metadata tag is assigned a lower weight due to the cancellation than a higher weight assigned the spoken metadata tag.

15. An apparatus for voice-based metadata tagging of video content, the apparatus comprising:

means for receiving a request to add a spoken metadata tag to be linked with a video content instance;

means for receiving a voice clip, wherein the voice clip comprises audio spoken by a user;

means for performing speech-to-text conversion of the voice clip to produce a proposed spoken metadata tag;

means for outputting the proposed spoken metadata tag for presentation;

means for receiving confirmation of the proposed spoken metadata tag to be the spoken metadata tag;

means for updating a crowdsourced metadata integration database to link the spoken metadata tag with the video content instance in response to the confirmation, wherein:

the means for updating the crowdsourced metadata integration data comprises:

means for updating the crowdsourced metadata integration database comprises:

means for determining a number of times that the spoken metadata tag has been submitted for the video content instance;

means for determining that the number of times exceeds a minimum tag threshold; and

means for linking the spoken metadata tag with the video content instance in response to the number of times being determined to exceed the minimum tag threshold;

means for determining that the number of times exceeds a presentation threshold; and

in response to determining that the number of times exceeds the presentation threshold, means for updating an EPG entry for the video content instance such that the spoken metadata tag is visually presented as part of the EPG entry; and

the crowdsourced metadata integration database is updated based on spoken metadata tags submitted by a plurality of content viewers;

means for receiving a content search;

means for providing content search results that are indicative of the video content instance, wherein the content search results are based at least in part on the spoken metadata tag being linked with the video content instance in the crowdsourced metadata integration database; and

means for outputting for presentation the content search results.

Assignments (2)
SECURITY INTEREST Recorded Nov 30, 2021
From: DISH BROADCASTING CORPORATION; DISH NETWORK L.L.C.; DISH TECHNOLOGIES L.L.C.
To: U.S. BANK, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 058295/0293 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2018
From: HENDERSON, JASON
To: DISH NETWORK L.L.C.
Reel/Frame 045461/0065 →