IP Library Granted Patent US 12,225,248
Granted Patent B2
US 12,225,248 · App. 18/389,315 · Granted Feb 11, 2025

Systems and methods for correcting errors in caption text

Inventors: Ajay Kumar Gupta (Andover, MA); Abhijit Satchidanand Savarkar (Andover, MA)
Assignee: Adeia Guides Inc.
H04N21/23424G06F40/166G06F40/232G06F40/40H04N21/234336H04N21/4884
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,225,248
App. No.
18/389,315
Granted
Feb 11, 2025
Kind
B2
Abstract

Systems and methods are described to address shortcomings in conventional systems by correcting an erroneous term in on-screen caption text for a media asset. In some aspects, the systems and methods identify the erroneous term in a text segment of the on-screen caption text, and identify one or more video frames of the media asset corresponding to the text segment. The systems and methods further identify a contextual term related to the erroneous term from the one or more video frames. By accessing a knowledge graph, the systems and methods identify a candidate correction based on the contextual term and a portion of the text segment. Lastly, the systems and methods replaces the erroneous term with the candidate correction.

Claims (53)

1. A method comprising:

identifying an erroneous term in a text portion of a video asset;

identifying a non-textual visual object in a video frame from the video asset that is associated with the erroneous term, wherein a descriptor of the identified non-textual visual object is not a candidate correction term for the erroneous term, and wherein the non-textual visual object comprises an image of a person;

analyzing the identified non-textual visual object in the video frame, using image recognition, to identify an object category based on the descriptor of the identified non-textual visual object, wherein the identified object category comprises an indication of an identity of the person;

identifying the candidate correction term for the erroneous term that belongs to the identified object category; and

replacing the erroneous term in the text portion of the video asset with the identified candidate correction term.

2. The method of claim 1 , wherein the identifying the erroneous term in the text portion of the video asset comprises performing natural language processing on the text portion of the video asset to compare the text portion of the video asset against a plurality of grammar rules.

3. The method of claim 1 , wherein the identifying the non-textual visual object in the video frame from the video asset that is associated with the erroneous term comprises extracting a first video frame at a position of the video asset corresponding to a position of a time-stamped text portion of the video asset.

4. The method of claim 1 , further comprising:

generating the text portion of the video asset by analyzing an audio stream of a media asset.

5. The method of claim 4 , wherein the text portion of the video asset is time-stamped, and wherein the video frame is extracted at a position of the video asset corresponding to a position of the erroneous term in a time-stamped text portion of the video asset.

6. The method of claim 1 , wherein the replacing the erroneous term in the text portion of the video asset with the identified candidate correction term comprises replacing the erroneous term with the identified candidate correction term while a live broadcast of the video asset is being generated for presentation.

7. The method of claim 1 , wherein the replacing the erroneous term in the text portion of the video asset with the identified candidate correction term comprises replacing the erroneous term with the identified candidate correction term when the video asset is being stored as on-demand content.

8. A method comprising:

identifying an erroneous term in a text portion of a video asset;

identifying an object in a video frame from the video asset, wherein a descriptor of the identified object is not a candidate correction term for the erroneous term;

analyzing the identified object in the video frame, using image recognition, to identify an object category based on the descriptor of the identified object;

identifying the candidate correction term for the erroneous term that belongs to the identified object category, wherein the identifying the candidate correction term that belongs to the identified object category comprises:

extracting a keyword from the text portion of the video asset;

searching a knowledge graph for nodes corresponding to the identified object category and the keyword;

analyzing the nodes for properties associated with the identified object category and the keyword; and

determining at least one other node based at least in part on the properties associated with the identified object category and the keyword, wherein the at least one other node corresponds to the candidate correction term that belongs to the identified object category; and

replacing the erroneous term in the text portion of the video asset with the identified candidate correction term.

9. The method of claim 8 , wherein the identifying the candidate correction term that belongs to the identified object category further comprises:

determining a plurality of candidate correction terms for the erroneous term from the knowledge graph;

assigning a weight to each candidate correction term of the plurality of candidate correction terms based on the determining; and

identifying a candidate correction term associated with a highest weight as the identified candidate correction term.

10. The method of claim 9 , wherein a more recent candidate correction term of the plurality of candidate correction terms is assigned a higher weight, and wherein the more recent candidate correction term is a candidate term associated with a more recent time-stamp, a candidate term that has been updated more recently or a candidate term that has gained popularity in recent searches.

11. A system comprising:

an input/output circuitry configured to receive a video asset; and

a control circuitry configured to:

identify an erroneous term in a text portion of the video asset;

identify a non-textual visual object in a video frame from the video asset that is associated with the erroneous term, wherein a descriptor associated with the identified non-textual visual object is not a candidate correction term for the erroneous term, and wherein the non-textual visual object comprises an image of a person;

analyze the identified non-textual visual object in the video frame, using image recognition, to identify an object category based on the descriptor associated with the identified non-textual visual object, wherein the identified object category comprises an indication of an identity of the person;

identify the candidate correction term for the erroneous term that belongs to the identified object category; and

replace the erroneous term in the text portion of the video asset with the identified candidate correction term.

12. The system of claim 11 , wherein the control circuitry is configured to identify the erroneous term in the text portion of the video asset by performing natural language processing on the text portion of the video asset to compare the text portion of the video asset against a plurality of grammar rules.

13. The system of claim 11 , wherein the control circuitry is configured to identify the non-textual visual object in the video frame from the video asset that is associated with the erroneous term by extracting a first video frame at a position of the video asset corresponding to a position of a time-stamped text portion of the video asset.

14. The system of claim 11 , wherein the control circuitry is further configured to:

generate the text portion of the video asset by analyzing an audio stream of a media asset.

15. The system of claim 14 , wherein the text portion of the video asset is time-stamped, and wherein the control circuitry is further configured to extract the video frame at a position of the video asset corresponding to a position of the erroneous term in a time-stamped text portion of the video asset.

16. The system of claim 11 , wherein the control circuitry is configured to identify the candidate correction term that belongs to the identified object category by:

extracting a keyword from the text portion of the video asset;

searching a knowledge graph for nodes corresponding to the identified object category and the keyword;

analyzing the nodes for properties associated with the identified object category and the keyword; and

determining at least one other node based on the properties associated with the identified object category and the keyword, wherein the at least one other node corresponds to the candidate correction term that belongs to the identified object category.

17. The system of claim 16 , wherein the control circuitry is configured to identify the candidate correction term that belongs to the identified object category by:

determining a plurality of candidate correction terms for the erroneous term from the knowledge graph;

assigning a weight to each candidate correction term of the plurality of candidate correction terms based on the determining; and

identifying a candidate correction term associated with a highest weight as the identified candidate correction term.

18. The system of claim 17 , wherein a more recent candidate correction term of the plurality of candidate correction terms is assigned a higher weight, and wherein the more recent candidate correction term is a candidate term associated with a more recent time-stamp, a candidate term that has been updated more recently or a candidate term that has gained popularity in recent searches.

19. The system of claim 11 , wherein the control circuitry is configured to replace the erroneous term in the text portion of the video asset with the identified candidate correction term by replacing the erroneous term with the identified candidate correction term while a live broadcast of the video asset is being generated for presentation.

20. The system of claim 11 , wherein the control circuitry is configured to replace the erroneous term in the text portion of the video asset with the identified candidate correction term by replacing the erroneous term with the identified candidate correction term when the video asset is being stored as on-demand content.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069085/0755 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2023
From: GUPTA, AJAY KUMAR; SAVARKAR, ABHIJIT SATCHIDANAND
To: ROVI GUIDES, INC.
Reel/Frame 065582/0175 →
Continuity (3)
Continuation 17063373 · Oct 5, 2020
Continuation 16067036
Related Publication 20240089516A1 · Mar 14, 2024
References Cited (35)
US 5617119A · Briggs et al. · 1997 [cited by applicant]
US 6239794B1 · Yuen et al. · 2001 [cited by applicant]
US 6473778B1 · Gibbon · 2002 [cited by applicant]
US 6564378B1 · Satterfield et al. · 2003 [cited by applicant]
US 7165098B1 · Boyer et al. · 2007 [cited by applicant]
US 7296218B2 · Dittrich · 2007 [cited by applicant]
US 7761892B2 · Ellis et al. · 2010 [cited by applicant]
US 8046801B2 · Ellis et al. · 2011 [cited by applicant]
US 20020161578A1 · Saindon · 2002 [cited by examiner]
US 20020174430A1 · Ellis et al. · 2002 [cited by applicant]
US 20050251827A1 · Ellis et al. · 2005 [cited by applicant]
US 20070118357A1 · Kasravi et al. · 2007 [cited by applicant]
US 20070118364A1 · Wise et al. · 2007 [cited by applicant]
US 20070118372A1 · Wise et al. · 2007 [cited by applicant]
US 20070118374A1 · Wise et al. · 2007 [cited by applicant]
US 20090185074A1 · Streijl · 2009 [cited by applicant]
US 20100121936A1 · Liu et al. · 2010 [cited by applicant]
US 20100153885A1 · Yates · 2010 [cited by applicant]
US 20110134321A1 · Berry et al. · 2011 [cited by applicant]
US 20110321100A1 · Tofighbakhsh · 2011 [cited by applicant]
US 20120304057A1 · Labsky · 2012 [cited by examiner]
US 20130019267A1 · Tofighbakhsh · 2013 [cited by applicant]
US 20150046148A1 · Oh et al. · 2015 [cited by applicant]
US 20150142704A1 · London · 2015 [cited by applicant]
US 20160035392A1 · Taylor et al. · 2016 [cited by applicant]
US 20160092447A1 · Venkataraman et al. · 2016 [cited by applicant]
US 20190215545A1 · Gupta et al. · 2019 [cited by applicant]
US 20210037274A1 · Gupta et al. · 2021 [cited by applicant]
CN 105654946A · 2016 [cited by applicant]
EP 1848192 · 2007 [cited by applicant]
JP 2007256714A · 2007 [cited by applicant]
JP 2016110087A · 2016 [cited by applicant]
WO 2015113578A1 · 2015 [cited by applicant]
WO WO2015113578 · 2015 [cited by applicant]
“Systems and Methods for Correcting Errors in Caption Text”, PCT International Search Report for International Application No. PCT/US2016/054689, mailed Feb. 8, 2017 (14 Pages), Feb. 8, 2017, 1-14. [cited by applicant]