IP Library Granted Patent US 9,036,083
Granted Patent B1
US 9,036,083 · App. 14/289,142 · Granted May 19, 2015

Text detection in video

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,036,083
App. No.
14/289,142
Granted
May 19, 2015
Kind
B1
Abstract

Techniques of detecting text in video are disclosed. In some embodiments, a portion of video content can be identified as having text. Text within the identified portion of the video content can be identified. A category for the identified text can be determined. In some embodiments, a determination is made as to whether the video content satisfies at least one predetermined condition, and the portion of video content is identified as having text in response to a determination that the video content satisfies the predetermined condition(s). In some embodiments, the predetermined condition(s) comprises at least one of a minimum level of clarity, a minimum level of contrast, and a minimum level of content stability across multiple frames. In some embodiments, additional information corresponding to the video content is determined based on the identified text and the determined category.

Claims (64)

1. A computer-implemented method comprising:

identifying, by a machine having a memory and at least one processor, a portion of video content as having text, the identifying comprising:

converting a frame of the video content to grayscale;

performing edge detection on the frame;

performing dilation on the frame to connect vertical edges within the frame;

binarizing the frame;

performing a connected component analysis on the frame to detect connected components within the frame;

merging the connected components into a plurality of text lines;

refining the plurality of text lines using horizontal and vertical projections;

filtering out at least one of the plurality of text lines based on a size of the at least one of the plurality of text lines to form a filtered set of text lines;

binarizing the filtered set of text lines; and

filtering out at least one of the text lines from the binarized filtered set of text lines based on at least one of a shape of components in the at least one of the text lines and a position of components in the at least one of the text lines to form the portion of the video content having text;

identifying text within the identified portion of the video content; and

determining a category for the identified text.

2. The method of claim 1 , further comprising determining whether the video content satisfies at least one predetermined condition, wherein performing the identifying of the portion of video content as having text is conditioned upon a determination that the video content satisfies the at least one predetermined condition.

3. The method of claim 2 , wherein the at least one predetermined condition comprises at least one of a minimum level of clarity, a minimum level of contrast, and a minimum level of content stability across multiple frames.

4. The method of claim 1 , further comprising determining additional information corresponding to the video content based on the identified text and the determined category.

5. The method of claim 4 , further comprising causing the additional information to be displayed on a media content device.

6. The method of claim 4 , further comprising storing the additional information in association with the video content or in association with an identified viewer of the video content.

7. The method of claim 4 , further comprising providing the additional information to a software application on a media content device.

8. The method of claim 7 , wherein the additional information comprises at least one of a uniform resource locator (URL), an identification of a user account, a metadata tag, and a phone number.

9. The method of claim 7 , wherein the media content device comprises one of a television, a laptop computer, a desktop computer, a tablet computer, and a smartphone.

10. The method of claim 1 , further comprising storing the identified text in association with the video content or in association with an identified viewer of the video content.

11. The method of claim 1 , wherein identifying text within the identified portion of the video content comprises performing optical character recognition on the identified portion of the video content.

12. The method of claim 1 , wherein determining the category for the identified text comprises:

parsing the identified text to determine a plurality of segments of the identified text; and

determining the category based on a stored association between at least one of the plurality of segments and the category.

13. The method of claim 1 , wherein the video content comprises a portion of a television program, a non-episodic movie, a webisode, user-generated content for a video-sharing website, or a commercial.

14. A system comprising:

a machine having a memory and at least one processor; and

at least one module on the machine, the at least one module being configured to:

identify a portion of video content as having text, the identifying comprising:

converting a frame of the video content to grayscale;

performing edge detection on the frame;

performing dilation on the frame to connect vertical edges within the frame;

binarizing the frame;

performing a connected component analysis on the frame to detect connected components within the frame;

merging the connected components into a plurality of text lines;

refining the plurality of text lines using horizontal and vertical projections;

filtering out at least one of the plurality of text lines based on a size of the at least one of the plurality of text lines to form a filtered set of text lines:

binarizing the filtered set of text lines; and

filtering out at least one of the text lines from the binarized filtered set of text lines based on at least one of a shape of components in the at least one of the text lines and a position of components in the at least one of the text lines to form the portion of the video content having text;

identify text within the identified portion of the video content; and

determine a category for the identified text.

15. The system of claim 14 , wherein the at least one module is further configured to:

determine whether the video content satisfies at least one predetermined condition; and

identify the portion of video content as having text in response to a determination that the video content satisfies the at least one predetermined condition.

16. The system of claim 15 , wherein the at least one predetermined condition comprises at least one of a minimum level of clarity, a minimum level of contrast, and a minimum level of content stability across multiple frames.

17. The system of claim 14 , wherein the at least one module is further configured to determine additional information corresponding to the video content based on the identified text and the determined category.

18. The system of claim 17 , wherein the at least one module is further configured to provide the additional information to a software application on a media content device.

19. A non-transitory machine-readable storage device, tangibly embodying a set of instructions that, when executed by at least one processor, causes the at least one processor to perform a set of operations comprising:

identifying a portion of video content as having text, the identifying comprising:

converting a frame of the video content to grayscale;

performing edge detection on the frame;

performing dilation on the frame to connect vertical edges within the frame;

binarizing the frame;

performing a connected component analysis on the frame to detect connected components within the frame;

merging the connected components into a plurality of text lines;

refining the plurality of text lines using horizontal and vertical projections;

filtering out at least one of the plurality of text lines based on a size of the at least one of the plurality of text lines to form a filtered set of text lines;

binarizing the filtered set of text lines; and

filtering out at least one of the text lines from the binarized filtered set of text lines based on at least one of a shape of components in the at least one of the text lines and a position of components in the at least one of the text lines to form the portion of the video content having text;

identifying text within the identified portion of the video content; and

determining a category for the identified text.

Assignments (12)
RELEASE (REEL 053473 / FRAME 0001) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063603/0001 →
RELEASE (REEL 054066 / FRAME 0064) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063605/0001 →
SECURITY INTEREST Recorded May 8, 2023
From: GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE, INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC
To: ARES CAPITAL CORPORATION
Reel/Frame 063574/0632 →
SECURITY INTEREST Recorded Apr 28, 2023
From: GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE, INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC
To: CITIBANK, N.A.
Reel/Frame 063561/0381 →
SECURITY AGREEMENT Recorded Jan 31, 2023
From: GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE, INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 063560/0547 →
RELEASE (REEL 042262 / FRAME 0601) Recorded Oct 13, 2022
From: CITIBANK, N.A.
To: GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC
Reel/Frame 061748/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE PATENTS LISTED ON SCHEDULE 1 RECORDED ON 6-9-2020 PREVIOUSLY RECORDED ON REEL 053473 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE SUPPLEMENTAL IP SECURITY AGREEMENT. Recorded Oct 7, 2020
From: A.C. NIELSEN (ARGENTINA) S.A.; A.C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A
Reel/Frame 054066/0064 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Jun 9, 2020
From: A. C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NIELSEN UK FINANCE I, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A.
Reel/Frame 053473/0001 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Apr 13, 2017
From: GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE DIGITAL VENTURES, LLC
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 042262/0601 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Feb 8, 2017
From: JPMORGAN CHASE BANK, N.A.
To: GRACENOTE, INC.; CASTTV INC.; TRIBUNE MEDIA SERVICES, LLC; TRIBUNE DIGITAL VENTURES, LLC
Reel/Frame 041656/0804 →
SECURITY AGREEMENT Recorded Nov 13, 2014
From: GRACENOTE, INC.; TRIBUNE BROADCASTING COMPANY, LLC; TRIBUNE DIGITAL VENTURES, LLC; TRIBUNE MEDIA COMPANY
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 034231/0333 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2014
From: ZHU, IRENE; HARRON, WILSON; CREMER, MARKUS K
To: GRACENOTE, INC.
Reel/Frame 033009/0448 →