IP Library Granted Patent US 12,374,138
Granted Patent B2
US 12,374,138 · App. 18/067,136 · Granted Jul 29, 2025

Text detection in videos

Inventors: Eliyahu Strugo (Tel Aviv, IL); Yonit Hoffman (Herzeliya, IL)
Assignee: Microsoft Technology Licensing, LLC
G06V30/19093G06V2201/09
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,374,138
App. No.
18/067,136
Granted
Jul 29, 2025
Kind
B2
Abstract

Systems and methods for detecting text in videos. To address problems with conventional Optical Character Recognition (OCR) systems, the present disclosure provides detection of text for improved OCR. Aspects of the present disclosure can, therefore, be utilized to detect a textual logo in videos, including when the text of the textual logo is clearly visible and when the text is inferred. Thus, examples capture appearance time of a textual logo from a video view perspective. Aspects use a multi-threshold pipeline for detecting video frames including the textual logo. A textual-visual scoring system is additionally used to leverage visual aspects of text in logos. A shot detection system is used to detect inferred text beyond a detected video frame. One or more verification models can be further applied.

Claims (56)

1. A computer-implemented method, comprising:

receiving predicted text derived from processing frames in a video via optical character recognition (OCR);

determining a visual distance of the predicted text to target text;

determining, within a shot, a set of detected frames that include the target text by applying a first distance threshold to the visual distance of the predicted text;

extending the set of detected frames within the shot by applying a second distance threshold to the visual distance of the predicted text within the shot;

determining a sequence of frames in the shot that includes the extended set of detected frames; and

outputting the sequence of frames as results including the target text.

2. The method of claim 1 , wherein extending the set of detected frames further comprises determining a frame in the set of detected frames that has a visual distance corresponding to a representative prediction, wherein the frame is included in the shot.

3. The method of claim 1 , wherein extending the set of detected frames further comprises:

determining boundaries of the detected frames; and

extending the boundaries to a beginning of the shot from an earliest detected frame and from a latest detected frame to an end of the shot.

4. The method of claim 1 , wherein the first distance threshold is stricter than the second distance threshold.

5. The method of claim 1 , wherein determining the visual distance comprises using optimal transport to visually compare characters in the predicted text to the target text.

6. The method of claim 1 , further comprising verifying the sequence of frames prior to outputting the sequence of frames.

7. The method of claim 6 , wherein verifying the sequence of frames comprises:

determining the predicted text is a frequently occurring word, wherein the target text is a textual logo; and

applying a weight to the visual distance to penalize the predicted text.

8. A system, comprising:

a processing system; and

memory storing instructions that, when executed by the processing system, cause the system to:

receive predicted text derived from processing frames in a video via optical character recognition (OCR);

determine a visual distance of the predicted text to target text;

determine, within a shot, a set of detected frames that include the target text by applying a first distance threshold to the visual distance of the predicted text;

extend the set of detected frames within a shot by applying a second distance threshold to the visual distance of the predicted text within the shot;

determine a sequence of frames in the shot that includes the extended set of detected frames; and

output the sequence of frames as results including the target text.

9. The system of claim 8 , wherein extending the set of detected frames comprises determining a frame in the set of detected frames has a visual distance corresponding to a representative prediction, wherein the frame is included in the shot.

10. The system of claim 8 , wherein extending the set of detected frames comprises:

determining boundaries of the detected frames; and

extending the boundaries to a beginning of the shot from a earliest detected frame and from a latest detected frame to an end of the shot.

11. The system of claim 8 , wherein the first distance threshold is stricter than the second distance threshold.

12. The system of claim 8 , wherein determining the visual distance comprises using optimal transport to visually compare characters in the predicted text to the target text.

13. The system of claim 8 , wherein the instructions cause the system to verify the sequence of frames prior to outputting the sequence of frames.

14. The system of claim 13 , wherein:

the target text is a textual logo; and

verifying the sequence of frames comprises:

determining the predicted text is a frequently occurring word; and

applying a weight to the visual distance to penalize the predicted text.

15. A computer readable media comprising instructions, which when executed by a computer, cause the computer to:

receive a plurality of predicted texts derived from processing frames in a video via optical character recognition (OCR);

determine a visual distance of each predicted text of the plurality of predicted texts to target text;

determine a set of detected frames that include the target text by applying a first distance threshold to the visual distance of each predicted text;

extend the set of detected frames within a shot by applying a second distance threshold to the visual distance of each predicted text within the shot, wherein the second distance threshold is less strict than the first distance threshold;

determine a sequence of frames in the shot that includes the extended set of detected frames; and

output the sequence of frames as results including the target text.

16. The computer readable media of claim 15 , wherein extending the set of detected frames comprises determining at least one frame in the set of detected frames has a visual distance corresponding to a representative prediction, wherein the at least one frame is included in the shot.

17. The computer readable media of claim 15 , wherein in extending the set of detected frames, the instructions further cause the computer to:

determine boundaries of the detected frames; and

extend the boundaries to a beginning of the shot from a earliest detected frame and from a latest detected frame to an end of the shot.

18. The computer readable media of claim 15 , wherein in determining the visual distance, the instructions cause the computer to use optimal transport to visually compare characters in the plurality of predicted texts to the target text.

19. The computer readable media of claim 15 , wherein the instructions cause the computer to verify the sequence of frames prior to outputting the sequence of frames.

20. The computer readable media of claim 15 , wherein:

the target text is a textual logo; and

verifying the sequence of frames comprises, for each predicted text in the sequence of frames:

determining the predicted text is a frequently occurring word; and

applying a weight to the visual distance to penalize the predicted text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2022
From: STRUGO, ELIYAHU; HOFFMAN, YONIT
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062153/0942 →
Continuity (1)
Related Publication 20240203146A1 · Jun 20, 2024
References Cited (22)
US 9179061B1 · Kraft et al. · 2015 [cited by applicant]
US 11195148B2 · Tang et al. · 2021 [cited by applicant]
US 20140193075A1 · Pavani · 2014 [cited by applicant]
US 20180109843A1 · Chang et al. · 2018 [cited by applicant]
US 20190236396A1 · Ronen · 2019 [cited by examiner]
US 20230042611A1 · Horner · 2023 [cited by applicant]
US 20230394860A1 · Tao · 2023 [cited by examiner]
CN 113052169A · 2021 [cited by examiner]
CN 113392689A · 2021 [cited by applicant]
CN 113704549A · 2021 [cited by applicant]
CN 113761235A · 2021 [cited by applicant]
KR 101304083B1 · 2013 [cited by applicant]
WO 0225575A2 · 2002 [cited by applicant]
Hua, et al., “Efficient Video Text Recognition Using Multiple Frame Integration”, In Proceedings of International Conference on Image Processing, vol. 2, Sep. 22, 2002, pp. 397-400. [cited by applicant]
“Scale-invariant feature transform”, Retrieved from: https://web.archive.org/web/20220704145527/https://en.wikipedia.org/wiki/Scale-invariant_feature_transform, Jun. 8, 2022, 20 Pages. [cited by applicant]
Kim, et al., “CLIP”, Retrieved from: https://web.archive.org/web/20220625172953/https://github.com/openai/CLIP, Jun. 25, 2022, 6 Pages. [cited by applicant]
Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, In Repository of arXiv:2103.00020v1, Feb. 26, 2021, 48 Pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2023/036894, mailed on Dec. 20, 2023, 15 pages. [cited by applicant]
Phan, et al., “Recognition of Video Text Through Temporal Integration”, 12th International Conference on Document Analysis and Recognition, pp. 589-593, Aug. 25, 2013. [cited by applicant]
Yin, et al., “Text Detection, Tracking and Recognition in Video: A Comprehensive Survey”, IEEE Transactions on Image Processing, vol. 25, Issue No. 6, Jun. 1, 2016, pp. 2752-2773. [cited by applicant]
Notice of Allowance mailed on Jul. 9, 2024, in U.S. Appl. No. 17/804,512, 10 pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/048121”, Mailed Date: Jan. 30, 2023, 10 Pages. [cited by applicant]