IP Library › Granted Patent US 12,254,707
Granted Patent B2
US 12,254,707 · App. 17/955,285 · Granted Mar 18, 2025

Pre-training for scene text detection

Inventors: Chuhui Xue (Singapore, SG); Wenqing Zhang (Singapore, SG); Yu Hao (Beijing, CN); Song Bai (Singapore, SG)
Assignees: LEMON INC.; BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
G06V20/63G06V30/18G06V30/19093G06V30/19147
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,707
App. No.
17/955,285
Granted
Mar 18, 2025
Kind
B2
Abstract

Embodiments of the present disclosure relate to a method, device and computer readable storage medium of scene text detection. In the method, a first visual representation of a first image is generated with an image encoding process. A first textual representation of a first text unit in the first image is generated with a text encoding process based on a first plurality of symbols obtained by masking a first symbol of a plurality of symbols in the first text unit. A first prediction of the masked first symbol is determined with a decoding process based on the first visual and textual representations. At least the image encoding process is updating according to at least a first training objective to increase at least similarity of the first prediction and the masked first symbol.

Claims (60)

1. A method of scene text detection, comprising:

generating, with an image encoding process, a first visual representation of a first image;

generating, with a text encoding process, based on a first plurality of symbols, a first textual representation of a first text unit in the first image, the first plurality of symbols obtained by masking a first symbol of a plurality of symbols in the first text unit;

determining, with a decoding process, a first prediction of the masked first symbol based on the first visual and textual representations; and

updating at least the image encoding process according to at least a first training objective to increase at least similarity of the first prediction and the masked first symbol.

2. The method of claim 1 , wherein generating the first textual representation comprises:

extracting a plurality of symbol representations from the first plurality of symbols; and

generating the first textual representation by aggregating the plurality of symbol representations.

3. The method of claim 1 , wherein the first text unit comprises a part of a text in the first image.

4. The method of claim 1 , further comprising:

generating, with the text encoding process, based on a second plurality of symbols, a second textual representation of a second text unit in the first image, the second plurality of symbols obtained by masking a second symbol of a plurality of symbols in the second text unit; and

determining, with the decoding process, a second prediction of the masked second symbol based on the first visual representation and the second textual representations,

where the first training objective is further to increase similarity of the second prediction and the masked second symbol.

5. The method of claim 1 , wherein updating the image encoding process comprises:

updating the image encoding process according to further a second training objective to increase at least similarity of the first visual and textual representations.

6. The method of claim 5 , further comprising:

generating, with the image encoding process, a second visual representation of a second image; and

generating, with the text encoding process, based on a third plurality of symbols, a third textual representation of a third text unit in the second image, the third plurality of symbols obtained by masking a third symbol of a plurality of symbols in the third text unit,

wherein the second training objective is further to increase similarity of the second visual representation and the third textual representation.

7. The method of claim 6 , wherein the second training objective is further to decrease similarity of the second visual representation and the first textual representation and similarity of the first visual representation and the third textual representation.

8. The method of claim 1 , wherein updating at least the image encoding process comprises:

updating the image encoding process along with the text encoding process.

9. An electronic device, comprising:

at least one processor; and

at least one memory communicatively coupled to the at least one processor and comprising computer-readable instructions that upon execution by the at least one processor cause the at least one processor to perform acts comprising:

generating, with an image encoding process, a first visual representation of a first image;

generating, with a text encoding process, based on a first plurality of symbols, a first textual representation of a first text unit in the first image, the first plurality of symbols obtained by masking a first symbol of a plurality of symbols in the first text unit;

determining, with a decoding process, a first prediction of the masked first symbol based on the first visual and textual representations; and

updating at least the image encoding process according to at least a first training objective to increase at least similarity of the first prediction and the masked first symbol.

10. The electronic device of claim 9 , wherein generating the first textual representation comprises:

extracting a plurality of symbol representations from the first plurality of symbols; and

generating the first textual representation by aggregating the plurality of symbol representations.

11. The electronic device of claim 9 , wherein the first text unit comprises a part of a text in the first image.

12. The electronic device of claim 9 , wherein the acts further comprise:

generating, with the text encoding process, based on a second plurality of symbols, a second textual representation of a second text unit in the first image, the second plurality of symbols obtained by masking a second symbol of a plurality of symbols in the second text unit; and

determining, with the decoding process, a second prediction of the masked second symbol based on the first visual representation and the second textual representations,

where the first training objective is further to increase similarity of the second prediction and the masked second symbol.

13. The electronic device of claim 9 , wherein updating the image encoding process comprises:

updating the image encoding process according to further a second training objective to increase at least similarity of the first visual and textual representations.

14. The electronic device of claim 13 , wherein the acts further comprise:

generating, with the image encoding process, a second visual representation of a second image; and

generating, with the text encoding process, based on a third plurality of symbols, a third textual representation of a third text unit in the second image, the third plurality of symbols obtained by masking a third symbol of a plurality of symbols in the third text unit,

wherein the second training objective is further to increase similarity of the second visual representation and the third textual representation.

15. The electronic device of claim 14 , wherein the second training objective is further to decrease similarity of the second visual representation and the first textual representation and similarity of the first visual representation and the third textual representation.

16. The electronic device of claim 9 , wherein updating at least the image encoding process comprises:

updating the image encoding process along with the text encoding process.

17. A non-transitory computer-readable storage medium, storing computer-readable instructions that upon execution by a computing device cause the computing device to perform acts comprising:

generating, with an image encoding process, a first visual representation of a first image;

generating, with a text encoding process, based on a first plurality of symbols, a first textual representation of a first text unit in the first image, the first plurality of symbols obtained by masking a first symbol of a plurality of symbols in the first text unit;

determining, with a decoding process, a first prediction of the masked first symbol based on the first visual and textual representations; and

updating at least the image encoding process according to at least a first training objective to increase at least similarity of the first prediction and the masked first symbol.

18. The computer-readable storage medium of claim 17 , wherein generating the first textual representation comprises:

extracting a plurality of symbol representations from the first plurality of symbols; and

generating the first textual representation by aggregating the plurality of symbol representations.

19. The computer-readable storage medium of claim 17 , wherein updating the image encoding process comprises:

updating the image encoding process according to further a second training objective to increase at least similarity of the first visual and textual representations.

20. The computer-readable storage medium of claim 19 , wherein the acts further comprise:

generating, with the image encoding process, a second visual representation of a second image; and

generating, with the text encoding process, based on a third plurality of symbols, a third textual representation of a third text unit in the second image, the third plurality of symbols obtained by masking a third symbol of a plurality of symbols in the third text unit,

wherein the second training objective is further to increase similarity of the second visual representation and the third textual representation and to decrease similarity of the second visual representation and the first textual representation and similarity of the first visual representation and the third textual representation.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2025
From: HAO, YU
To: SHANGHAI INFINITE MEMORY SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 070109/0995 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2025
From: TIKTOK PTE. LTD.
To: LEMON INC.
Reel/Frame 070110/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2025
From: TIKTOK PTE. LTD.
To: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 070110/0011 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2025
From: SHANGHAI INFINITE MEMORY SCIENCE AND TECHNOLOGY CO., LTD.
To: LEMON INC.
Reel/Frame 070110/0017 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2025
From: SHANGHAI INFINITE MEMORY SCIENCE AND TECHNOLOGY CO., LTD.
To: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 070110/0021 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2025
From: XUE, CHUHUI; ZHANG, WENQING; BAI, SONG
To: TIKTOK PTE. LTD.
Reel/Frame 070109/0986 →
Continuity (1)
Related Publication 20240119743A1 · Apr 11, 2024
References Cited (5)
US 20150055854A1 · Marchesotti · 2015 [cited by examiner]
US 20160210532A1 · Soldevila · 2016 [cited by examiner]
CN 114565751A · 2022 [cited by applicant]
CN 111858882B · 2022 [cited by applicant]
CN 114998678A · 2022 [cited by applicant]
Cited By (1)
US 12,579,820