IP Library › Granted Patent US 12,731,381
Granted Patent B2
US 12,731,381 · App. 18/398,945 · Granted Sep 8, 2026

Image encoding learning and application

Inventors: Quan Cui (Beijing, CN); Hao Wu (Beijing, CN); Cheng Yang (Beijing, CN)
Assignee: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
G06V10/774G06V10/50G06V10/82G06V10/98G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,381
App. No.
18/398,945
Granted
Sep 8, 2026
Kind
B2
Abstract

Embodiments of the present disclosure provide a solution for image encoding learning and application. A method for image encoding learning comprises: extracting an image feature representation of a sample image using an image encoder to be trained; extracting a text feature representation of a sample text sequence using a text encoder, the sample text sequency being associated with the sample image; generating, using the text encoder, a predicted text sequence based on the text feature representation and the image feature representation; and training the image encoder at least based on a text error between the predicted text sequence and the sample text sequence.

Claims (60)

1 . A method for data encoding learning, comprising:

extracting an image feature representation of a sample image using an image encoder to be trained;

extracting a text feature representation of a sample text sequence using a text encoder, the sample text sequency being associated with the sample image;

generating, using the text encoder, a predicted text sequence based on the text feature representation and the image feature representation; and

training the image encoder at least based on a text error between the predicted text sequence and the sample text sequence;

wherein training the image encoder comprises:

jointly training the image encoder and the text encoder at least based on the text error.

2 . The method of claim 1 , wherein training the image encoder comprises:

generating, using the image encoder, a predicted image based on the image feature representation; and

training the image encoder further based on an image error between the predicted image sequence and the sample image sequence.

3 . The method of claim 1 , wherein the method further comprises:

providing the trained image encoder for a downstream task, wherein the text encoder is discarded.

4 . The method of claim 1 , wherein extracting the image feature representation comprises:

masking at least one image block of the sample image; and

extracting, using the image encoder, the image feature representation from at least one unmasked image block of the sample image.

5 . The method of claim 1 , wherein the sample text sequence comprises a plurality of text units, and wherein extracting the text feature representation comprises: for a given text unit of the plurality of text units,

extracting a text feature representation for the given text unit from the given text unit and at least one text unit preceding the given text unit in the sample text sequence.

6 . The method of claim 5 , wherein generating the predicted text sequence comprises: for a given text unit of the plurality of text units,

determining a predicted text unit from the text feature representation for the given text unit and the image feature representation, the predicted text unit a prediction of a text unit following the given text unit in the sample text sequence.

7 . The method of claim 5 , wherein the given text unit is a last text unit in the sample text sequence, and the predicted text unit is a prediction of an end of the sample text sequence.

8 . The method of claim 1 , wherein the sample text sequence comprises a plurality of text units, and wherein generating the predicted text sequence comprises:

determining self-attention weights for the sample image based on the image feature representation and the text feature representation; and

generating the predicted text sequence based on the image feature representation and the self-attention weights.

9 . The method of claim 1 , wherein the text encoder comprises a transformer block, and the text feature representation is defined as a query feature input to the transformer block, and the image feature representation is defined as a key feature and a value feature input to the transformer block.

10 . The method of claim 1 , further comprising:

extracting, using the trained image encoder, an image feature representation of a target image; and

performing a predetermined vision task for the target image based on the image feature representation of the target image.

11 . An electronic device, comprising:

at least one processing unit; and

at least one memory coupled to the at least one processing unit and storing instructions executable by the at least one processing unit, the instructions, upon execution by the at least one processing unit, causing the device to perform acts comprising:

extracting an image feature representation of a sample image using an image encoder to be trained;

extracting a text feature representation of a sample text sequence using a text encoder, the sample text sequency being associated with the sample image;

generating, using the text encoder, a predicted text sequence based on the text feature representation and the image feature representation; and

training the image encoder at least based on a text error between the predicted text sequence and the sample text sequence;

wherein training the image encoder comprises:

jointly training the image encoder and the text encoder at least based on the text error.

12 . The electronic device of claim 11 , wherein training the image encoder comprises:

generating, using the image encoder, a predicted image based on the image feature representation; and

training the image encoder further based on an image error between the predicted image sequence and the sample image sequence.

13 . The electronic device of claim 11 , wherein the acts further comprises:

providing the trained image encoder for a downstream task, wherein the text encoder is discarded.

14 . The electronic device of claim 11 , wherein extracting the image feature representation comprises:

masking at least one image block of the sample image; and

extracting, using the image encoder, the image feature representation from at least one unmasked image block of the sample image.

15 . The electronic device of claim 11 , wherein the sample text sequence comprises a plurality of text units, and wherein extracting the text feature representation comprises: for a given text unit of the plurality of text units,

extracting a text feature representation for the given text unit from the given text unit and at least one text unit preceding the given text unit in the sample text sequence.

16 . The electronic device of claim 15 , wherein generating the predicted text sequence comprises: for a given text unit of the plurality of text units,

determining a predicted text unit from the text feature representation for the given text unit and the image feature representation, the predicted text unit a prediction of a text unit following the given text unit in the sample text sequence.

17 . The electronic device of claim 15 , wherein the given text unit is a last text unit in the sample text sequence, and the predicted text unit is a prediction of an end of the sample text sequence.

18 . The electronic device of claim 11 , wherein the sample text sequence comprises a plurality of text units, and wherein generating the predicted text sequence comprises:

determining self-attention weights for the sample image based on the image feature representation and the text feature representation; and

generating the predicted text sequence based on the image feature representation and the self-attention weights.

19 . The electronic device of claim 11 , wherein the text encoder comprises a transformer block, and the text feature representation is defined as a query feature input to the transformer block, and the image feature representation is defined as a key feature and a value feature input to the transformer block.

20 . A non-transitory computer-readable storage medium, having a computer program stored thereon which, upon execution by a processor, implements acts comprising:

extracting an image feature representation of a sample image using an image encoder to be trained;

extracting a text feature representation of a sample text sequence using a text encoder, the sample text sequency being associated with the sample image;

generating, using the text encoder, a predicted text sequence based on the text feature representation and the image feature representation; and

training the image encoder at least based on a text error between the predicted text sequence and the sample text sequence;

wherein training the image encoder comprises:

jointly training the image encoder and the text encoder at least based on the text error.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2026
From: YANG, CHENG
To: CHENGDU GUANGHEXINHAO TECHNOLOGY CO., LTD.
Reel/Frame 075509/0114 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2026
From: CHENGDU GUANGHEXINHAO TECHNOLOGY CO., LTD.; SHANGHAI SUIXUNTONG ELECTRONIC TECHNOLOGY CO., LTD.; SHANGHAI GEWU ZHIYUAN NETWORK TECHNOLOGY CO., LTD.
To: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 075509/0188 →
Priority Claims (1)
CN 202310032610.9 · Jan 10, 2023 · national
Continuity (1)
Related Publication 20240185578A1 · Jun 6, 2024
References Cited (29)
US 11714849B2 · Zhou · 2023 [cited by examiner]
US 12333837B2 · He · 2025 [cited by examiner]
US 12424010B2 · Lv · 2025 [cited by examiner]
US 20200117854A1 · Lu · 2020 [cited by examiner]
US 20210203997A1 · Veselov · 2021 [cited by examiner]
US 20210248309A1 · Zhang · 2021 [cited by examiner]
US 20210358177A1 · Park · 2021 [cited by examiner]
US 20220172080A1 · Chaudhury · 2022 [cited by examiner]
US 20220284321A1 · Yuan · 2022 [cited by examiner]
US 20220415071A1 · Zhang et al. · 2022 [cited by applicant]
US 20230019211A1 · Wang · 2023 [cited by examiner]
US 20230162490A1 · Zhang · 2023 [cited by examiner]
US 20230178076A1 · Abramson · 2023 [cited by examiner]
US 20230245418A1 · Zhang · 2023 [cited by examiner]
US 20240212327A1 · Geng · 2024 [cited by examiner]
CN 108959322A · 2018 [cited by applicant]
CN 108228686B · 2021 [cited by applicant]
CN 112560652A · 2021 [cited by applicant]
CN 113591902A · 2021 [cited by applicant]
CN 113792112A · 2021 [cited by applicant]
CN 113836333A · 2021 [cited by applicant]
CN 114494011A · 2022 [cited by applicant]
CN 114626392A · 2022 [cited by applicant]
CN 115062782A · 2022 [cited by applicant]
CN 116503636A · 2023 [cited by examiner]
CN 114998673B · 2023 [cited by examiner]
WO 2022155994A1 · 2022 [cited by applicant]
Office Action for Chinese Patent Application No. 202310032610.9, mailed on Dec. 23, 2025, 17 pages. [cited by applicant]
Office Action for Chinese Patent Application No. 202310032610.9, mailed on Jun. 16, 2026, 21 pages. [cited by applicant]