IP Library Granted Patent US 12,475,728
Granted Patent B2
US 12,475,728 · App. 18/001,031 · Granted Nov 18, 2025

Label identification method and apparatus, device, and medium

Inventors: Jinlai Liu (Beijing, CN); Bin Wen (Beijing, CN); Changhu Wang (Beijing, CN)
Assignee: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
G06V20/70G06F40/30G06V10/7715G06V10/806G06V20/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,728
App. No.
18/001,031
Granted
Nov 18, 2025
Kind
B2
Abstract

Provided are a label identification method and apparatus, a device, and a medium. The method includes: obtaining a target feature of a first image, in which the target feature characterizes a visual feature of the first image and a word feature of at least one label; and identifying a label of the first image from the at least one label based on the target feature. By characterizing the visual feature of the first image and the target feature of the word feature of the at least one label, the label of the first image is identified from the at least one label. Thus, identification accuracy of the label can be improved.

Claims (54)

1 . A label identification method, comprising:

obtaining a target feature of a first image, wherein the target feature characterizes a visual feature of the first image and a word feature of at least one label; and

identifying a label of the first image from the at least one label based on the target feature; wherein:

at least one visual feature vector corresponding to at least one image channel of the first image is used to generate the visual feature of the first image; and

at least one word feature vector corresponding to the at least one label is used to generate the word feature of the at least one label.

2 . The method according to claim 1 , wherein:

the at least one word feature vector is used to direct the at least one visual feature vector to generate the visual feature of the first image; and

the at least one visual feature vector is used to direct the at least one word feature vector to generate the word feature of the at least one label.

3 . The method according to claim 1 , wherein a number of image channels in the at least one image channel is equal to a number of labels in the at least one label.

4 . The method according to claim 1 , wherein a number of dimensions of each of the at least one visual feature vector is equal to a number of dimensions of each of the at least one word feature vector.

5 . The method according to claim 1 , wherein said obtaining the target feature of the first image comprises:

determining a first matrix based on the at least one visual feature vector;

determining a second matrix based on the at least one word feature vector; and

determining the target feature based on the first matrix and the second matrix.

6 . The method according to claim 5 , wherein:

the first matrix is a product of a first weight and a matrix defined by the at least one visual feature vector; and

the second matrix is a product of a second weight and a matrix defined by the at least one word vector,

wherein the first weight is a weight determined based on a score of the at least one word feature vector; and

the second weight is a weight determined based on a score of the at least one visual feature vector.

7 . The method according to claim 1 , wherein said identifying the label of the first image from the at least one label based on the target feature comprises:

mapping the target feature to a probability distribution feature, wherein a number of dimensions of the probability distribution feature is equal to a number of dimensions of the at least one label, and a value at a position in the probability distribution feature represents a confidence of a label corresponding to the value at the position;

determining a first numerical value greater than a predetermined threshold in the probability distribution feature; and

determining a label corresponding to the first numerical value as the label of the first image.

8 . The method according to claim 7 , wherein the semantic relationship comprises at least one of a generic-specific semantic relationship, a similar semantic relationship, or an opposite semantic relationship.

9 . The method according to claim 1 , wherein a semantic relationship applies between the labels of the at least one label.

10 . The method according to claim 1 , wherein the label of the first image comprises the at least one label.

11 . The method according to claim 1 , wherein said obtaining the target feature of the first image comprises:

obtaining the target feature using a prediction model with the visual feature of the first image as an input.

12 . An electronic device, comprising:

a memory having a computer program stored thereon; and

a processor configured to call and execute the computer program stored in the memory to:

obtain a target feature of a first image, wherein the target feature characterizes a visual feature of the first image and a word feature of at least one label; and

identify a label of the first image from the at least one label based on the target feature; wherein:

at least one visual feature vector corresponding to at least one image channel of the first image is used to generate the visual feature of the first image; and

at least one word feature vector corresponding to the at least one label is used to generate the word feature of the at least one label.

13 . A computer-readable storage medium, having a computer program stored thereon, wherein the computer program causes a computer to:

obtain a target feature of a first image, wherein the target feature characterizes a visual feature of the first image and a word feature of at least one label; and

identify a label of the first image from the at least one label based on the target feature; wherein:

at least one visual feature vector corresponding to at least one image channel of the first image is used to generate the visual feature of the first image; and

at least one word feature vector corresponding to the at least one label is used to generate the word feature of the at least one label.

14 . The electronic device according to claim 13 , wherein:

the at least one word feature vector is used to direct the at least one visual feature vector to generate the visual feature of the first image; and

the at least one visual feature vector is used to direct the at least one word feature vector to generate the word feature of the at least one label.

15 . The electronic device according to claim 13 , wherein a number of image channels in the at least one image channel is equal to a number of labels in the at least one label.

16 . The electronic device according to claim 13 , wherein a number of dimensions of each of the at least one visual feature vector is equal to a number of dimensions of each of the at least one word feature vector.

17 . The electronic device according to claim 13 , wherein said obtaining the target feature of the first image comprises:

determining a first matrix based on the at least one visual feature vector;

determining a second matrix based on the at least one word feature vector; and

determining the target feature based on the first matrix and the second matrix.

18 . The electronic device according to claim 17 , wherein:

the first matrix is a product of a first weight and a matrix defined by the at least one visual feature vector; and

the second matrix is a product of a second weight and a matrix defined by the at least one word vector,

wherein the first weight is a weight determined based on a score of the at least one word feature vector, and

the second weight is a weight determined based on a score of the at least one visual feature vector.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2023
From: LIU, JINLAI
To: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO. LTD.
Reel/Frame 064628/0919 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2023
From: WEN, BIN; WANG, CHANGHU
To: DOUYIN VISION CO., LTD.
Reel/Frame 064628/0947 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2023
From: DOUYIN VISION CO., LTD.
To: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO. LTD.
Reel/Frame 064628/0957 →
Priority Claims (1)
CN 202011086888.7 · Oct 12, 2020 · national
Continuity (1)
Related Publication 20230230400A1 · Jul 20, 2023
References Cited (18)
US 11531863B1 · Zweig · 2022 [cited by examiner]
US 11694460B1 · Luo · 2023 [cited by examiner]
US 20170124116A1 · League · 2017 [cited by applicant]
US 20180373952A1 · Bui · 2018 [cited by examiner]
US 20220117218A1 · Sibley · 2022 [cited by examiner]
US 20230222821A1 · Delp, III · 2023 [cited by examiner]
US 20230360421A1 · Luo · 2023 [cited by examiner]
US 20240054803A1 · Negi · 2024 [cited by examiner]
US 20240265547A1 · Cui · 2024 [cited by examiner]
CN 110472642A · 2019 [cited by applicant]
CN 111090763A · 2020 [cited by applicant]
CN 111177444A · 2020 [cited by applicant]
CN 111507355A · 2020 [cited by applicant]
CN 112347290A · 2021 [cited by applicant]
Second Office Action issued Nov. 24, 2023 in Chinese Application No. 202011086888.7, with English Translation (10 pages). [cited by applicant]
Written Opinion for International Application No. PCT/CN2021/117459, mailed Nov. 26, 2021, 8 Pages. [cited by applicant]
Search Report mailed Nov. 26, 2021 for PCT Application No. PCT/CN2021/117459 English Translation (6 pages). [cited by applicant]
First Office Action issued May 5, 2023 in Chinese Application No. 202011086888.7, with English Translation (14 pages). [cited by applicant]