IP Library Granted Patent US 12,475,680
Granted Patent B2
US 12,475,680 · App. 18/116,602 · Granted Nov 18, 2025

Method of training image representation model

Inventors: Minyoung Mun (Suwon-si, KR); Seijoon Kim (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06V10/761G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,680
App. No.
18/116,602
Granted
Nov 18, 2025
Kind
B2
Abstract

A method generates an anchor image embedding vector for an anchor image using an image representation model, determine first similarities between the anchor image and negative samples of the anchor image using first image embedding vectors for the negative samples and the generated anchor image embedding vector, determine second similarities between the anchor image and positive samples of the anchor image using second image embedding vectors for the positive samples and the generated anchor image embedding vector, obtain one of a vector corresponding to a label of the anchor image and third similarities between the label of the anchor image and labels of the negative samples, determine a loss value for the anchor image based on the determined first similarities, and the determined second similarities, and one of the obtained third similarities and a fourth similarity.

Claims (59)

1 . A training method performed by a computing apparatus, the training method comprising:

generating an anchor image embedding vector for an anchor image using an image representation model;

determining first similarities between the anchor image and negative samples of the anchor image using first image embedding vectors for the negative samples and the generated anchor image embedding vector;

determining second similarities between the anchor image and positive samples of the anchor image using second image embedding vectors for the positive samples and the generated anchor image embedding vector;

obtaining one of a vector corresponding to a label of the anchor image and third similarities between the label of the anchor image and labels of the negative samples;

determining a loss value for the anchor image based on (i) the determined first similarities, (ii) the determined second similarities, and (iii) one of the obtained third similarities and a fourth similarity, wherein the fourth similarity is a similarity between the obtained vector and the generated anchor image embedding vector; and

updating weights of the image representation model based on the determined loss value.

2 . The training method of claim 1 , wherein

the positive samples and the anchor image belong to a same class, and the negative samples do not belong to the class.

3 . The training method of claim 1 , wherein the determining of the loss value comprises:

applying the obtained third similarities as weights to each of the determined first similarities;

calculating normalized values for the obtained third similarities; and

determining the loss value using a result of applying the obtained third similarities as weights to each of the determined first similarities, the calculated normalized values, and the determined second similarities.

4 . The training method of claim 1 , further comprising:

determining similarities of pairings of labels of respective images in a training data set using an embedding model;

generating a first dictionary to store the similarities for the pairings;

forming a batch of images extracted from the training data set;

forming an image set corresponding to the batch by performing augmentation on the images in the formed batch; and

retrieving, from the first dictionary, similarities for respective pairings of labels of the batch.

5 . The training method of claim 4 , wherein

the obtaining comprises obtaining the third similarities from among the retrieved similarities.

6 . The training method of claim 1 , wherein

the third similarities are similarities between the vector corresponding to the label of the anchor image and vectors corresponding to the labels of the negative samples, and

the vector corresponding to the label of the anchor image and the vectors corresponding to the labels of the negative samples are generated by an embedding model.

7 . The training method of claim 1 , wherein the determining of the loss value comprises:

determining an initial loss value using the determined first similarities and the determined second similarities;

applying a first weight to the determined initial loss value;

applying a second weight to the fourth similarity; and

determining the loss value by subtracting the fourth similarity to which the second weight is applied from the initial loss value to which the first weight is applied.

8 . The training method of claim 7 , wherein the sum of the first weight and the second weight is 1.

9 . The training method of claim 1 , further comprising:

generating vectors respectively corresponding to labels of a training data set using an embedding model;

generating a second dictionary to store the generated vectors;

forming a batch by extracting images from the training data set;

forming an image set corresponding to the batch by performing augmentation on the images in the formed batch; and

retrieving vectors corresponding to labels of the batch from the second dictionary.

10 . The training method of claim 9 , wherein

the obtaining comprises obtaining the vector corresponding to the label of the anchor image from among the retrieved vectors.

11 . A computing apparatus, comprising:

a memory configured to store one or more instructions; and

a processor configured to execute the stored instructions,

wherein, when the instructions are executed, the processor is configured to:

generate an anchor image embedding vector for an anchor image using an image representation model,

determine first similarities between the anchor image and negative samples of the anchor image using first image embedding vectors for the negative samples and the generated anchor image embedding vector,

determine second similarities between the anchor image and positive samples of the anchor image using second image embedding vectors for the positive samples and the generated anchor image embedding vector,

obtain one of a vector corresponding to a label of the anchor image and third similarities between the label of the anchor image and labels of the negative samples,

determine a loss value for the anchor image based on (i) the determined first similarities, (iii) the determined second similarities, and (iii) one of the obtained third similarities and a fourth similarity, wherein the fourth similarity is a similarity between the obtained vector and the generated anchor image embedding vector, and

update weights of the image representation model based on the determined loss value.

12 . The computing apparatus of claim 11 , wherein the positive samples and the anchor image belong to a same class, and the negative samples and the anchor image do not belong to the class.

13 . The computing apparatus of claim 11 , wherein the processor is configured to apply the obtained third similarities as weights to each of the determined first similarities, calculate normalized values for the obtained third similarities, and determine the loss value using a result of applying the obtained third similarities as weights to each of the determined first similarities, the calculated normalized values, and the determined second similarities.

14 . The computing apparatus of claim 11 , wherein the processor is configured to determine similarities of pairings of labels of respective images in a training data set using an embedding model, generate a first dictionary to store the similarities for the pairings, form a batch of images extracted from the training data set, form an image set corresponding to the batch by performing augmentation on the images in the formed batch, and retrieve, from the first dictionary, similarities for respective pairings of labels of the batch.

15 . The computing apparatus of claim 14 , wherein the processor is configured to obtain the third similarities from among the retrieved similarities.

16 . The computing apparatus of claim 11 , wherein

the third similarities are similarities between the vector corresponding to the label of the anchor image and vectors corresponding to the labels of the negative samples, and

the vector corresponding to the label of the anchor image and the vectors corresponding to the labels of the negative samples are generated by an embedding model.

17 . The computing apparatus of claim 11 , wherein the processor is configured to determine an initial loss value using the determined first similarities and the determined second similarities, apply a first weight to the determined initial loss value, apply a second weight to the fourth similarity, and determine the loss value by subtracting the fourth similarity to which the second weight is applied from the initial loss value to which the first weight is applied.

18 . The computing apparatus of claim 17 , wherein the sum of the first weight and the second weight is 1.

19 . The computing apparatus of claim 11 , wherein the processor is configured to generate vectors respectively corresponding to labels of a training data set using an embedding model, generate a second dictionary to store the generated vectors, form a batch by extracting images from the training data set, form an image set corresponding to the formed batch by performing augmentation on the images in the formed batch, and retrieve vectors corresponding to labels of the formed batch from the second dictionary.

20 . The computing apparatus of claim 19 , wherein the processor is configured to obtain the vector corresponding to the label of the anchor image from among the retrieved vectors.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2026
From: GHANG, WHAN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 074799/0581 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2023
From: MUN, MINYOUNG; KIM, SEIJOON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062860/0050 →
Priority Claims (1)
KR 10-2022-0111092 · Sep 2, 2022 · national
Continuity (1)
Related Publication 20240078785A1 · Mar 7, 2024
References Cited (14)
US 10102454B2 · Merler · 2018 [cited by examiner]
US 10832062B1 · Evans · 2020 [cited by examiner]
US 11347975B2 · Krishnan · 2022 [cited by examiner]
US 11354778B2 · Chen · 2022 [cited by examiner]
US 20200090039A1 · Song · 2020 [cited by examiner]
US 20220147838A1 · Gu · 2022 [cited by examiner]
US 20220156593A1 · Liu · 2022 [cited by examiner]
US 20230011503A1 · Duke · 2023 [cited by examiner]
Chen, Ting, et al. “A simple framework for contrastive learning of visual representations.” International conference on machine learning. PmLR. (Year: 2020). [cited by examiner]
Robinson, Joshua, et al. “Contrastive learning with hard negative samples.” arXiv preprint arXiv:2010.04592 (2020). (Year: 2020). [cited by examiner]
Khosla, Prannay, et al. “Supervised contrastive learning.” Advances in neural information processing systems 33. (Year: 2022). [cited by examiner]
Chen, Ting, et al. “A simple framework for contrastive learning of visual representations.” [cited by applicant]
Robinson, Joshua, et al. “Contrastive learning with hard negative samples.” arXiv preprint arXiv:2010.04592 vol. 2. (2020). pp 1-28. [cited by applicant]
Khosla, Prannay, et al. “Supervised contrastive learning.” [cited by applicant]