IP Library Granted Patent US 12,530,873
Granted Patent B2
US 12,530,873 · App. 18/162,498 · Granted Jan 20, 2026

Training method

Inventors: Shota Isobe (Kyoto, JP); Yutaka Yoshihama (Osaka, JP)
Assignee: Panasonic Automotive Systems Co., Ltd.
G06V10/774G06V10/764G06V10/776G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,873
App. No.
18/162,498
Granted
Jan 20, 2026
Kind
B2
Abstract

A training method according to the present disclosure is a training method performed by a computer to train a neural network model including a first network branch for representation learning with use of supervised contrastive learning, and a second network branch for training of computer vision tasks including a classification task and a task other than the classification task. The training method includes: performing label processing for converting labels of M image data items into labels applicable to the representation learning, as labels of the computer vision tasks; and training an encoder network model and a first model with use of a first loss function for use in the supervised contrastive learning, the labels of the M image data items on which the label processing has been performed, and embedding vectors of the M image data items.

Claims (34)

1 . A training method performed by a computer to train a neural network model that includes a first network branch for representation learning with use of supervised contrastive learning, and a second network branch for training of computer vision tasks that include a classification task and a task other than the classification task, the neural network model including: an encoder network model shared by the first network branch and the second network branch; a first model included in only the first network branch; and a second model included in only the second network branch, the training method comprising:

obtaining N image data that is one or more image data items and one or more labels in one-to-one association with the N image data from a data set that includes one or more preprovided image data items and one or more preprovided labels, N denoting an integer greater than or equal to 1;

performing data augmentation processing on the N image data obtained and the one or more labels obtained, which are in one-to-one association with the N image data, to obtain M image data items and labels in one-to-one association with the M image data items, M denoting an integer multiple of N;

extracting, by the encoder network model, feature representations of the M image data items from the M image data items;

projecting, by the first model, the feature representations of the M image data items that are extracted, onto embedding vectors for use in the supervised contrastive learning;

performing label processing for converting the labels of the M image data items into labels applicable to the representation learning, as labels of the computer vision tasks;

training the encoder network model and the first model with use of a first loss function for use in the supervised contrastive learning, the labels of the M image data items on which the label processing has been performed, and the embedding vectors of the M image data items;

obtaining the M image data items resulting from the data augmentation processing;

extracting, by the encoder network model, feature representations of the M image data items from the M image data items obtained;

inferring, by the second model, labels of the M image data items from the feature representations of the M image data items that are extracted; and

training the encoder network model and the second model with use of a second loss function for use in the training, the labels of the M image data items that are inferred, and the labels of the M image data items,

wherein the training of the encoder network model and the first model and the training of the encoder network model and the second model are simultaneously performed, and

wherein in the label processing, by

(i) converting, as the labels of the computer vision tasks, the labels of the M image data items into applicable representations in which a value greater than or equal to 2 is allowed for a value of each of dimensions, and

(ii) applying a step function that converts a value greater than β for each of the dimensions in the applicable representations into 1, the applicable representations being applicable to the representation learning, β denoting an arbitrary number,

the labels of the M image data items are converted into the applicable representations in which the value for each of the dimensions is 0 or 1.

2 . The training method according to claim 1 ,

wherein in the training of the encoder network model and the first model, when the M image data items include two different image data items that are obtained by performing the data augmentation processing on a same image data item,

supervised contrastive loss is calculated by using the first loss function.

3 . The training method according to claim 2 ,

wherein in the label processing, when the applicable representations resulting from converting the labels of the M image data items include a value of 1 for each of two or more of the dimensions,

the value of 1 for each of the two or more of the dimensions is further converted into 0.

4 . The training method according to claim 1 ,

wherein in the training of the encoder network model and the first model,

a loss based on a vector similarity is calculated by using the first loss function.

5 . The training method according to claim 1 ,

wherein the first network branch:

includes a third model;

causes the third model to output a third embedding vector obtained by the third model predicting a second embedding vector from a first embedding vector, the first embedding vector being one of embedding vectors of two image data items that are output by the first model, the second embedding vector being a remaining one of the embedding vectors of the two image data items;

performs the label processing for converting labels of the two image data items into one-hot representations in each of which a class dimension is used, the class dimension having a dimension count that is a class count of a class label used in the classification task; and

trains the encoder network model, the first model, and the third model with use of the first loss function for use in the supervised contrastive learning, the labels of the two image data items on which the label processing has been performed, the second embedding vector, and the third embedding vector.

6 . The training method according to claim 5 ,

wherein in the training of the encoder network model, the first model, and the third model,

a loss based on a cosine similarity is calculated by using the first loss function.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2024
From: PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO., LTD.
To: PANASONIC AUTOMOTIVE SYSTEMS CO., LTD.
Reel/Frame 066709/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2023
From: ISOBE, SHOTA; YOSHIHAMA, YUTAKA
To: PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO., LTD.
Reel/Frame 064380/0832 →
Priority Claims (1)
JP 2022-029757 · Feb 28, 2022 · national
Continuity (1)
Related Publication 20230274533A1 · Aug 31, 2023
References Cited (17)
US 12314861B2 · Li · 2025 [cited by examiner]
US 20170228641A1 · Sohn · 2017 [cited by applicant]
US 20180336683A1 · Feng · 2018 [cited by examiner]
US 20190025848A1 · Kolouri · 2019 [cited by examiner]
US 20210357698A1 · Murasaki et al. · 2021 [cited by applicant]
US 20220164600A1 · Cheng · 2022 [cited by examiner]
US 20220269946A1 · Zhou · 2022 [cited by examiner]
US 20230153629A1 · Krishnan · 2023 [cited by examiner]
US 20230169332A1 · Karthik · 2023 [cited by examiner]
JP 2019509551A · 2019 [cited by applicant]
JP 2020047055A · 2020 [cited by applicant]
WO WO2017136060A1 · 2017 [cited by applicant]
WO WO2018015080A1 · 2018 [cited by examiner]
WO WO2021216310A1 · 2021 [cited by examiner]
Wang, Peng, et al. “Contrastive learning based hybrid networks for long-tailed image classification.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021. (Year: 2021). [cited by examiner]
Gunel et al., “Supervised Contrastive Learning for Pre-Trained Language Model Fine-Tuning,” International Conference on Learning Representations, Vienna, Austria, May 4, 2021. (15 pages). [cited by applicant]
Wang et al., “Contrastive Learning based Hybrid Networks for Long-Tailed Image Classification,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, Jun. 20-25, 2021, pp. 943-9… [cited by applicant]