IP Library › Granted Patent US 12,579,820
Granted Patent B2
US 12,579,820 · App. 18/213,980 · Granted Mar 17, 2026

Learning apparatus and learning method

Inventors: Naoki Hosomi (Wako, JP); Teruhisa Misu (San Jose, CA); Shumpei Hatanaka (Yokohama, JP); Wei Yang (Yokohama, JP); Komei Sugiura (Yokohama, JP)
Assignees: HONDA MOTOR CO., LTD.; KEIO UNIVERSITY
G06V20/58G06V10/806
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,820
App. No.
18/213,980
Granted
Mar 17, 2026
Kind
B2
Abstract

A learning apparatus for performing machine learning includes an acquisition unit configured to acquire teaching data including input data and correct answer data, the input data including an input image that contains a reference object and an input text that relatively designates a target position by referring to the reference object; a generation unit configured to input the input data to a model to generate output data for specifying the target position, a reference position that is a position of the reference object, and a positional relationship of the target position with respect to the reference position; and an update unit configured to update a parameter of the model to reduce a loss obtained by inputting the output data and the correct answer data to a loss function. The loss function is based on at least two errors of a first error between the target position specified by the output data and the target position specified by the correct answer data, a second error between the reference position specified by the output data and the reference position specified by the correct answer data, and a third error between the positional relationship specified by the output data and the positional relationship specified by the correct answer data.

Claims (46)

1 . A learning apparatus for performing machine learning, comprising a memory storing instructions and a processor configured to execute the instructions to:

acquire teaching data including input data and correct answer data, the input data including an input image that contains a reference object and an input text that relatively designates a target position by referring to the reference object;

input the input data to a model to generate output data for specifying the target position, a reference position that is a position of the reference object, and a positional relationship of the target position with respect to the reference position; and

update a parameter of the model to reduce a loss obtained by inputting the output data and the correct answer data to a loss function,

wherein the loss function is based on at least two errors of

a first error between the target position specified by the output data and the target position specified by the correct answer data,

a second error between the reference position specified by the output data and the reference position specified by the correct answer data, and

a third error between the positional relationship specified by the output data and the positional relationship specified by the correct answer data, and

wherein the model includes

an image encoding layer for encoding the input image, and

a text encoding layer for encoding the input text,

a part of features determined by the text encoding layer is input to the image encoding layer, and

a part of features determined by the image encoding layer is input to the text encoding layer.

2 . The learning apparatus according to claim 1 , wherein the input image includes an image imaged by a camera of a vehicle.

3 . The learning apparatus according to claim 1 , wherein the input text is expressed by a natural language.

4 . The learning apparatus according to claim 1 , wherein the loss function is based on at least the first error.

5 . The learning apparatus according to claim 1 , wherein the loss function is based on all of the first error, the second error, and the third error.

6 . The learning apparatus according to claim 1 , wherein the image encoding layer and the text encoding layer have an identical layer structure.

7 . The learning apparatus according to claim 1 ,

wherein the machine learning of the model is first machine learning, and

a parameter of the model at a start point in time of training is a parameter of the image encoding layer determined by second machine learning with the input image as input data and a position of the reference object as correct answer data.

8 . A learning apparatus for performing machine learning, comprising a memory storing instructions and a processor configured to execute the instructions to:

acquire teaching data including input data and correct answer data, the input data including an input image that contains a reference object and an input text that relatively designates a target position by referring to the reference object;

generate output data for specifying the target position, a reference position that is a position of the reference object, and a positional relationship of the target position with respect to the reference position by inputting the input data to a model; and

update a parameter of the model to reduce a loss obtained by inputting the output data and the correct answer data to a loss function,

wherein the loss function is based on an error between the positional relationship specified by the output data and the positional relationship specified by the correct answer data, and

wherein the model includes

an image encoding layer for encoding the input image, and

a text encoding layer for encoding the input text,

a part of features determined by the text encoding layer is input to the image encoding layer, and

a part of features determined by the image encoding layer is input to the text encoding layer.

9 . A non-transitory computer readable storage medium for storing a program causing a computer to function as the learning apparatus according to claim 1 .

10 . A learning method of performing machine learning, comprising:

acquiring teaching data including input data and correct answer data, the input data including an input image that contains a reference object and an input text that relatively designates a target position by referring to the reference object;

inputting the input data to a model to generate output data for specifying the target position, a reference position that is a position of the reference object, and a positional relationship of the target position with respect to the reference position; and

updating a parameter of the model to reduce a loss obtained by inputting the output data and the correct answer data to a loss function,

wherein the loss function is based on at least two errors of

a first error between the target position specified by the output data and the target position specified by the correct answer data,

a second error between the reference position specified by the output data and the reference position specified by the correct answer data, and

a third error between the positional relationship specified by the output data and the positional relationship specified by the correct answer data, or

wherein the loss function is based on the third error, and

wherein the model includes

an image encoding layer for encoding the input image, and

a text encoding layer for encoding the input text,

a part of features determined by the text encoding layer is input to the image encoding layer, and

a part of features determined by the image encoding layer is input to the text encoding layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2023
From: HOSOMI, NAOKI; MISU, TERUHISA; HATANAKA, SHUMPEI; YANG, WEI; SUGIURA, KOMEI
To: HONDA MOTOR CO., LTD; KEIO UNIVERSITY
Reel/Frame 065647/0534 →
Continuity (1)
Related Publication 20240428597A1 · Dec 26, 2024
References Cited (22)
US 10977501B2 · Mao · 2021 [cited by examiner]
US 11468688B2 · Cheng · 2022 [cited by examiner]
US 11574142B2 · Lin · 2023 [cited by examiner]
US 11663294B2 · Liu · 2023 [cited by examiner]
US 11978271B1 · Kharbanda · 2024 [cited by examiner]
US 12254707B2 · Xue · 2025 [cited by examiner]
US 12271792B2 · Li · 2025 [cited by examiner]
US 12394085B2 · Chen · 2025 [cited by examiner]
US 20200202145A1 · Mao et al. · 2020 [cited by applicant]
US 20210326609A1 · Mao et al. · 2021 [cited by applicant]
US 20240257536A1 · Ferroni · 2024 [cited by examiner]
US 20250095393A1 · Yang · 2025 [cited by examiner]
CN 116310920A · 2023 [cited by applicant]
JP 2009193097A · 2009 [cited by applicant]
JP 2015149013A · 2015 [cited by applicant]
JP 2022513866A · 2022 [cited by applicant]
WO 2020132082A1 · 2020 [cited by applicant]
WO 2023101679A1 · 2023 [cited by applicant]
Hosomi et al., Multimodal Target Localization With Landmark-Aware Positioning for Urban Mobility, 2024 IEEE 2377-3766, IEEE Robotics and Automation Letter, vol. 10, No. 1, Jan. 2025, pp. 716-723. (Year: 2025). [cited by examiner]
Dou et al., Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone, Nov. 18, 2022, https://arxiv.org/pdf/2206.07643.pdf. [cited by applicant]
Hatanaka, S. et al., Target Position Prediction Using UNITER Regressor for Understanding Navigation Instructions in Urban Areas, Keio University, Honda R&D Co., Ltd., Honda Research Institute USA, Aug. 19, 2022, pp. 34-… [cited by applicant]
International Search Report for PCT/JP2024/018221 dated Jul. 23, 2024. [cited by applicant]