IP Library › Granted Patent US 12,633,101
Granted Patent B2
US 12,633,101 · App. 18/198,680 · Granted May 19, 2026

Information processing apparatus, information processing method, learning method and moving object for predicting a region in an image corresponding to utterance

Inventors: Naoki Hosomi (Wako, JP); Teruhisa Misu (San Jose, CA); Shumpei Hatanaka (Yokohama, JP); Wei Yang (Yokohama, JP); Komei Sugiura (Yokohama, JP)
Assignees: HONDA MOTOR CO., LTD.; KEIO UNIVERSITY
G06V10/806G06T7/11G06V10/764G06V10/7715G06V10/774G06V10/82G06V10/945G06V20/70G10L15/02G10L15/16G10L15/22G06T2207/20021G06T2207/20081G06T2207/20084G06T2207/20092G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,101
App. No.
18/198,680
Granted
May 19, 2026
Kind
B2
Abstract

An information processing apparatus in embodiments performs at least one trained machine learning model that includes an encoder and a decoder. The encoder receives inputs of text information including designation of a place, a first image that is an image captured by an image capturing apparatus and that includes the place, and a second image obtained by dividing a region for every identical object in the first image, and outputs tri-modal features that have been generated to include visual features of the first image that has been captured, visual features of the second image obtained by dividing the region, and language features of the text information. The decoder outputs a region on the first image corresponding to the designation of the place in the text information, by using the tri-modal features.

Claims (44)

1 . An information processing apparatus comprising

at least one processor configured to perform at least one trained machine learning model, wherein

the at least one trained machine learning model includes:

an encoder configured to receive inputs of text information including designation of a place, a first image that is an image captured by an image capturing apparatus and that includes the place, and a second image obtained by dividing a region for every identical object in the first image, and configured to output tri-modal features that have been generated to include visual features of the first image that has been captured, visual features of the second image obtained by dividing the region, and language features of the text information; and

a decoder configured to output a region on the first image corresponding to the designation of the place in the text information, by using the tri-modal features.

2 . The information processing apparatus according to claim 1 , wherein

the encoder includes a first encoder block and a second encoder block, the first encoder block being configured to output bi-modal features corresponding to the first image, the bi-modal features being obtained by fusing the language features of the text information into the visual features of the first image by use of an attention mechanism, the second encoder block being configured to output bi-modal features corresponding to the second image, the bi-modal features being obtained by fusing the language features of the text information into the visual features of the second image by use of an attention mechanism; and

the encoder either concatenates or adds the bi-modal features corresponding to the first image and the bi-modal features corresponding to the second image, and outputs the tri-modal features.

3 . The information processing apparatus according to claim 2 , wherein the encoder further outputs a visual feature map in which the bi-modal features corresponding to the first image are merged with the visual features of the first image, and a visual feature map in which the bi-modal features corresponding to the second image are merged with the visual features of the second image.

4 . The information processing apparatus according to claim 3 , wherein the encoder includes:

a first layer encoder configured to output the tri-modal features, the visual feature map related to the first image, and the visual feature map related to the second image; and

a second layer encoder configured to generate bi-modal features in which the visual feature maps that have been input from the first layer encoder and the language features of the text information are respectively fused, and configured to output tri-modal features in which the bi-modal features that have been generated are combined together.

5 . The information processing apparatus according to claim 4 , wherein a spatial size of the tri-modal features output from the second layer encoder is smaller than a spatial size of the tri-modal features output from the first layer encoder.

6 . The information processing apparatus according to claim 5 , wherein the decoder includes:

a second layer decoder configured to decode by using the tri-modal features output from the second layer encoder; and

a first layer decoder configured to decode features in which the tri-modal features output from the first layer encoder are incorporated into the features decoded by the second layer decoder.

7 . The information processing apparatus according to claim 1 , further comprising a classification model configured to classify a state of a subject displayed in the first image as a subtask, by using the tri-modal features.

8 . The information processing apparatus according to claim 2 , further comprising a classification model configured to classify a state of a subject displayed in the first image as a subtask, by using the bi-modal features corresponding to the second image.

9 . The information processing apparatus according to claim 4 , further comprising a classification model configured to classify a state of a subject displayed in the first image as a subtask, by using the tri-modal features generated by the second layer encoder.

10 . The information processing apparatus according to claim 4 , further comprising a classification model configured to classify a state of a subject displayed in the first image as a subtask, by using the bi-modal features generated by the second layer encoder.

11 . The information processing apparatus according to claim 7 , wherein the classification model configured to classify the state of the subject classifies at least any of whether the state of the subject is day or night, which time zone, what weather, and whether there is sunshine.

12 . The information processing apparatus according to claim 3 , wherein the encoder performs processing by using a language gate on the bi-modal features corresponding to the second image, adds the visual features of the second image, and outputs the visual feature map related to the second image.

13 . The information processing apparatus according to claim 1 , wherein the visual features of the first image are generated by inputting the first image into a transformer corresponding to the first image, and the visual features of the second image are generated by inputting the second image into a transformer corresponding to the second image.

14 . The information processing apparatus according to claim 1 , wherein the text information includes either utterance by voice or text that has been input.

15 . A learning method performed by an information processing apparatus for training at least one machine learning model to generate at least one trained machine learning model, wherein

the at least one machine learning model each includes a neural network, and the at least one machine learning model includes:

an encoder configured to receive inputs of text information including designation of a place, a first image that is an image captured by an image capturing apparatus and that includes the place, and a second image obtained by dividing a region for every identical object in the first image, configured to generate bi-modal features corresponding to each image in which language features of the text information are fused with visual features of each image by use of an attention mechanism, and configured to output tri-modal features in which the bi-modal features that have been generated are combined together; and

a decoder configured to output a region on the first image corresponding to the designation of the place in the text information, by using either the tri-modal features or the bi-modal features corresponding to each image,

the learning method comprising

changing a weighting parameter of the neural network to reduce a value of a loss function using a loss with use of an output from the decoder and a correct answer indicating the region on the first image.

16 . The learning method according to claim 15 , wherein the at least one machine learning model further includes a classification model configured to classify a state of a subject displayed in the first image as a subtask, by using either features including the visual features of the second image or the tri-modal features,

the learning method further comprising

changing the weighting parameter of the neural network to reduce the value of the loss function using the loss with use of the output from the decoder and the correct answer indicating the region on the first image, and a loss with use of an output from the classification model and a correct answer indicating the state of the subject.

17 . A moving object comprising:

an image capturing apparatus;

an acquisition unit configured to acquire text information; and

at least one processor configured to perform processing of at least one trained machine learning model, wherein

the at least one trained machine learning model includes:

an encoder configured to receive inputs of text information including designation of a place, a first image that is an image captured by the image capturing apparatus and that includes the place, and a second image obtained by dividing a region for every identical object in the first image, and configured to output tri-modal features that have been generated to include visual features of the first image that has been captured, visual features of the second image obtained by dividing the region, and language features of the text information; and

a decoder configured to output a region on the first image corresponding to the designation of the place in the text information, by using the tri-modal features.

18 . An information processing method performed by an information processing apparatus for performing at least one trained machine learning model,

the information processing method comprising:

receiving inputs of text information including designation of a place, a first image that is an image captured by an image capturing apparatus and that includes the place, and a second image obtained by dividing a region for every identical object in the first image, and encoding tri-modal features that have been generated to include visual features of the first image that has been captured, visual features of the second image obtained by dividing the region, and language features of the text information; and

decoding a region on the first image corresponding to the designation of the place in the text information, by using the tri-modal features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2024
From: HOSOMI, NAOKI; MISU, TERUHISA; HATANAKA, SHUMPEI; YANG, WEI; SUGIURA, KOMEI
To: HONDA MOTOR CO., LTD.; KEIO UNIVERSITY
Reel/Frame 066666/0214 →
Continuity (1)
Related Publication 20240386711A1 · Nov 21, 2024
References Cited (13)
US 11151406B2 · Huang et al. · 2021 [cited by applicant]
US 20240142995A1 · Toyoshi et al. · 2024 [cited by applicant]
JP 2019216442A · 2019 [cited by examiner]
JP 2020032844A · 2020 [cited by applicant]
JP 2020054651A · 2020 [cited by examiner]
JP 2020135852A · 2020 [cited by applicant]
JP 2020190930A · 2020 [cited by applicant]
WO 2022181130A1 · 2022 [cited by applicant]
Liu et al., RoBERTa: A Robustly Optimized BERT Pretraining Approach, Jul. 26, 2019, Paul G. Allen School of Computer Science & Engineering, University of Washington, Seattle, WA, https://arxiv.org/abs/1907.11692v1. [cited by applicant]
Cheng et al., Masked-attention Mask Transformer for Universal Image Segmentation, CVPR 2022, https://bowenc0221.github.io/mask2former/. [cited by applicant]
Yang et al., LAVT: Language-Aware Vision Transformer for Referring Image Segmentation, pp. 18155-18165, https://ppenaccess.thecvf.com/content/CVPR2022/papers/Yang_LAVT_Language-Aware_Vision_Transformer_for_Referring_Ima… [cited by applicant]
Liu et al., Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, pp. 10012-10022, https://openaccess.thecvf.com/content/ICCV2021/papers/Liu_Swin_Transformer_Hierarchical_Vision_Transformer_Using_Shif… [cited by applicant]
International Search Report and Written Opinion for PCT/JP2024/018220 mailed Jul. 30, 2024. [cited by applicant]