IP Library › Granted Patent US 12,597,176
Granted Patent B2
US 12,597,176 · App. 18/029,093 · Granted Apr 7, 2026

Image generator and method of image generation

Inventors: Hiya Roy (Tokyo, JP); Mitsuru Nakazawa (Tokyo, JP); Bjorn Stenger (Tokyo, JP)
Assignee: RAKUTEN GROUP, INC.
G06T11/001G06T7/11G06V10/454G06V10/82G06T2207/20084G06T2207/20132G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,176
App. No.
18/029,093
Granted
Apr 7, 2026
Kind
B2
Abstract

Provided is an information-processing device including: a CPU; and a memory storing instructions for causing the information-processing device, when executed by the CPU, to: output an intermediate heatmap for input of an input image by using at least one of a plurality of machine learning models; and generate a heatmap based on an attribute of the input image, which is provided independently of the input image, and the intermediate heatmaps.

Claims (38)

1 . An information-processing device, comprising:

a CPU; and

a memory storing instructions for causing the information-processing device, when executed by the CPU, to:

output an intermediate heatmap for input of an input image by using at least one of a plurality of machine learning models;

generate a heatmap based on an attribute of the input image and the intermediate heatmaps;

wherein the attribute of the input image is provided independently of the input image; and

wherein each of the plurality of machine learning models is configured to output the intermediate heatmap irrespective of any other machine learning models among the plurality of machine learning models.

2 . The information-processing device according to claim 1 , wherein the instructions further cause the information-processing device to:

select, as a machine learning model to which the input image is to be input, at least one machine learning model from the plurality of machine learning models based on the attribute.

3 . The information-processing device according to claim 1 , wherein the instructions further cause the information-processing device to:

select at least one intermediate heatmap from a plurality of the intermediate heatmaps output from the plurality of machine learning models based on the attribute.

4 . The information-processing device according to claim 1 , wherein the instructions further cause the information-processing device to:

generate the heatmap by giving a weight to each of a plurality of the intermediate heatmaps and combining the plurality of the weighted intermediate heatmaps.

5 . The information-processing device according to claim 4 , wherein the instructions further cause the information-processing device to:

determine at least a part of the weights based on the attribute.

6 . The information-processing device according to claim 1 , wherein the instructions further cause the information-processing device to:

cut out a principal portion being a part of the input image based on the heatmap.

7 . The information-processing device according to claim 1 , wherein the plurality of machine learning models includes at least one of a first machine learning model that outputs a click through rate prediction or a second machine learning model that outputs aesthetic values.

8 . The information-processing device according to claim 1 , wherein each of the plurality of machine learning models is configured to output a heatmap independent of the attribute.

9 . The information-processing device according to claim 2 , wherein at least one of the unselected machine learning models from the plurality of machine learning models does not generate a heat map.

10 . The information-processing device according to claim 3 , wherein only a subset of the plurality of generated intermediate heat maps are selected and used to generate the heatmap.

11 . The information-processing device according to claim 6 , wherein the instructions cause the information-processing device to:

cut out the principal portion using a sliding window by setting a plurality of windows; extracting a plurality of candidate windows among the plurality of windows; and selecting a candidate window among the plurality of candidate windows as the principal portion.

12 . The information-processing device according to claim 6 , wherein the instructions cause the information-processing device to:

cut out the principal portion using a machine learning model to directly output a size, a shape, and a position of the principal portion.

13 . The information-processing device according to claim 1 , wherein the heatmap is an evaluation image.

14 . The information-processing device according to claim 1 , wherein a resolution of the heatmap does not match a resolution of the input image.

15 . The information-processing device according to claim 1 , wherein the heatmap is a combination of a plurality of the intermediate heatmaps.

16 . An information-processing method of causing a computer to execute:

outputting, through use of one or a plurality of machine learning models, one or a plurality of intermediate heatmaps for input of an input image;

generating a heatmap based on an attribute of the input image and the one or the plurality of intermediate heatmaps;

wherein the attribute of the input image is provided independently of the input image; and

wherein each of the plurality of machine learning models is configured to output the intermediate heatmap irrespective of any other machine learning models among the plurality of machine learning models.

17 . A non-transitory computer-readable information recording medium storing an information-processing program for causing a computer to:

output, through use of one or a plurality of machine learning models, one or a plurality of intermediate heatmaps for input of an input image;

generate a heatmap based on an attribute of the input image and the one or the plurality of intermediate heatmaps;

wherein the attribute of the input image is provided independently of the input image; and

wherein each of the plurality of machine learning models is configured to output the intermediate heatmap irrespective of any other machine learning models among the plurality of machine learning models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2023
From: ROY, HIYA; NAKAZAWA, MITSURU; STENGER, BJORN
To: RAKUTEN GROUP, INC.
Reel/Frame 063158/0125 →
Continuity (1)
Related Publication 20240362831A1 · Oct 31, 2024
References Cited (53)
US 12004871B1 · Fazeli · 2024 [cited by examiner]
US 12406023B1 · Alvarez Lopez et al. · 2025 [cited by applicant]
US 20090208118A1 · Csurka · 2009 [cited by applicant]
US 20150170053A1 · Miao · 2015 [cited by examiner]
US 20170345196A1 · Tanaka · 2017 [cited by examiner]
US 20190050681A1 · Tate · 2019 [cited by applicant]
US 20190057515A1 · Teixeira et al. · 2019 [cited by applicant]
US 20190370587A1 · Burachas et al. · 2019 [cited by applicant]
US 20200074634A1 · Kecskemethy et al. · 2020 [cited by applicant]
US 20200380302A1 · Nakazawa · 2020 [cited by examiner]
US 20210009080A1 · Hu et al. · 2021 [cited by applicant]
US 20210133861A1 · Kumar · 2021 [cited by examiner]
US 20210192772A1 · Tate · 2021 [cited by applicant]
US 20210249118A1 · Papagiannakis · 2021 [cited by examiner]
US 20210344936A1 · Liu · 2021 [cited by applicant]
US 20210374403A1 · Yumiba · 2021 [cited by examiner]
US 20210390700A1 · Lee · 2021 [cited by examiner]
US 20220138490A1 · Tate · 2022 [cited by applicant]
US 20220180528A1 · Dundar · 2022 [cited by examiner]
US 20220207875A1 · Kopparapu · 2022 [cited by examiner]
US 20220269895A1 · Barkan · 2022 [cited by examiner]
US 20220269996A1 · Nogami · 2022 [cited by applicant]
US 20220277472A1 · Birchfield · 2022 [cited by examiner]
US 20220327155A1 · Luo · 2022 [cited by examiner]
US 20220382802A1 · Sharifi · 2022 [cited by examiner]
US 20220391771A1 · Huang · 2022 [cited by examiner]
US 20230069310A1 · Myronenko · 2023 [cited by examiner]
US 20230153374A1 · Podlozhnyuk · 2023 [cited by examiner]
US 20230162051A1 · Lv · 2023 [cited by examiner]
US 20230410484A1 · Udayamurthy · 2023 [cited by examiner]
US 20240290054A1 · Yin · 2024 [cited by examiner]
CN 111629212A · 2020 [cited by applicant]
CN 112802034A · 2021 [cited by applicant]
JP 2000075889A · 2000 [cited by applicant]
JP 2019032773A · 2019 [cited by applicant]
JP 2020516427A · 2020 [cited by applicant]
JP 2020149641A · 2020 [cited by applicant]
JP 2021005301A · 2021 [cited by applicant]
JP 2021081793A · 2021 [cited by applicant]
JP 2021103347A · 2021 [cited by applicant]
JP 2021516646A · 2021 [cited by applicant]
WO 2021130856A1 · 2021 [cited by applicant]
Author = {Wu, Chenyun and Lin, Zhe and Cohen, Scott and Bui, Trung and Maji, Subhransu}, title = {PhraseCut: Language-Based Image Segmentation in the Wild}, month = {June}, year = {2020} (Year: 2020). [cited by examiner]
Yu Xiang et al. “Subcategory-aware Convolutional Neural Networks for Object Proposals and Detection”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 16, 2016, XP0806958… [cited by applicant]
Search Report of Dec. 15, 2023, for related EP Patent Application No. 21959402.5, pp. 1-11. [cited by applicant]
Wenguan Wang, Jianbing Shen, “Deep Cropping via Attention Box Prediction and Aesthetics Assessment”, [online], International Conference on Computer Vision (ICCV-2017); Published Oct. 2017; Internet <URL:https://openacce… [cited by applicant]
International Search Report for PCT/JP2021/036195 dated Nov. 22, 2021, p. 1-9. [cited by applicant]
Office Action of Jan. 10, 2023, for corresponding JP Application No. 2022-557886 with a partial translation of the Office Action, p. 1-5. [cited by applicant]
Chenyun Wu et.al. “PhraseCut: Language-based Image Segmentation in the Wild”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 3, 2020, XP081731210, pp. 1-17. [cited by applicant]
Search Report of Sep. 11, 2023, for corresponding EP Patent Application No. 21950363.8, pp. 1-9. [cited by applicant]
Office Action of Jun. 10, 2025, for related U.S. Appl. No. 18/029,668, pp. 1-35. [cited by applicant]
Xiang et al. “Subcategory-aware Convolutional Neural Networks for Object Proposals and Detection”, 2017 IEEE winter conference on applications of computer vision (WACV), IEEE, 2017, https://ieeexplore.ieee.org/abstract/… [cited by applicant]
Office Action of Dec. 3, 2025, for related U.S. Appl. No. 18/029,668, pp. 1-19. [cited by applicant]