IP Library › Granted Patent US 12,361,682
Granted Patent B2
US 12,361,682 · App. 18/008,730 · Granted Jul 15, 2025

Information processing apparatus, information processing method, and non-transitory computer readable medium

Inventors: Josue Cuevas Juarez (Tokyo, JP); Hasan Arslan (Tokyo, JP); Rajasekhar Sanagavarapu (Tokyo, JP)
Assignee: Rakuten Group, Inc.
G06V10/7715G06V10/44G06V20/60G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,682
App. No.
18/008,730
Granted
Jul 15, 2025
Kind
B2
Abstract

An information processing apparatus includes a processor configured to read program code stored in memory and operate as instructed by the program code. The program code includes acquisition code configured to cause the at least one processor to acquire a red, blue, green (RGB) image including an object. The program code includes converting code configured to cause the at least one processor to apply a discrete cosine transform (DCT) to the RGB image to generate image coefficients corresponding to an YCbCr image comprising Luma (Y) elements and Chroma (Cb, Cr) elements. The program code includes prediction code configured to cause the at least one processor to predict various attributes relating to the object by inputting the image coefficient into a learning model. The learning model is a learning model that is stored in the at least one memory and shared between a plurality of different objects including the object.

Claims (40)

1. An information processing apparatus comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code including:

acquisition code configured to cause the at least one processor to acquire a red, blue, green (RGB) image including an object;

converting code configured to cause the at least one processor to apply a discrete cosine transform (DCT) to the RGB image to generate image coefficients corresponding to an YCbCr image comprising Luma (Y) elements and Chroma (Cb, Cr) elements; and

prediction code configured to cause the at least one processor to predict various attributes relating to the object by inputting the image coefficient into a learning model,

wherein the learning model is a learning model that is stored in the at least one memory and shared between a plurality of different objects including the object, and wherein the learning model includes:

a plurality of estimation layers that estimate a plurality of attribute values for a plurality of attributes relating to the plurality of different objects, and

an output layer that concatenates and outputs the plurality of attribute values outputted from the plurality of estimation layers.

2. The information processing apparatus according to claim 1 ,

wherein the learning model is composed of a first part and a second part,

the first part receives the image coefficients as an input and outputs a feature vector expressing features of the object,

the second part includes the plurality of estimation layers and the output layer,

the plurality of estimation layers receive the feature vector as an input and output a value indicating an object type of the object and the plurality of attribute values, and

the output layer concatenates and outputs the value indicating the object type of the object and the plurality of attribute values outputted from the plurality of estimation layers.

3. The information processing apparatus according to claim 2 ,

wherein the prediction code is further configured to cause the at least one processor to predict the various attributes from the plurality of attribute values outputted from the second part of the learning model.

4. The information processing apparatus according to claim 2 ,

wherein at least one valid attribute value out of the plurality of attribute values is set in advance in keeping with the value indicated by the object type, and

wherein the prediction code is further configured to cause the at least one processor to acquire, from the plurality of attribute values, the at least one valid attribute value in keeping with a value indicated by the object type, and predict attributes corresponding to the at least one valid attribute value as the various attributes.

5. The information processing apparatus according to claim 1 ,

wherein in a case where the RGB object image includes a plurality of objects, the prediction code is further configured to cause the at least one processor to predict various attributes that relate to each of the plurality of objects.

6. The information processing apparatus according to claim 1 ,

wherein the image coefficients are concatenated data produced by size matching of the Y elements, the Cb elements, and the Cr elements out of the data produced by the discrete cosine transform.

7. The information processing apparatus according to claim 1 ,

wherein the program code further comprises output configured to cause the at least one processor to output the various attributes.

8. An information processing method performed by at least one processor, the method comprising:

acquiring a red, blue, green (RGB) object image including an object;

applying a discrete cosine transform (DCT) to the RGB image to generate image coefficients corresponding to an YCbCr image comprising Luma (Y) elements and Chroma (Cb, Cr) elements; and

predicting various attributes relating to the object by inputting the image coefficients into a learning model,

wherein the learning model is a learning model that is stored in a memory and shared between a plurality of different objects including the object, and wherein the learning model includes:

a plurality of estimation layers that estimate a plurality of attribute values for a plurality of attributes relating to the plurality of different objects; and

an output layer that concatenates and outputs the plurality of attribute values outputted from the plurality of estimation layers.

9. A non-transitory computer readable medium having instructions stored therein, which when executed by a processor, cause the processor to execute a method comprising:

acquiring a red, blue, green (RGB) an object image including an object;

applying a discrete cosine transform (DCT) to the RGB image to generate image coefficients corresponding to an YCbCr image comprising Luma (Y) elements and Chroma (Cb, Cr) elements; and

predicting various attributes relating to the object by inputting a learning model to the object image,

wherein the learning model is a learning model that is stored in memory and shared between a plurality of different objects including the object and includes:

a plurality of estimation layers that estimate a plurality of attribute values for a plurality of attributes relating to the plurality of different objects; and

an output layer that concatenates and outputs the plurality of attribute values outputted from the plurality of estimation layers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2022
From: CUEVAS JUAREZ, JOSUE; ARSLAN, HASAN; SANAGAVARAPU, RAJASEKHAR
To: RAKUTEN GROUP, INC.
Reel/Frame 062080/0298 →
Continuity (1)
Related Publication 20240303969A1 · Sep 12, 2024
References Cited (10)
US 20170372515A1 · Hauswiesner · 2017 [cited by examiner]
US 20200320769A1 · Chen · 2020 [cited by examiner]
JP 2005101712A · 2005 [cited by applicant]
JP 2016139189A · 2016 [cited by applicant]
“Shanxin Yuan et al., NTIRE 2020 Challenge on Image Demoireing: Methods and Results, 2020, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 460-461” (Year: 2020). [cited by examiner]
“Mohammad Moosazadeh et. al., Robust Image Watermarking Algorithm using DCT Coefficients Relation in YCoCg-R Color Space, Sep. 2016, 2016 Eighth International Conference on Information and Knowledge Technology, Iran” (Y… [cited by examiner]
Communication, issued in European Application No. 21943330.7. [cited by applicant]
Tsapatsoulis N., et al. “Object Classification Using the MPEG-7 Visual Descriptors: An Experimental Evaluation Using State of the Art Data Classifiers” Sep. 14, 2009 (Sep. 14, 2009), SAT 2015 18th International Conferen… [cited by applicant]
Tereshchenko S.N. et al. “Features of Applying Pretrained Convolutional Neural Networks to Graphic Image Steganalysis”, Optoelectronics, Instrumentation and Data Processing, Pleiades Publishing, Moscow, vol. 57, No. 4, … [cited by applicant]
Temburwar S. et al. “Deep Learning Based Image Retrieval in the JPEG Compressed Domain”, Shrikant Temburwar et al: “Deep Learning Based Image Retrieval in the JPEG Compressed Domain”, arxiv.org, Cornell University Libra… [cited by applicant]