IP Library Granted Patent US 12,711,805
Granted Patent B2
US 12,711,805 · App. 18/449,192 · Granted Aug 18, 2026

Ethical human-centric image dataset

Inventors: Jerone Andrews (Tokyo, JP); Alice Xiang (Seattle, WA)
Assignee: SONY GROUP CORPORATION
G06V40/171G06F3/04845G06T7/70G06V10/25G06V20/70G06V40/172G06T2207/20081G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,805
App. No.
18/449,192
Granted
Aug 18, 2026
Kind
B2
Abstract

A diverse dataset of human images can be created by collecting a plurality of images from a plurality of diverse people. A first graphical user interface requires a user to provide subject data, instrument data and environment data as metadata for each of the plurality of images. A second graphical user interface requires a user to form a bounding box about a face of a subject in each of the plurality of images. A third graphical user interface requires annotators to provide annotations for each of the plurality of images. The dataset may be used for training or evaluating machine learning or artificial intelligence systems, such as systems for body and face detection, body and face landmark detection, body and face parsing, face alignment, face recognition, face verification, image editing and image synthesis.

Claims (56)

1 . A computer-implemented method for constructing a dataset of human images, comprising:

collecting a plurality of images from a plurality of diverse people, each of the plurality of diverse people being a primary subject;

providing a first graphical user interface requiring the primary subject to provide subject data, instrument data and environment data as metadata for each of the plurality of images; and

storing the plurality of images as the dataset, wherein

the subject data includes demographic information, physical characteristics, actions and head pose;

the plurality of images includes from 4 to 10 images of each primary subject;

each of the plurality of images per each primary subject are captured at least one day apart; and

the system adjusts the storing of the plurality of images into the dataset to meet selected specifications related to diversity of the subject data, the instrument data and the environment data.

2 . The computer-implemented method of claim 1 , further comprising providing a second graphical user interface permitting a user to form a bounding box about a face of a subject in each of the plurality of images.

3 . The computer-implemented method of claim 1 , further comprising providing a third graphical user interface permitting annotators to provide annotations for each of the plurality of images.

4 . The computer-implemented method of claim 1 , wherein no more than 20% of the plurality of images contain subject-subject interaction annotations where the primary subject is annotated as interacting with a secondary subject.

5 . The computer-implemented method of claim 1 , wherein each of the plurality of images per each primary subject are taken in a different location and at a different time of day.

6 . The computer-implemented method of claim 1 , further comprising obtaining explicit informed consent from each of the plurality of diverse people, wherein the explicit informed consent is provided as metadata for each of the plurality of images.

7 . The computer-implemented method of claim 3 , wherein the annotators are demographically diverse with respect to age, pronouns and ancestry.

8 . The computer-implemented method of claim 3 , wherein the annotations include segmentation labels for each part of a subject's body in each of the plurality of images.

9 . The computer-implemented method of claim 1 , wherein the physical characteristics include age, pronouns, nationality, residence, ancestry and disability; physical characteristics including skin tone, eye color, head hair type, head hair style, head hair color, facial hair style, facial hair color, height, weight, and facial marks.

10 . The computer-implemented method of claim 1 , wherein the actions include body pose, subject-object interaction and subject-subject interaction.

11 . The computer-implemented method of claim 1 , wherein the environment data includes illumination, scene, camera position and camera distance.

12 . The computer-implemented method of claim 1 , further comprising providing an output illustrating the diversity of the dataset with respect to each of the subject data, the instrument data and the environment data.

13 . A computer-implemented method for training or evaluating commercial machine learning or artificial intelligence systems in an unconstrained setting, the method comprising:

creating a diverse dataset of human images by:

collecting a plurality of images from a plurality of diverse people each of the plurality of diverse people being a primary subject;

providing a first graphical user interface requiring the primary subject to provide subject data, instrument data and environment data as metadata for each of the plurality of images;

providing a second graphical user interface requiring a user to form a bounding box about a face of a subject in each of the plurality of images;

providing a third graphical user interface requiring annotators to provide annotations for each of the plurality of images; and

storing the plurality of images as the dataset; and

training or evaluating the machine learning or artificial intelligence system by using the diverse dataset in the machine learning or artificial intelligence system, wherein:

the plurality of images includes from 4 to 10 images of each primary subject;

each of the plurality of images per each primary subject are captured at least one day apart;

at least 50% of the plurality of images per each primary subject are captured at least seven days apart; and

the system adjusts the storing of the plurality of images into the dataset to meet selected specifications related to diversity of the subject data, the instrument data and the environment data.

14 . The computer-implemented method of claim 13 , wherein the machine learning or artificial intelligence system is operable for one or more of body and face detection, body and face landmark detection, body and face parsing, face alignment, face recognition, face verification, image editing and image synthesis.

15 . The computer-implemented method of claim 13 , wherein:

no more than 20% of the plurality of images contain subject-subject interaction annotations where the primary subject is annotated as interacting with a secondary subject; and

each of the plurality of images per each primary subject are taken in a different location and at a different time of day.

16 . The computer-implemented method of claim 13 , further comprising obtaining explicit informed consent from each of the plurality of diverse people, wherein the explicit informed consent is provided as metadata for each of the plurality of images.

17 . The computer-implemented method of claim 13 , wherein the annotators are demographically diverse with respect to age, pronouns and ancestry.

18 . The computer-implemented method of claim 13 , wherein:

the subject data includes demographic information, physical characteristics, actions and head pose;

the physical characteristics include age, pronouns, nationality, residence, ancestry and disability; physical characteristics including skin tone, eye color, head hair type, head hair style, head hair color, facial hair style, facial hair color, height, weight, and facial marks;

the actions include body pose, subject-object interaction and subject-subject interaction; and

the environment data includes illumination, scene, camera position and camera distance.

19 . A computer-implemented method for constructing a dataset of human images, comprising:

collecting a plurality of images from a plurality of diverse people, each of the plurality of diverse people being a primary subject;

providing a first graphical user interface requiring the primary subject to provide subject data, instrument data and environment data as metadata for each of the plurality of images;

providing a second graphical user interface requiring a user to form a bounding box about a face of a subject in each of the plurality of images;

providing a third graphical user interface requiring annotators to provide annotations for each of the plurality of images; and

storing the plurality of images as the dataset, wherein:

the subject data includes demographic information, physical characteristics, actions and head pose;

the plurality of images includes from 4 to 10 images of each primary subject;

each of the plurality of images per each primary subject are captured at least one day apart; and

the system adjusts the storing of the plurality of images into the dataset to meet selected specifications related to diversity of the subject data, the instrument data and the environment data.

20 . The computer-implemented method of claim 19 , wherein:

the physical characteristics include age, pronouns, nationality, residence, ancestry and disability; physical characteristics including skin tone, eye color, head hair type, head hair style, head hair color, facial hair style, facial hair color, height, weight, and facial marks;

the actions include body pose, subject-object interaction and subject-subject interaction; and

the environment data includes illumination, scene, camera position and camera distance.