IP Library › Granted Patent US 12,738,008
Granted Patent B2
US 12,738,008 · App. 18/603,752 · Granted Sep 15, 2026

Computer-readable recording medium storing region detection program, apparatus, and method

Inventors: Fan Yang (Edogawa, JP); Shigeyuki Odashima (Tama, JP)
Assignee: Fujitsu Limited
G06V10/25G06T7/60G06T7/70G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,008
App. No.
18/603,752
Granted
Sep 15, 2026
Kind
B2
Abstract

Avoiding False Negatives in vision-based multi-view gymnast detection presents a significant challenge. Detection results with False Negatives can further impact subsequent processing, such as multi-view 3D pose estimation. Therefore, interpolating missing bounding boxes is a desirable solution. Assuming that calibrated camera parameters are known, our method interpolates missing bounding boxes when a gymnast's 2D bounding boxes are detected in more than two views but are absent in others. This method primarily involves three steps: 1) Inferring the vertical 3D body center line using detected cross-view 2D bounding boxes and camera parameters through 2D-to-3D projection; 2) Obtaining the average 3D gymnast scale from pre-acquired data, and then calculating the 3D horizontal scale based on the 3D vertical scale obtained in step 1; 3) Interpolating the missing 2D bounding boxes using the inferred 3D vertical line and horizontal scale through 3D-to-2D projection.

Claims (24)

1 . A non-transitory computer-readable recording medium storing a region detection program for causing a computer to execute a process comprising:

acquiring images each which is captured by each of a plurality of imaging apparatuses that are arranged in the same horizontal plane and capture the respective images of a person from respective different directions;

detecting a region indicating the person from each of the images by inputting the images to a machine learning model which is generated in advance by a machine learning so as to detect the region indicating the person; and

interpolating, based on a first region of the person which is detected from a first image in which the region indicating the person is detected by the machine learning model in the images and a parameter of each of the plurality of imaging apparatuses, a second region indicating the person in a second image in which the region indicating the person is not detected by the machine learning model in the images, a width of the second region being estimated based on a height of the first region, and a height of the second region estimated by converting an end point of a vertical center line of the first region into a coordinate of an end point of a vertical center line of the person in a three dimensional space based on the parameter of an imaging apparatus which captures the first image and converting the coordinate into a coordinate in the second image based on the parameter of an imaging apparatus which captures the second image, and statistical information regarding a posture of the person, and the statistical information regarding a posture of the person being a mean of a sum of a height and a width of a rectangular parallelepiped surrounding the person in each of a plurality of different postures of the person.

2 . The non-transitory computer-readable recording medium according to claim 1 , wherein

a length of the vertical center line of the person in the three dimensional space is set as a height of the person in the three dimensional space, a difference between the mean indicated by the statistical information and the height of the person in the three dimensional space is estimated as a width of the person in the three dimensional space and the width of the second region is estimated based on a ratio of the height and the width of the person in the three dimensional space and the height of the second region.

3 . The non-transitory computer-readable recording medium according to claim 1 , wherein

the plurality of imaging apparatuses are arranged in a same vertical plane, and

a height of the second region is estimated based on a width of the first region, a width of the second region, and statistical information regarding a posture of the person.

4 . A region detection apparatus comprising:

a memory; and

a processor coupled to the memory and configured to:

acquire images each which is captured by each of a plurality of imaging apparatuses that are arranged in the same horizontal plane and capture the respective images of a person from respective different directions;

detected a region indicating the person from each of the images by inputting the images to a machine learning model which is generated in advance by a machine learning so as to detect the region indicating the person; and

interpolate, based on a first region of the person which is detected from a first image in which the region indicating the person is detected by the machine learning model in the images and a parameter of each of the plurality of imaging apparatuses, a second region indicating the person in a second image in which the region indicating the person is not detected by the machine learning model in the images, a width of the second region being estimated based on a height of the first region, and a height of the second region estimated by converting an end point of a vertical center line of the first region into a coordinate of an end point of a vertical center line of the person in a three dimensional space based on the parameter of an imaging apparatus which captures the first image and converting the coordinate into a coordinate in the second image based on the parameter of an imaging apparatus which captures the second image, and statistical information regarding a posture of the person, and the statistical information regarding a posture of the person being a mean of a sum of a height and a width of a rectangular parallelepiped surrounding the person in each of a plurality of different postures of the person.

5 . The region detection apparatus according to claim 4 , wherein

a length of the vertical center line of the person in the three dimensional space is set as a height of the person in the three dimensional space, a difference between the mean indicated by the statistical information and the height of the person in the three dimensional space is estimated as a width of the person in the three dimensional space and the width of the second region is estimated based on a ratio of the height and the width of the person in the three dimensional space and the height of the second region.

6 . The region detection apparatus according to claim 4 , wherein

the plurality of imaging apparatuses are arranged in a same vertical plane, and

a height of the second region is estimated based on a width of the first region, a width of the second region, and statistical information regarding a posture of the person.

7 . A region detection method for executing a process comprising:

acquiring images each which is captured by each of a plurality of imaging apparatuses that are arranged in the same horizontal plane and capture the respective images of a person from respective different directions;

detecting a region indicating the person from each of the images by inputting the images to a machine learning model which is generated in advance by a machine learning so as to detect the region indicating the person; and

interpolating, based on a first region of the person which is detected from a first image in which the region indicating the person is detected by the machine learning model in the images and a parameter of each of the plurality of imaging apparatuses, a second region indicating the person in a second image in which the region indicating the person is not detected by the machine learning model in the images, a width of the second region being estimated based on a height of the first region, and a height of the second region estimated by converting an end point of a vertical center line of the first region into a coordinate of an end point of a vertical center line of the person in a three dimensional space based on the parameter of an imaging apparatus which captures the first image and converting the coordinate into a coordinate in the second image based on the parameter of an imaging apparatus which captures the second image, and statistical information regarding a posture of the person, and the statistical information regarding a posture of the person being a mean of a sum of a height and a width of a rectangular parallelepiped surrounding the person in each of a plurality of different postures of the person.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2024
From: YANG, FAN; ODASHIMA, SHIGEYUKI
To: FUJITSU LIMITED
Reel/Frame 066753/0136 →
Continuity (2)
Continuation PCTJP2021037958 · Oct 13, 2021
Related Publication 20240242464A1 · Jul 18, 2024
References Cited (25)
US 9692964B2 · Steinberg · 2017 [cited by examiner]
US 11113887B2 · Kopeinigg · 2021 [cited by examiner]
US 11227194B2 · Zhang · 2022 [cited by examiner]
US 11270347B2 · Lehman · 2022 [cited by examiner]
US 11282282B2 · Simpkinson · 2022 [cited by examiner]
US 11373318B1 · McKennoch · 2022 [cited by examiner]
US 11804045B2 · Kudo · 2023 [cited by examiner]
US 20120327197A1 · Yamashita et al. · 2012 [cited by applicant]
US 20160292533A1 · Uchiyama et al. · 2016 [cited by applicant]
US 20180341835A1 · Siminoff · 2018 [cited by examiner]
US 20190114824A1 · Martinez · 2019 [cited by applicant]
US 20190122378A1 · Aswin · 2019 [cited by applicant]
US 20190287310A1 · Kopeinigg · 2019 [cited by examiner]
US 20190356885A1 · Ribeiro · 2019 [cited by examiner]
US 20220284672A1 · Teranishi et al. · 2022 [cited by applicant]
JP 2002290962A · 2002 [cited by applicant]
JP 2009143722A · 2009 [cited by applicant]
JP 2018132847A · 2018 [cited by applicant]
JP 2021071749A · 2021 [cited by applicant]
WO 2021100681A1 · 2021 [cited by applicant]
EESR—Extended European Search Report mailed on Oct. 28, 2024 for corresponding European Patent Application No. 21960618.3 [10 pages]. **JP2021-071749. [cited by applicant]
WIPO, International Search Report with English-language translation and Written Opinion mailed Dec. 14, 2021, in connection with PCT/JP2021/037958. [cited by applicant]
Hideo Saito, et al., “View Interpolation of Multiple Cameras Based on Projective Geometry”, 7 pages, 2002 Internet URL: https://www.researchgate.net/profile/Makoto-Kimura-7/publication/2840433_View_Interpolation_of_Mult… [cited by applicant]
OpenCV, Intel Corporation, Ver.3.4, 1 page, 2017 “Triangulation—Structure from Motion” Internet URL: https://docs.opencv.org/3.4/d0/dbd/group_triangulation.html. [cited by applicant]
EPOA—Office Action mailed on Dec. 3, 2025 for European Patent Application No. 21960618.3 (6 pages). *WO2021/100681A1 EPOA was previously submitted in the IDS filed on Nov. 12, 2024. [cited by applicant]