IP Library › Granted Patent US 12,464,237
Granted Patent B2
US 12,464,237 · App. 18/442,261 · Granted Nov 4, 2025

Inference apparatus, image capturing apparatus, training apparatus, inference method, training method, and storage medium

Inventors: Hideyuki Hamano (Kanagawa, JP); Akihiko Kanda (Kanagawa, JP); Kuniaki Sugitani (Kanagawa, JP); Yohei Matsui (Kanagawa, JP)
Assignee: CANON KABUSHIKI KAISHA
H04N23/675G06T7/50H04N23/61H04N23/635G06T2207/20081G06T2207/20084G06T2207/30196H04N23/672
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,464,237
App. No.
18/442,261
Granted
Nov 4, 2025
Kind
B2
Abstract

There is provided an inference apparatus. An inference unit performs inference with use of a machine learning model based on a subject region including a subject within an image obtained through shooting, and on a plurality of distance information pieces detected from a plurality of focus detection regions inside the subject region, thereby generating an inference result indicating a distance information piece corresponding to the subject or a distance information range corresponding to the subject. The machine learning model is a model that has been trained to suppress a contribution made to the inference result by one or more distance information pieces that are not based on the subject among the plurality of distance information pieces.

Claims (79)

1 . An image capturing apparatus comprising at least one processor and/or at least one circuit which functions as:

an image capturing unit configured to generate an image through shooting;

a first detection unit configured to detect a subject region from the image;

a second detection unit configured to detect a plurality of distance information pieces from a plurality of focus detection regions inside the subject region; and

an inference unit configured to perform inference with use of a machine learning model based on the subject region including a subject within the image obtained through shooting, and on the plurality of distance information pieces detected from the plurality of focus detection regions inside the subject region, thereby generating an inference result indicating a distance information piece corresponding to the subject or a distance information range corresponding to the subject,

wherein the machine learning model is a model that has been trained to suppress a contribution made to the inference result by one or more distance information pieces that are not based on the subject among the plurality of distance information pieces.

2 . The image capturing apparatus according to claim 1 , wherein

the subject region includes a first region where the subject exists, and a second region where the subject does not exist, and

the one or more distance information pieces that are not based on the subject among the plurality of distance information pieces include one or more distance information pieces corresponding to one or more focus detection regions corresponding to the second region among the plurality of focus detection regions.

3 . The image capturing apparatus according to claim 1 , wherein

the one or more distance information pieces that are not based on the subject among the plurality of distance information pieces include one or more distance information pieces with a detection error that exceeds a predetermined extent among the plurality of distance information pieces.

4 . The image capturing apparatus according to claim 1 , wherein the at least one processor and/or the at least one circuit further functions as

a first adjustment unit configured to adjust a focus of an optical system used in the shooting based on the inference result.

5 . The image capturing apparatus according to claim 4 , wherein

the inference result indicates the distance information range corresponding to the subject, and

the first adjustment unit adjusts the focus of the optical system based on a distance information piece that is included in the distance information range among the plurality of distance information pieces.

6 . The image capturing apparatus according to claim 1 , wherein

the inference result indicates the distance information range corresponding to the subject, and

the at least one processor and/or the at least one circuit further functions as

a second adjustment unit configured to adjust a diaphragm of an optical system used in the shooting based on the distance information range.

7 . The image capturing apparatus according to claim 6 , wherein

the second adjustment unit adjusts the diaphragm of the optical system based on the distance information range so that the subject is included in a depth of field.

8 . The image capturing apparatus according to claim 1 , wherein

the inference result indicates the distance information range corresponding to the subject, and

the at least one processor and/or the at least one circuit further functions as

a display unit configured to display information for giving notice of a focus detection region corresponding to a distance information piece that is included in the distance information range among the plurality of distance information pieces.

9 . The image capturing apparatus according to claim 1 , wherein

the plurality of distance information pieces detected from the plurality of focus detection regions are a plurality of defocus amounts detected from the plurality of focus detection regions.

10 . The image capturing apparatus according to claim 9 , wherein

the inference result indicates a defocus amount corresponding to the subject or a defocus amount range corresponding to the subject.

11 . The image capturing apparatus according to claim 1 , wherein

the first detection unit detects a plurality of subject regions each of which partially includes each of a plurality of subjects,

the second detection unit detects a plurality of distance information pieces for each of the plurality of subject regions, and

the inference unit generates an inference result for each of the plurality of subjects.

12 . The image capturing apparatus according to claim 11 , wherein

the plurality of subjects correspond to a plurality of parts of a subject of a predetermined type.

13 . An inference apparatus comprising at least one processor and/or at least one circuit which functions as:

an image capturing unit configured to generate an image through shooting;

an obtainment unit configured to obtain the image obtained through shooting, information of a subject region including a subject within the image, and a plurality of distance information pieces that respectively correspond to a plurality of regions inside the subject region; and

an inference unit configured to perform inference with use of a machine learning model using, as inputs, the image, the information of the subject region, and the plurality of distance information pieces, thereby generating an inference result indicating a distance information piece corresponding to the subject or a distance information range corresponding to the subject.

14 . The inference apparatus according to claim 13 , wherein

the obtainment unit obtains a plurality of distance information pieces corresponding to a plurality of parts of the subject, and

the inference unit generates an inference result indicating a plurality of distance information ranges corresponding to the plurality of parts.

15 . An image capturing apparatus, comprising:

the inference apparatus according to claim 14 ,

wherein the at least one processor and/or the at least one circuit further functions as:

a determination unit configured to, based on priority degrees of the plurality of parts, determine a part to be focused on from the plurality of distance information ranges corresponding to the plurality of parts output from the inference apparatus.

16 . The image capturing apparatus according to claim 15 , wherein

the determination unit compares the plurality of distance information ranges corresponding to the plurality of parts with one another, thereby determining the part to be focused on from the plurality of distance information ranges.

17 . A training apparatus comprising at least one processor and/or at least one circuit which functions as:

an image capturing unit configured to generate an image through shooting;

a first detection unit configured to detect a subject region from the image; and

a second detection unit configured to detect a plurality of distance information pieces from a plurality of focus detection regions inside the subject region;

an inference unit configured to perform inference with use of a machine learning model based on the subject region including a subject within the image obtained through shooting, and on the plurality of distance information pieces detected from the plurality of focus detection regions inside the subject region, thereby generating an inference result indicating a distance information piece corresponding to the subject or a distance information range corresponding to the subject; and

a training unit configured to train the machine learning model so that the inference result approaches ground truth information to which a contribution made by one or more distance information pieces that are not based on the subject among the plurality of distance information pieces is suppressed.

18 . An inference method executed by an inference apparatus, comprising:

generating an image through shooting;

detecting a subject region from the image;

detecting a plurality of distance information pieces from a plurality of focus detection regions inside the subject region; and

performing inference with use of a machine learning model based on the subject region including a subject within the image obtained through shooting, and on the plurality of distance information pieces detected from the plurality of focus detection regions inside the subject region, thereby generating an inference result indicating a distance information piece corresponding to the subject or a distance information range corresponding to the subject,

wherein the machine learning model is a model that has been trained to suppress a contribution made to the inference result by one or more distance information pieces that are not based on the subject among the plurality of distance information pieces.

19 . A training method executed by a training apparatus, comprising:

generating an image through shooting;

detecting a subject region from the image;

detecting a plurality of distance information pieces from a plurality of focus detection regions inside the subject region;

performing inference with use of a machine learning model based on the subject region including a subject within the image obtained through shooting, and on the plurality of distance information pieces detected from the plurality of focus detection regions inside the subject region, thereby generating an inference result indicating a distance information piece corresponding to the subject or a distance information range corresponding to the subject; and

training the machine learning model so that the inference result approaches ground truth information to which a contribution made by one or more distance information pieces that are not based on the subject among the plurality of distance information pieces is suppressed.

20 . A non-transitory computer-readable storage medium which stores a program for causing a computer to execute an inference method comprising:

generating an image through shooting;

detecting a subject region from the image;

detecting a plurality of distance information pieces from a plurality of focus detection regions inside the subject region;

performing inference with use of a machine learning model based on the subject region including a subject within the image obtained through shooting, and on the plurality of distance information pieces detected from the plurality of focus detection regions inside the subject region, thereby generating an inference result indicating a distance information piece corresponding to the subject or a distance information range corresponding to the subject,

wherein the machine learning model is a model that has been trained to suppress a contribution made to the inference result by one or more distance information pieces that are not based on the subject among the plurality of distance information pieces.

21 . A non-transitory computer-readable storage medium which stores a program for causing a computer to execute a training method comprising:

generating an image through shooting;

detecting a subject region from the image;

detecting a plurality of distance information pieces from a plurality of focus detection regions inside the subject region;

performing inference with use of a machine learning model based on the subject region including a subject within the image obtained through shooting, and on the plurality of distance information pieces detected from the plurality of focus detection regions inside the subject region, thereby generating an inference result indicating a distance information piece corresponding to the subject or a distance information range corresponding to the subject; and

training the machine learning model so that the inference result approaches ground truth information to which a contribution made by one or more distance information pieces that are not based on the subject among the plurality of distance information pieces is suppressed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2024
From: HAMANO, HIDEYUKI; KANDA, AKIHIKO; SUGITANI, KUNIAKI; MATSUI, YOHEI
To: CANON KABUSHIKI KAISHA
Reel/Frame 066658/0848 →
Priority Claims (1)
JP 2023-024624 · Feb 20, 2023 · national
Continuity (1)
Related Publication 20240284047A1 · Aug 22, 2024
References Cited (20)
US 10666851B2 · Kikuchi · 2020 [cited by examiner]
US 11094075B1 · Lanman · 2021 [cited by examiner]
US 11468687B2 · Takami · 2022 [cited by examiner]
US 11627245B2 · Hoshino · 2023 [cited by examiner]
US 11758269B2 · Tsuji · 2023 [cited by examiner]
US 11900651B2 · Yoneyama · 2024 [cited by examiner]
US 12327368B2 · Komatsu · 2025 [cited by examiner]
US 20090146046A1 · Katsuda · 2009 [cited by examiner]
US 20180213149A1 · Ohtsubo · 2018 [cited by examiner]
US 20190296063A1 · Ichimiya · 2019 [cited by examiner]
US 20190387175A1 · Kikuchi · 2019 [cited by applicant]
US 20210182577A1 · Takami · 2021 [cited by examiner]
US 20210303846A1 · Yoneyama · 2021 [cited by examiner]
US 20220094840A1 · Hongu · 2022 [cited by examiner]
US 20220264022A1 · Tsuji · 2022 [cited by examiner]
US 20220294991A1 · Hoshino · 2022 [cited by applicant]
US 20220392096A1 · Komatsu · 2022 [cited by applicant]
JP H03167536A · 1991 [cited by applicant]
JP 2022137760A · 2022 [cited by applicant]
European Search Report issued on Jul. 8, 2024, that issued in the corresponding European Patent Application No. 24157792.3. [cited by applicant]