IP Library › Granted Patent US 12,272,114
Granted Patent B2
US 12,272,114 · App. 17/192,973 · Granted Apr 8, 2025

Learning method, storage medium, and image processing device

Inventors: Nao Mishima (Inagi, JP); Masako Kashiwagi (Ageo, JP)
Assignee: KABUSHIKI KAISHA TOSHIBA
G06V10/56G06F18/2148G06N7/00G06N20/00G06T7/50G06T2207/20081G06V10/467
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,114
App. No.
17/192,973
Granted
Apr 8, 2025
Kind
B2
Abstract

According to one embodiment, a learning method of causing a statistical model for outputting a distance to a subject to learn by using an image including the subject as an input is provided. The method includes acquiring an image for learning including a subject having an already known shape, acquiring a first distance to the subject included in the image for learning, from the image for learning, and causing the statistical model to learn by restraining the first distance with the shape of the subject included in the image for learning.

Claims (55)

1. A learning method of causing a statistical model for outputting a distance to a subject to learn by using an image including the subject as an input, the method comprising:

acquiring an image for learning including a subject having an already known shape;

acquiring a first distance to the subject included in the image for learning, from the image for learning; and

causing the statistical model to learn by restraining the first distance with the shape of the subject included in the image for learning, wherein

the causing the statistical model to learn comprises:

correcting the first distance to a second distance, based on the shape of the subject included in the image for learning; and

causing the statistical model to learn the image for learning and the second distance,

the image for learning comprises a first image for learning to which a correct label is assigned and a second image for learning to which the correct label is not assigned,

the first and second images for learning include subjects of the same shape,

the acquiring the first distance comprises acquiring a first distance to the subject included in the first image for learning, from the second image for learning,

the first distance is corrected to a second distance, based on the shape of the subject included in the second image for learning, and

the causing the statistical model to learn comprises updating a parameter of the statistical model so as to minimize a value obtained by adding a difference between a relative value of the second distance and a relative value of the third distance output from the statistical model by inputting the second image for learning to the statistical model, to a difference between a correct label and the third distance output from the statistical model by inputting the first image for learning to the statistical model.

2. The method of claim 1 , wherein

the correcting comprises correcting the first distance to the second distance by fitting a parameter used to express the shape of the subject to the first distance.

3. The method of claim 2 , wherein

the shape of the subject is expressed by any function including the parameter.

4. The method of claim 1 , wherein

the causing the statistical model to learn comprises updating a parameter of the statistical model so as to minimize a difference between the second distance and a third distance output from the statistical model by inputting the image for learning to the statistical model.

5. The method of claim 1 , wherein

the causing the statistical model to learn comprises regularizing the statistical model.

6. The method of claim 5 , wherein

the regularizing the statistical model comprises updating a parameter of the statistical model so as to minimize a difference between a relative value of the second distance and a relative value of a third distance output from the statistical model by inputting the image for learning to the statistical model.

7. The method of claim 1 , wherein

the statistical model is generated by learning bokeh which occurs in an image affected by aberration of an optical system and which changes nonlinearly in accordance with the distance to the subject included in the image.

8. The method of claim 1 , wherein

the statistical model is generated by learning bokeh which occurs in an image generated by light passed through a filter and which changes nonlinearly in accordance with the distance to the subject included in the image.

9. The method of claim 1 , wherein

the acquiring comprises acquiring a distance output from the statistical model by inputting the image for learning to the statistical model.

10. The method of claim 1 , wherein

the acquiring comprises acquiring a distance, based on a marker assigned to the subject included in the image for learning.

11. A non-transitory computer-readable storage medium having stored thereon a computer program which is executable by a computer and causes a statistical model for outputting a distance to a subject to learn by using an image including the subject as an input, the computer program comprising instructions capable of causing the computer to execute functions of:

acquiring an image for learning including a subject having an already known shape;

acquiring a distance to the subject included in the image for learning, from the image for learning; and

causing the statistical model to learn by restraining the acquired distance with the shape of the subject included in the image for learning, wherein

the causing the statistical model to learn comprises:

correcting the first distance to a second distance, based on the shape of the subject included in the image for learning; and

causing the statistical model to learn the image for learning and the second distance,

the image for learning comprises a first image for learning to which a correct label is assigned and a second image for learning to which the correct label is not assigned,

the first and second images for learning include subjects of the same shape,

the acquiring the first distance comprises acquiring a first distance to the subject included in the first image for learning, from the second image for learning,

the first distance is corrected to a second distance, based on the shape of the subject included in the second image for learning, and

the causing the statistical model to learn comprises updating a parameter of the statistical model so as to minimize a value obtained by adding a difference between a relative value of the second distance and a relative value of the third distance output from the statistical model by inputting the second image for learning to the statistical model, to a difference between a correct label and the third distance output from the statistical model by inputting the first image for learning to the statistical model.

12. An image processing device for causing a statistical model for outputting a distance to a subject to learn by using an image including the subject as an input, the device comprising:

a processor configured to:

acquire the image for learning including a subject having an already known shape;

acquire a distance to the subject included in the image for learning, from the image for learning; and

cause the statistical model to learn by restraining the acquired distance with the shape of the subject included in the image for learning, wherein

the processor is configured to:

correct the first distance to a second distance, based on the shape of the subject included in the image for learning; and

cause the statistical model to learn the image for learning and the second distance,

the image for learning comprises a first image for learning to which a correct label is assigned and a second image for learning to which the correct label is not assigned,

the first and second images for learning include subjects of the same shape,

the processor is configured to acquire a first distance to the subject included in the first image for learning, from the second image for learning,

the first distance is corrected to a second distance, based on the shape of the subject included in the second image for learning, and

the processor is configured to update a parameter of the statistical model so as to minimize a value obtained by adding a difference between a relative value of the second distance and a relative value of the third distance output from the statistical model by inputting the second image for learning to the statistical model, to a difference between a correct label and the third distance output from the statistical model by inputting the first image for learning to the statistical model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2025
From: MISHIMA, NAO; KASHIWAGI, MASAKO
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 069865/0980 →
Priority Claims (1)
JP 2020-069159 · Apr 7, 2020 · national
Continuity (1)
Related Publication 20210312233A1 · Oct 7, 2021
References Cited (20)
US 20190014262A1 · Yamaguchi et al. · 2019 [cited by applicant]
US 20200036895A1 · Midorikawa · 2020 [cited by examiner]
US 20200051264A1 · Mishima et al. · 2020 [cited by applicant]
US 20200294260A1 · Kashiwagi et al. · 2020 [cited by applicant]
US 20210082146A1 · Kashiwagi et al. · 2021 [cited by applicant]
JP 2018132477A · 2018 [cited by applicant]
JP 201915575A · 2019 [cited by applicant]
JP 201929021A · 2019 [cited by applicant]
JP 202026990A · 2020 [cited by applicant]
JP 2020026990A · 2020 [cited by applicant]
JP 2020148483A · 2020 [cited by applicant]
JP 202143115A · 2021 [cited by applicant]
Sattar, “Human detection and distance estimation with monocular camera using YOLOv3 neural network”, University of Tartu 2019 (Year: 2019). [cited by examiner]
Shah “A Novel Local Surface Description for Automatic 3D Object Recognition in Low Resolution Cluttered Scenes”, IEEE 2013 ( Year: 2013). [cited by examiner]
Travnik, (“On Bokeh”, https://jtra.cz/stuff/essays/bokeh/index.html, originally published 2011) (Year: 2011). [cited by examiner]
Chatzitofis (“DeepMoCap: Deep Optical Motion Capture Using Multiple Depth Sensors and Retro-Reflectors”, MDPI 2018) (Year: 2018). [cited by examiner]
Sattar, “Human Detection and distance estimation with monocular camera using YOLOv3 neural network”, University of Tartu, Jun. 14, 2016, pp. 1-43. [cited by applicant]
Lee, “Pseudo-Label: The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks,” ICML 2013 Workshop: Challenges in Representation Learning (WREPL), 2013, 7 pages. [cited by applicant]
Mishima et al., “Physical Cue based Depth-Sensing by Color Coding with Deaberration Network”, BMVC, 2019, pp. 1-13, https://bmvc2019.org/wp-content/uploads/papers/0156-paper.pdf. [cited by applicant]
Romero-Ramirez et al., “Speeded Up Detection of Squared Fiducial Markers”, Image and Vision Computing, vol. 76, 2018, 14 pages, DOI: 10.1016/j.imavis.2018.05.004. [cited by applicant]