IP Library Granted Patent US 12,333,785
Granted Patent B2
US 12,333,785 · App. 17/620,608 · Granted Jun 17, 2025

Learning data generation device, learning data generation method, and program

Inventors: Dan Mikami (Tokyo, JP); Mariko Isogawa (Tokyo, JP); Hiroko Yabushita (Tokyo, JP); Yoshinori Kusachi (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06V10/774G06N5/041G06N20/00G06T7/00G06T7/20G06V10/757G06V20/42G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,785
App. No.
17/620,608
Granted
Jun 17, 2025
Kind
B2
Abstract

A learning data generation device for generating learning data for learning a recognizer capable of estimating a contour of a sphere making spinning motion, with high accuracy, the sphere being recorded in a single camera video image, is provided. The learning data generation device includes: a spinning rate estimation unit that receives an input of a learning video image in which motion of a spinning sphere is recorded and an initial value of a size of a contour of the recorded sphere in the video image, sets a plurality of set values of the size of the contour based on the initial value, and obtains an estimated value of a spinning rate of the sphere based on the learning video image, for each of the set values; a contour determination unit that receives an input of a true value of the spinning rate of the sphere, the true value being obtained in advance for the learning video image, and determines at least any of a plurality of the set values respectively corresponding to a plurality of the estimated values selected in order of closeness to the true value, as a determined value of the contour; and a learning data output unit that outputs the learning video image and the determined value as learning data.

Claims (21)

1. A learning data generation device comprising:

processing circuitry configured to

receive an input of a learning video image in which motion of a spinning sphere is recorded and an initial value of a size of a contour of the recorded sphere in the video image, set a plurality of set values of the size of the contour based on the initial value, and obtain an estimated value of a spinning rate of the sphere based on the learning video image, for each of the set values;

receive an input of a true value of the spinning rate of the sphere, the true value being obtained in advance for the learning video image, and determine at least one of a plurality of the set values respectively corresponding to a plurality of the estimated values selected in order of closeness to the true value, as a determined value of the contour; and

output the learning video image and the determined value as learning data.

2. The learning data generation device according to claim 1 , wherein

the processing circuitry estimates a spinning state of the sphere by, using the learning video image at a time t and the learning video image at a time t+t c where t c is a predetermined integer of no less than 1, selecting, from among a plurality of hypotheses of the spinning state, a hypothesis of the spinning state, such that a likelihood of an image of the sphere resulting from the sphere in the learning video image at a certain time being spun for t c unit time based on the hypothesis of the spinning state is high.

3. The learning data generation device according to claim 2 , wherein

the processing circuitry estimates the spinning state of the sphere by, using the learning video images at times t 1 , t 2 , . . . , t K and the learning video images at times t 1 +t c , t 2 +t c , . . . , t K +t c , selecting, from among a plurality of hypotheses of the spinning state, a hypothesis of the spinning state, such that a likelihood of an image of the sphere resulting from the sphere in the learning video images at the times t 1 , t 2 , . . . , t K being spun for t c unit time based on the hypothesis of the spinning state is high.

4. The learning data generation device according to claim 2 , wherein

the processing circuitry repeatedly perform processing for, for each of the plurality of hypotheses of the spinning state, calculating a likelihood of an image of the sphere resulting from the sphere in the learning video image at the time t or the learning video images at the times t 1 , t 2 , . . . , t K being spun for t c unit time based on the hypothesis of the spinning state, and processing for newly generating a plurality of likely hypotheses of the spinning state based on the calculated likelihoods.

5. The learning data generation device according to claim 4 , wherein the processing for newly generating a plurality of likely hypotheses of the spinning state based on the calculated likelihoods, the processing being performed by the processing circuitry, is processing for newly generating a plurality of hypotheses by repeating, a plurality of times, processing for determining a hypothesis, from among the plurality of hypotheses of the spinning state, in such a manner that a hypothesis, based on which the calculated likelihood of the hypothesis is higher, is determined with a higher probability and determining a spinning state having a value obtained by addition of a random number to a value of the spinning state in the determined hypothesis, as a new hypothesis.

6. A learning data generation method comprising:

a step of receiving an input of a learning video image in which motion of a spinning sphere is recorded and an initial value of a size of a contour of the recorded sphere in the video image, setting a plurality of set values of the size of the contour based on the initial value, and obtaining an estimated value of a spinning rate of the sphere based on the learning video image, for each of the set values;

a step of receiving an input of a true value of the spinning rate of the sphere, the true value being obtained in advance for the learning video image, and determining at least one of a plurality of the set values respectively corresponding to a plurality of the estimated values selected in order of closeness to the true value, as a determined value of the contour; and

a learning data step of outputting the learning video image and the determined value as learning data.

7. A non-transitory computer readable medium that stores a program that causes a computer to function as the learning data generation device according to claim 1 .

8. A non-transitory computer readable medium that stores a program that causes a computer to function as the learning data generation device according to claim 2 .

9. A non-transitory computer readable medium that stores a program that causes a computer to function as the learning data generation device according to claim 3 .

10. A non-transitory computer readable medium that stores a program that causes a computer to function as the learning data generation device according to claim 4 .

11. A non-transitory computer readable medium that stores a program that causes a computer to function as the learning data generation device according to claim 5 .

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0623 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2021
From: MIKAMI, DAN; ISOGAWA, MARIKO; YABUSHITA, HIROKO; KUSACHI, YOSHINORI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 058423/0242 →
Continuity (1)
Related Publication 20220375203A1 · Nov 24, 2022
References Cited (14)
US 10118696B1 · Hoffberg · 2018 [cited by examiner]
US 11230375B1 · Hoffberg · 2022 [cited by examiner]
US 11712637B1 · Hoffberg · 2023 [cited by examiner]
US 12118672B1 · Reinsel · 2024 [cited by examiner]
US 20210008417A1 · Tuxen · 2021 [cited by examiner]
CN 109087336A · 2018 [cited by examiner]
CN 110458281A · 2019 [cited by examiner]
JP 2002202317A · 2002 [cited by examiner]
JP 2014182032A · 2014 [cited by examiner]
CN 109087336 A (machine translation) (Year: 2018). [cited by examiner]
CN 110458281 A (machine translation) (Year: 2019). [cited by examiner]
JP 2014182032 A (machine translation) (Year: 2014). [cited by examiner]
JP 2002202317 A (machine translation) (Year: 2002). [cited by examiner]
He et al. (2017) “Mask R-CNN”, IEEE International Conference on Computer Vision (ICCV). [cited by applicant]