Learning apparatus, estimation apparatus, learning method, estimation method, and program and non-transitory storage medium
The present invention provides a learning apparatus ( 10 ) including an acquisition unit ( 11 ) that acquires learning data associating a training image including a person with a correct answer label indicating a position of each person, a correct answer label indicating whether each of a plurality of keypoints of a body of each of the persons is visible in the training image, and a correct answer label indicating a position, within the training image, of the keypoint being visible in the training image among a plurality of the keypoints, and a learning unit ( 12 ) that learns, based on the learning data, an estimation model.
1 . A learning apparatus comprising:
at least one memory configured to store one or more instructions; and
at least one processor configured to execute the one or more instructions to:
acquire learning data associating a training image including at least one person with a correct answer label indicating a position of each of the at least one person, a correct answer label indicating whether each of a plurality of keypoints of a body of each of the at least one person is visible in the training image, and a correct answer label indicating a position, within the training image, of a keypoint being visible in the training image among a plurality of the keypoints; and
learn, based on the learning data, an estimation model that estimates information indicating a position of each of the at least one person, information indicating whether each of a plurality of the keypoints of each of the at least one person included in a processing image is visible in the processing image, and information being related to a position of each of keypoints for computing a position, within the processing image, of the keypoint being visible in the processing image,
wherein the estimation model does not output the information being related to a position of each of the keypoints for computing a position, within the processing image, regarding a subset of the plurality of keypoints indicated to be invisible in the information indicating whether each of the plurality of the keypoints of each of the at least one person included in the processing image is visible in the processing image.
2 . The learning apparatus according to claim 1 , wherein,
in the correct answer label, a position, within the training image, of the keypoint being invisible in the training image is not indicated.
3 . The learning apparatus according to claim 1 , wherein
the at least one processor is further configured to execute the one or more instructions to
estimate, based on the estimation model being learned, information indicating a position of each of the at least one person, information indicating whether each of a plurality of the keypoints of each of the at least one person included in the processing image is visible in the processing image, and information being related to a position of each of keypoints for computing a position of each of the plurality of the keypoints within the training image,
adjust a first parameter of the estimation model in such a way as to minimize a difference between an estimation result of information indicating a position of each of the at least one person and information indicating a position of each of the at least one person indicated by the correct answer label,
adjust a second parameter of the estimation model in such a way as to minimize a difference between an estimation result of information indicating whether each of a plurality of the keypoints of each of the at least one person included in the processing image is visible in the processing image, and information indicating whether each of a plurality of the keypoints of a body of each of the at least one person indicated by the correct answer label is visible in the training image, and
adjust a third parameter of the estimation model in such a way as to minimize a difference between an estimation result of information being related to a position of each of keypoints for computing a position of each of the plurality of the keypoints within the training image, and information being related to a position of each of keypoints acquired from a position, within the training image, of the keypoint being visible in the training image among a plurality of the keypoints indicated by the correct answer label, for only a keypoint being visible in the training image indicated by the correct answer label.
4 . The learning apparatus according to claim 1 , wherein
the correct answer label further indicates a state of each invisible keypoint for each of the at least one person in the training image, and
the estimation model further estimates the state of each of the invisible keypoints for each of the at least one person in the processing image.
5 . The learning apparatus according to claim 4 , wherein
the state includes a state of being located outside an image, a state of being located within an image but hidden by another object, and a state of being located within an image but hidden by an own part.
6 . The learning apparatus according to claim 4 , wherein
the state indicates a number of objects hiding the keypoint being invisible in the training image or the processing image.
7 . An estimation apparatus comprising:
at least one memory configured to store one or more instructions; and
at least one processor configured to execute the one or more instructions to:
estimate a position, within a processing image, of each of a plurality of keypoints of each of the at least one person included in the processing image, by using an estimation model learned by the learning apparatus according to claim 1 .
8 . The estimation apparatus according to claim 7 , wherein
the at least one processor is further configured to execute the one or more instructions to estimate, by using the estimation model, whether each of a plurality of the keypoints of each of the at least one person included in the processing image is visible in the processing image, and estimate, by using a result of the estimation, a position, within the processing image, of each of a plurality of keypoints for each of the at least one person included in the processing image.
9 . The estimation apparatus according to claim 8 , wherein the at least one processor is further configured to execute the one or more instructions to;
output a type of an invisible keypoint for each of the at least one person, based on the estimate whether each of the plurality of keypoints of each of the at least one person included in the processing image is visible in the processing image, or
represent a type of the invisible keypoint as an object modeled on a person and display the object for each of the at least one person.
10 . The estimation apparatus according to claim 8 , wherein the at least one processor is further configured to execute the one or more instructions to:
determine an invisible keypoint, based on the estimate as to whether each of a plurality of the keypoints of each of the at least one person included in the processing image is visible in the processing image, determine a visible keypoint being directly connected to the determined invisible keypoint, based on a previously defined connection relation of a plurality of keypoints to a person, and
estimate a position of the determined invisible keypoint in the processing image, based on a position of the determined visible keypoint within the processing image.
11 . The estimation apparatus according to claim 7 , wherein
compute information indicating, for each estimated person, at least one of a degree at which a body of a person is visible in the processing image, and a degree at which a body of a person is hidden in the processing image, based on at least one of a number of the keypoints estimated to be visible in the processing image and a number of keypoints estimated to be invisible in the processing image, with respect to each estimated person.
12 . The estimation apparatus according to claim 11 , wherein
the at least one processor is further configured to execute the one or more instructions to display, for each of the at least one person, information indicating at least one of the computed degree at which a body of a person is visible, and the computed degree at which a body of a person is hidden, based on a center position of each of the at least one person or a specified keypoint position.
13 . The estimation apparatus according to claim 11 , wherein the at least one processor is further configured to execute the one or more instructions to:
convert, into information indicating hiding absent or hiding present for each of the at least one person, based on a specified threshold value, information indicating at least one of the computed degree at which a body of a person is visible, and the computed degree at which a body of a person is hidden, and
display the information indicating hiding absent or hiding present for each of the at least one person, based on a center position of each of the at least one person or a specified keypoint position.
14 . The estimation apparatus according to claim 7 , wherein the at least one processor is further configured to execute the one or more instructions to:
compute a maximum value for each of the at least one person in a number of objects hiding each keypoint for each of the at least one person,
compute the computed maximum value as a state of a way of overlapping for each of the at least one person, and
display, for each of the at least one person, the computed state of a way of overlapping for each of the at least one person, based on a center position of each of the at least one person or a position of a specified keypoint, or display a keypoint on a person with a color corresponding to a state of a way of overlapping for each of the at least one person.
15 . An estimation method of executing,
by a computer,
estimating a position, within a processing image, of each of a plurality of keypoints of each of the at least one person included in the processing image, by using an estimation model learned by the learning apparatus according to claim 1 .
16 . A non-transitory storage medium storing a program causing a computer to:
estimate a position, within a processing image, of each of a plurality of keypoints of each of the at least one person included in the processing image, by using an estimation model learned by the learning apparatus according to claim 1 .
17 . A learning method executed by a computer, the method comprising:
acquiring learning data associating a training image including at least one person with a correct answer label indicating a position of each of the at least one person, a correct answer label indicating whether each of a plurality of keypoints of a body of each of the at least one person is visible in the training image, and a correct answer label indicating a position, within the training image, of the keypoint being visible in the training image among a plurality of the keypoints; and
learning, based on the learning data, an estimation model that estimates information indicating a position of each of the at least one person, information indicating whether each of a plurality of the keypoints of each of the at least one person included in a processing image is visible in the processing image, and information being related to a position of each of keypoints for computing a position, within the processing image, of the keypoint being visible in the processing image,
wherein the estimation model does not output the information being related to a position of each of the keypoints for computing a position, within the processing image, regarding a subset of the plurality of keypoints indicated to be invisible in the information indicating whether each of the plurality of the keypoints of each of the at least one person included in the processing image is visible in the processing image.
18 . A non-transitory storage medium storing a program causing a computer to execute instructions, the instructions comprising:
acquire learning data associating a training image including at least one person with a correct answer label indicating a position of each of the at least one person, a correct answer label indicating whether each of a plurality of keypoints of a body of each of the at least one person is visible in the training image, and a correct answer label indicating a position, within the training image, of the keypoint being visible in the training image among a plurality of the keypoints; and
learn, based on the learning data, an estimation model that estimates information indicating a position of each of the at least one person, information indicating whether each of a plurality of the keypoints of each of the at least one person included in a processing image is visible in the processing image, and information being related to a position of each of keypoints for computing a position, within the processing image, of the keypoint being visible in the processing image,
wherein the estimation model does not output the information being related to a position of each of the keypoints for computing a position, within the processing image, regarding a subset of the plurality of keypoints indicated to be invisible in the information indicating whether each of the plurality of the keypoints of each of the at least one person included in the processing image is visible in the processing image.