Key point smoothing method based on frame up-sampling
There is provided a frame up-sampling-based key point smoothing method. The key point smoothing method according to an embodiment up-samples frames based on which key points are extracted, and smooths the up-sampled frames. Accordingly, key points are generated by smoothing key point extraction frames after up-sampling, so that a time-series error occurring when key points are extracted can be reduced and time-series stability can be enhanced, and quality of an application service provided subsequently can be improved.
1 . A key point generation method comprising:
extracting facial key points on a frame basis;
temporally up-sampling key point extraction frames based on which the facial key points are extracted on a time axis; and
smoothing the temporally up-sampled key point extraction frames on the time axis,
wherein the smoothing comprises smoothing a predetermined number of frames at predetermined intervals, wherein the predetermined interval is determined based on a frame rate before up-sampling, wherein the predetermined number is determined based on an increase rate of a frame rate caused by up-sampling, and wherein the method includes reducing a time-series error of the facial key points based on a result of the smoothing.
2 . The key point generation method of claim 1 , further comprising acquiring a speech signal,
wherein the extracting comprises extracting facial key points from the acquired speech signal on a frame basis.
3 . The key point generation method of claim 2 , wherein the extracting comprises inputting a speech signal to a machine learning model that is trained to extract facial key points from a speech signal, and extracting the facial key points.
4 . The key point generation method of claim 1 , further comprising providing an application service by using the smoothed key point extraction frames.
5 . A key point generation system comprising:
one or more processors configured to:
extract facial key points on a frame basis;
temporally up-sample key point extraction frames based on which the facial key points are extracted on a time axis; and
smooth the temporally up-sampled key point extraction frames on the time axis,
wherein the smoothing comprises smoothing a predetermined number of frames at predetermined intervals, wherein the predetermined interval is determined based on a frame rate before up-sampling, wherein the predetermined number is determined based on an increase rate of a frame rate caused by up-sampling, and wherein the one or more processors are configured to reduce a time-series error of the facial key points based on a result of the smoothing.
6 . The system of claim 5 , wherein the one or more processors are further configured to acquire a speech signal and extract facial key points from the acquired speech signal on a frame basis.
7 . The system of claim 6 , wherein, for the extracting, the one or more processors are configured to input a speech signal to a machine learning model that is trained to extract facial key points from a speech signal, and extract the facial key points.
8 . The system of claim 5 , wherein the one or more processors are configured to provide an application service by using the smoothed key point extraction frames.
9 . A key point smoothing method comprising:
temporally up-sampling key point extraction frames based on which facial key points are extracted on a frame basis on a time axis; and
smoothing the temporally up-sampled key point extraction frames on the time axis,
wherein the smoothing comprises smoothing a predetermined number of frames at predetermined intervals,
wherein the predetermined interval is determined based on a frame rate before up-sampling,
wherein the predetermined number is determined based on an increase rate of a frame rate caused by up-sampling, and
wherein the method includes reducing a time-series error of the facial key points based on a result of the smoothing.