Method, system, and computer-readable medium for providing depression preliminary diagnosis information by using machine learning model
View Patent ↗The present invention relates to a method, a system, and a computer-readable medium for providing depression preliminary diagnosis information by using a machine learning model, in which a result of analysis on depression with respect to an answer video performed by a person to be evaluated and supporting information therefor are provided to medical staff, users, counselors and the like through a special user interface, so as to more efficiently determine the depression.
1 . A method for providing depression preliminary diagnosis information by using a machine learning model performed on a computing device having at least one processor and at least one memory, the method comprising:
deriving the depression preliminary diagnosis information by using a machine-learned model for at least one answer video captured using a camera and a microphone in a patient terminal; and
providing the depression preliminary diagnosis information to a user,
wherein the machine-learned model outputs, for each of a plurality of sections of the answer video, a degree of depression,
wherein a first display screen displayed by the providing the depression preliminary diagnosis information to the user includes:
an answer video layer for displaying the answer video;
a script layer for displaying question information related to the answer video and answer text information extracted from the answer video; and
a depression graph layer for displaying the degree of depression output according to a time axis,
wherein the depression graph layer includes:
a first reference display element for indicating that the degree of depression exceeds a predetermined first reference and a second reference display element indicating that the degree of depression exceeds a predetermined second reference higher than the predetermined first reference, and
wherein a video time point of the answer video layer is changeable to a time point corresponding to a position of a script part selected according to an input by the user in the script layer, and
the video time point of the answer video layer is changeable to a time point corresponding to a position on the time axis selected according to an input by the user in the depression graph layer;
wherein the deriving includes:
extracting multiple words, multiple image frames, and multiple pieces of voice information extracted from text of a voice from the answer video;
deriving multiple pieces of first feature information from the words, multiple pieces of second feature information from the image frames, and multiple pieces of third feature information from the pieces of voice information by using each detailed machine learning model or algorithm; and
deriving derived information including the degree of depression by using an artificial neural network considering sequence data from the pieces of first feature information, the pieces of second feature information, and the pieces of third feature information.
2 . The method of claim 1 , wherein the diagnosing includes deriving the depression preliminary diagnosis information by using the machine-learned model from the answer video, based on multiple words extracted from text of a voice, multiple image frames corresponding to the words, respectively, and multiple pieces of voice information corresponding to the words, respectively.
3 . The method of claim 1 , wherein the diagnosing includes:
deriving the depression preliminary diagnosis information by using the machine-learned model, based on at least two pieces of information among multiple words, multiple image frames, and voice information extracted from text of voice extracted from the answer video.
4 . The method claim 1 , wherein the artificial neural network considering the sequence data corresponds to a recurrent neural network or a transformer-based machine learning model based on an attention mechanism, and the first feature information, the second feature information, and the third feature information corresponding to each of the words are input, in a merged form, to the recurrent neural network or the transformer-based machine learning model.
5 . The method of claim 1 , wherein a part of the answer text displayed on the script layer is displayed with change in at least one of highlight, font, size, color, and underline according to at least one of the degree of depression and a type of depression in the depression preliminary diagnosis information corresponding to the part of the answer text.
6 . The method of claim 5 , wherein the part of the answer text displayed with the change in at least one of the highlight, font, size, color, and underline corresponds to a condition that the degree of depression of the depression preliminary diagnosis information corresponding to the part of the answer text is equal to or greater than a predetermined reference, and
the video time point of the answer video layer is changed to a time point corresponding to a position of the part of the answer text when the user selects the part of the answer text displayed with the change in at least one of the highlight, font, size, color, and underline.
7 . The method of claim 1 , wherein the first reference display element indicates a detailed type of the depression.
8 . The method of claim 1 , wherein the first display screen further includes a scroll layer for displaying scroll information of the script layer,
the scroll layer displays an information display element at each time point displayed according to at least one of the degree of depression and a type of depression in the depression preliminary diagnosis information, and
the time point is shifted to a position of answer text corresponding to the selected information display element in the script layer when the user selects the information display element.
9 . The method of claim 1 , wherein a second display screen displayed by the providing the depression preliminary diagnosis information to the user includes:
the answer video layer for displaying the answer video; and
a summary script layer for displaying summary answer text information extracted from the answer video, and
wherein the summary answer text information includes at least one part of the answer text in which the degree of depression in the depression preliminary diagnosis information corresponding to the part of the answer text is equal to or greater than a predetermined reference, and
the video time point of the answer video layer is changeable to a time point corresponding to a position of a script part selected according to an input by the user in the summary script layer.
10 . The method of claim 9 , wherein the second display screen further includes at least one selection input element corresponding to each of the summary answer text information, and
a part of an answer video corresponding to a part of the selection input element or the summary answer text information selected according to an input of the user is played.
11 . A device for providing depression preliminary diagnosis information by using a machine learning model implemented as a computing device having at least one processor and at least one memory, the device comprising:
a diagnosing unit for deriving the depression preliminary diagnosis information by using a machine-learned model for at least one answer video captured using a camera and a microphone in a patient terminal; and
a providing unit for providing the depression preliminary diagnosis information to a user,
wherein the machine-learned model outputs, for each of a plurality of sections of the answer video, a degree of depression,
wherein a first display screen displayed by the providing the depression preliminary diagnosis information to the user includes:
an answer video layer for displaying the answer video;
a script layer for displaying question information related to the answer video and answer text information extracted from the answer video; and
a depression graph layer for displaying the degree of depression output according to a time axis,
wherein the depression graph layer includes:
a first reference display element for indicating that the degree of depression exceeds a predetermined first reference; and
a second reference display element indicating that the degree of depression exceeds a predetermined second reference higher than the predetermined first reference, and
wherein a video time point of the answer video layer is changeable to a time point corresponding to a position of a script part selected according to an input by the user in the script layer, and
the video time point of the answer video layer is changeable to a time point corresponding to a position on the time axis selected according to an input by the user in the depression graph layer;
wherein the deriving includes:
extracting multiple words, multiple image frames, and multiple pieces of voice information extracted from text of a voice from the answer video;
deriving multiple pieces of first feature information from the words, multiple pieces of second feature information from the image frames, and multiple pieces of third feature information from the pieces of voice information by using each detailed machine learning model or algorithm; and
deriving derived information including the degree of depression by using an artificial neural network considering sequence data from the pieces of first feature information, the pieces of second feature information, and the pieces of third feature information.
12 . A non-transitory computer-readable medium recording program instructions for implementing a method for providing depression preliminary diagnosis information by using a machine learning model performed on a computing device having at least one processor and at least one memory, wherein the method comprising:
deriving the depression preliminary diagnosis information by using a machine-learned model for at least one answer video captured using a camera and a microphone in a patient terminal; and
providing the depression preliminary diagnosis information to a user,
wherein the machine-learned model outputs, for each of a plurality of sections of the answer video, a degree of depression,
wherein a first display screen displayed by the providing the depression preliminary diagnosis information to the user includes:
an answer video layer for displaying the answer video;
a script layer for displaying question information related to the answer video and answer text information extracted from the answer video; and
a depression graph layer for displaying the degree of depression output according to a time axis,
wherein the depression graph layer includes:
a first reference display element for indicating that the degree of depression exceeds a predetermined first reference; and
a second reference display element indicating that the degree of depression exceeds a predetermined second reference higher than the predetermined first reference, and
wherein a video time point of the answer video layer is changeable to a time point corresponding to a position of a script part selected according to an input by the user in the script layer, and
the video time point of the answer video layer is changeable to a time point corresponding to a position on the time axis selected according to an input by the user in the depression graph layer;
wherein the deriving includes:
extracting multiple words, multiple image frames, and multiple pieces of voice information extracted from text of a voice from the answer video;
deriving multiple pieces of first feature information from the words, multiple pieces of second feature information from the image frames, and multiple pieces of third feature information from the pieces of voice information by using each detailed machine learning model or algorithm; and
deriving derived information including the degree of depression by using an artificial neural network considering sequence data from the pieces of first feature information, the pieces of second feature information, and the pieces of third feature information.