Electronic device and controlling method of electronic device
An electronic device and a method of controlling the electronic device are provided. The electronic device includes a microphone, a display, memory, and a processor configured to obtain, based on a user voice being received through the microphone while a first user interface (UI) screen is being displayed, text information corresponding to the user voice by inputting the user voice, obtain first information including information on a command and information on an execution target of the command included in the text information, obtain second information including information on functions corresponding to the plurality of objects and information on texts included in the plurality of objects, identify whether a target object corresponding to the user voice is present from among the plurality of objects, and control the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.
1 . An electronic device comprising:
a microphone;
a display;
memory storing instructions; and
at least one processor communicatively coupled to the microphone, the display, and the memory,
wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:
obtain, based on a user voice being received through the microphone while a first user interface (UI) screen comprising a plurality of objects is being displayed on the display, text information corresponding to the user voice by inputting the user voice to a voice recognition model,
obtain, based on the text information, first information comprising information on a command and information on an execution target of the command comprised in the text information,
identify a variable area in which information displayed within the first UI screen is changeable and a constant area which is different from the variable area based on metadata on the first UI screen,
obtain, based on the metadata on the first UI screen, second information comprising information on functions corresponding to at least one object comprised in the constant area among the plurality of objects and information on texts comprised in the at least one object comprised in the constant area among the plurality of objects,
identify, based on a comparison result between the first information and the second information, whether a target object corresponding to the user voice is present from among the plurality of objects, and
control, based on the target object being identified, the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.
2 . The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to:
identify, based on the information on the execution target corresponding to a first text comprised in the plurality of objects, the target object comprised with the first text from among the plurality of objects as the target object.
3 . The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to:
identify, based on the information on the execution target corresponding to a second text comprised in the plurality of objects, and the command corresponding to one function from among a plurality of functions corresponding to the plurality of objects, an object corresponding to the one function from among the plurality of objects as the target object.
4 . The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to:
control, based on the target object not being identified, the display to maintain displaying of the first UI screen.
5 . The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to:
control, based on the target object not being identified, the display to display a third UI screen for performing an operation corresponding to the command.
6 . The electronic device of claim 1 ,
wherein a UI graph comprising a plurality of nodes corresponding to a type of a plurality of UI screens and a plurality of edges showing a connection relationship between the plurality of nodes according to an operation performed by a conversion between the plurality of UI screens is stored for each application in the memory, and
wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to:
obtain, based on the target object being identified, a first embedding vector corresponding to the second information based on the metadata on the first UI screen,
identify a first node corresponding to the first UI screen from among the plurality of nodes based on the comparison result between the first information and the second information and a comparison result between the first embedding vector corresponding to the second information and a second embedding vector corresponding to the plurality of nodes, respectively,
identify, based on the information on the command and the information on the execution target of the command, a second node for performing the operation corresponding to the command from among at least one node which is connected to the first node, and
control the display to display the second UI screen corresponding to the second node.
7 . A method performed by an electronic device, the method comprising:
obtaining, based on a user voice being received while a first user interface (UI) screen comprising a plurality of objects is displayed on a display of the electronic device, text information corresponding to the user voice by inputting the user voice in a voice recognition model;
obtaining, based on the text information, first information comprising information on a command and information on an execution target of the command comprised in the text information;
identifying a variable area in which information displayed within the first UI screen is changeable and a constant area which is different from the variable area based on metadata on the first UI screen;
obtaining, based on the metadata on the first UI screen, second information comprising information on functions corresponding to at least one object comprised in the constant area among the plurality of objects and information on texts comprised in the at least one object comprised in the constant area among the plurality of objects;
identifying, based on a comparison result between the first information and the second information, whether a target object corresponding to the user voice is present from among the plurality of objects; and
controlling, based on the target object being identified, the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.
8 . The method of claim 7 , wherein the identifying of whether the target object is present comprises identifying, based on the information on the execution target corresponding to a first text comprised in the plurality of objects, an object comprised with the first text from among the plurality of objects as the target object.
9 . The method of claim 7 , wherein the identifying of whether the target object is present comprises identifying, based on the information on the execution target corresponding to a second text comprised in the plurality of objects, and the command corresponding to one function from among a plurality of functions corresponding to the plurality of objects, an object corresponding to the one function from among the plurality of objects as the target object.
10 . The method of claim 7 , further comprising:
controlling, based on the target object not being identified, the display to maintain displaying of the first UI screen.
11 . The method of claim 7 , further comprising:
controlling, based on the target object not being identified, the display to display a third UI screen for performing an operation corresponding to the command.
12 . The method of claim 7 , further comprising:
obtaining a UI graph comprising a plurality of nodes corresponding to a type of a plurality of UI screens and a plurality of edges showing a connection relationship between the plurality of nodes according to an operation performed by a conversion between the plurality of UI screens;
obtaining, based on the target object being identified, a first embedding vector corresponding to the second information based on the metadata on the first UI screen;
identifying a first node corresponding to the first UI screen from among the plurality of nodes based on the comparison result between the first information and the second information and a comparison result between the first embedding vector corresponding to the second information and a second embedding vector corresponding to the plurality of nodes, respectively;
identifying, based on information on the command and information on the execution target of the command, a second node for performing the operation corresponding to the command from among at least one node which is connected to the first node; and
controlling the display to display the second UI screen corresponding to the second node.
13 . One or more non-transitory computer-readable storage media storing computer-executable instructions that, when executed by at least one processor of an electronic device individually or collectively, cause the electronic device to perform operations, the operations comprising:
obtaining, based on a user voice being received while a first user interface (UI) screen comprising a plurality of objects is displayed on a display of the electronic device, text information corresponding to the user voice by inputting the user voice in a voice recognition model;
obtaining, based on the text information, first information comprising information on a command and information on an execution target of the command comprised in the text information;
identifying a variable area in which information displayed within the first UI screen is changeable and a constant area which is different from the variable area based on metadata on the first UI screen;
obtaining, based on the metadata on the first UI screen, second information comprising information on functions corresponding to at least one object comprised in the constant area among the plurality of objects and information on texts comprised in the at least one object comprised in the constant area among the plurality of objects;
identifying, based on a comparison result between the first information and the second information, whether a target object corresponding to the user voice is present from among the plurality of objects; and
controlling, based on the target object being identified, the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.
14 . The one or more non-transitory computer-readable storage media of claim 13 , wherein the identifying of whether the target object is present comprises identifying, based on the information on the execution target corresponding to a first text comprised in the plurality of objects, an object comprised with the first text from among the plurality of objects as the target object.
15 . The one or more non-transitory computer-readable storage media of claim 13 , wherein the identifying of whether the target object is present comprises identifying, based on the information on the execution target corresponding to a second text comprised in the plurality of objects, and the command corresponding to one function from among a plurality of functions corresponding to the plurality of objects, an object corresponding to the one function from among the plurality of objects as the target object.
16 . The one or more non-transitory computer-readable storage media of claim 13 , the operations further comprising:
controlling, based on the target object not being identified, the display to maintain displaying of the first UI screen.
17 . The one or more non-transitory computer-readable storage media of claim 13 , the operations further comprising:
controlling, based on the target object not being identified, the display to display a third UI screen for performing an operation corresponding to the command.
18 . The one or more non-transitory computer-readable storage media of claim 13 , the operations further comprising:
obtaining a UI graph comprising a plurality of nodes corresponding to a type of a plurality of UI screens and a plurality of edges showing a connection relationship between the plurality of nodes according to an operation performed by a conversion between the plurality of UI screens;
obtaining, based on the target object being identified, a first embedding vector corresponding to the second information based on the metadata on the first UI screen;
identifying a first node corresponding to the first UI screen from among the plurality of nodes based on the comparison result between the first information and the second information and a comparison result between the first embedding vector corresponding to the second information and a second embedding vector corresponding to the plurality of nodes, respectively;
identifying, based on information on the command and information on the execution target of the command, a second node for performing the operation corresponding to the command from among at least one node which is connected to the first node; and
controlling the display to display the second UI screen corresponding to the second node.