Method, apparatus, electronic device and readable medium for presenting
A method, apparatus, electronic device, and storage medium for presenting is provided. The method includes: determining operation information of a user from a video, the operation information comprising a motion trajectory of a hand of the user in the video (S 110 ); generating a profile of an object based on the motion trajectory (S 120 ); searching for corresponding information that matches the profile in a database (S 130 ); presenting an image based on the corresponding information (S 140 ).
1 . A method for presenting, comprising:
determining operation information of a user from a video, the operation information comprising a motion trajectory of a hand of the user in the video;
generating a profile of an object based on the motion trajectory;
searching for corresponding information that matches the profile in a database; and
presenting an image based on the corresponding information,
wherein the determining operation information of a user comprises:
collecting a plurality of frames of images in the video;
performing semantic partitioning on the plurality of frames of images to extract a hand region in the plurality of frames of images;
generating the motion trajectory based on the hand region in the plurality of frames of images;
performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and
correcting the motion trajectory based on the non-hand region in the plurality of frames of images.
2 . The method of claim 1 , wherein searching for corresponding information that matches the profile in a predetermined database comprises:
determining a template object associated with the profile by a generative adversarial network (GAN); and
searching for corresponding information of the template object in the predetermined database.
3 . The method of claim 1 , further comprising:
performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and
correcting the motion trajectory based on the non-hand region in the plurality of frames of images.
4 . The method of claim 1 , further comprising:
recognizing, based on the hand region in the plurality of frames of images, at least one of a hand posture and a hand-held object; and
determining a category of the object based on the at least one of the hand postures and the hand-held object.
5 . The method of claim 1 , before searching for corresponding information that matches the profile in a predetermined database, the method further comprising:
recognizing a keyword from a speech stream of the user by an automatic speech recognition (ASR) model; and
determining a category of the object based on the keyword.
6 . The method of claim 4 , wherein the searching for corresponding information that matches the profile in a predetermined database comprises:
filtering corresponding information of a template object that is consistent with the category from the predetermined database; and
searching for the corresponding information that matches the profile from the corresponding information of the template object that is consistent with the category.
7 . The method of claim 5 , wherein the recognizing a keyword from a speech stream of the user by an ASR model comprises:
segmenting the speech stream to obtain segments of the speech stream, and storing the segments of the speech stream in a buffer;
recognizing keywords from respective ones of the segments by the ASR model, and determining confidences of the keywords; and
determining a keyword with a highest confidence as the keyword in the speech stream.
8 . The method of claim 1 , wherein the determining operation information of a user comprises:
collecting, by a real-time communication (RTC) module, an audio frame, and a video frame of the display process; and
transmitting, with a callback function, the video frame and the audio frame to an ASR module and a visual special effect module through a RTC channel.
9 . The method of claim 1 , after generating an image based on the corresponding information, the method further comprising:
performing traffic pushing, by Aiortc, on the image to a target device.
10 . The method of claim 1 , wherein the corresponding information comprises rendering information of the display object; and
generating an image based on the corresponding information, comprises:
rendering the profile based on the rendering information to obtain the image.
11 . The method of claim 1 , wherein the image comprises a template object and explanation information of the template object; and
after generating an image based on the corresponding information, the method further comprises:
displaying the template object in a first region and displaying the explanation information in a second region.
12 . An electronic device, comprising:
a processor; and
a storage device storing a program,
wherein the program, when executed by the processor, causes the processor to perform the method for presenting comprising:
determining operation information of a user from a video, the operation information comprising a motion trajectory of a hand of the user in the video;
generating a profile of an object based on the motion trajectory;
searching for corresponding information that matches the profile in a database; and
presenting an image based on the corresponding information,
wherein the determining operation information of a user comprises:
collecting a plurality of frames of images in the video;
performing semantic partitioning on the plurality of frames of images to extract a hand region in the plurality of frames of images;
generating the motion trajectory based on the hand region in the plurality of frames of images;
performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and
correcting the motion trajectory based on the non-hand region in the plurality of frames of images.
13 . A non-transitory computer-readable storage medium having a computer program thereon, the computer program, when executed by a processor, causing the method for presenting comprising:
determining operation information of a user from a video, the operation information comprising a motion trajectory of a hand of the user in the video;
generating a profile of an object based on the motion trajectory;
searching for corresponding information that matches the profile in a database; and
presenting an image based on the corresponding information,
wherein the determining operation information of a user comprises:
collecting a plurality of frames of images in the video;
performing semantic partitioning on the plurality of frames of images to extract a hand region in the plurality of frames of images;
generating the motion trajectory based on the hand region in the plurality of frames of images;
performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and
correcting the motion trajectory based on the non-hand region in the plurality of frames of images.
14 . The electronic device of claim 12 , wherein searching for corresponding information that matches the profile in a predetermined database comprises:
determining a template object associated with the profile by a generative adversarial network (GAN); and
searching for corresponding information of the template object in the predetermined database.
15 . The electronic device of claim 12 , wherein the method further comprises:
performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and
correcting the motion trajectory based on the non-hand region in the plurality of frames of images.
16 . The electronic device of claim 12 , wherein the method further comprises:
recognizing, based on the hand region in the plurality of frames of images, at least one of a hand posture and a hand-held object; and
determining a category of the object based on the at least one of the hand postures and the hand-held object.
17 . The electronic device of claim 12 , wherein the method further comprises, before searching for corresponding information that matches the profile in a predetermined database:
recognizing a keyword from a speech stream of the user by an automatic speech recognition (ASR) model; and
determining a category of the object based on the keyword.
18 . The electronic device of claim 16 , wherein the searching for corresponding information that matches the profile in a predetermined database comprises:
filtering corresponding information of a template object that is consistent with the category from the predetermined database; and
searching for the corresponding information that matches the profile from the corresponding information of the template object that is consistent with the category.