IP Library › Granted Patent US 12,724,821
Granted Patent B2
US 12,724,821 · App. 18/703,226 · Granted Sep 1, 2026

Method, apparatus, electronic device and readable medium for presenting

Inventor: Jiyuan Tian (Beijing, CN)
Assignee: Beijing Zitiao Network Technology Co., Ltd.
G06F16/538G06F16/532G06T7/11G06T7/20G06T11/00G06V10/764G06V40/11G06V40/28G10L15/04G10L15/08G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,821
App. No.
18/703,226
Granted
Sep 1, 2026
Kind
B2
Abstract

A method, apparatus, electronic device, and storage medium for presenting is provided. The method includes: determining operation information of a user from a video, the operation information comprising a motion trajectory of a hand of the user in the video (S 110 ); generating a profile of an object based on the motion trajectory (S 120 ); searching for corresponding information that matches the profile in a database (S 130 ); presenting an image based on the corresponding information (S 140 ).

Claims (81)

1 . A method for presenting, comprising:

determining operation information of a user from a video, the operation information comprising a motion trajectory of a hand of the user in the video;

generating a profile of an object based on the motion trajectory;

searching for corresponding information that matches the profile in a database; and

presenting an image based on the corresponding information,

wherein the determining operation information of a user comprises:

collecting a plurality of frames of images in the video;

performing semantic partitioning on the plurality of frames of images to extract a hand region in the plurality of frames of images;

generating the motion trajectory based on the hand region in the plurality of frames of images;

performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and

correcting the motion trajectory based on the non-hand region in the plurality of frames of images.

2 . The method of claim 1 , wherein searching for corresponding information that matches the profile in a predetermined database comprises:

determining a template object associated with the profile by a generative adversarial network (GAN); and

searching for corresponding information of the template object in the predetermined database.

3 . The method of claim 1 , further comprising:

performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and

correcting the motion trajectory based on the non-hand region in the plurality of frames of images.

4 . The method of claim 1 , further comprising:

recognizing, based on the hand region in the plurality of frames of images, at least one of a hand posture and a hand-held object; and

determining a category of the object based on the at least one of the hand postures and the hand-held object.

5 . The method of claim 1 , before searching for corresponding information that matches the profile in a predetermined database, the method further comprising:

recognizing a keyword from a speech stream of the user by an automatic speech recognition (ASR) model; and

determining a category of the object based on the keyword.

6 . The method of claim 4 , wherein the searching for corresponding information that matches the profile in a predetermined database comprises:

filtering corresponding information of a template object that is consistent with the category from the predetermined database; and

searching for the corresponding information that matches the profile from the corresponding information of the template object that is consistent with the category.

7 . The method of claim 5 , wherein the recognizing a keyword from a speech stream of the user by an ASR model comprises:

segmenting the speech stream to obtain segments of the speech stream, and storing the segments of the speech stream in a buffer;

recognizing keywords from respective ones of the segments by the ASR model, and determining confidences of the keywords; and

determining a keyword with a highest confidence as the keyword in the speech stream.

8 . The method of claim 1 , wherein the determining operation information of a user comprises:

collecting, by a real-time communication (RTC) module, an audio frame, and a video frame of the display process; and

transmitting, with a callback function, the video frame and the audio frame to an ASR module and a visual special effect module through a RTC channel.

9 . The method of claim 1 , after generating an image based on the corresponding information, the method further comprising:

performing traffic pushing, by Aiortc, on the image to a target device.

10 . The method of claim 1 , wherein the corresponding information comprises rendering information of the display object; and

generating an image based on the corresponding information, comprises:

rendering the profile based on the rendering information to obtain the image.

11 . The method of claim 1 , wherein the image comprises a template object and explanation information of the template object; and

after generating an image based on the corresponding information, the method further comprises:

displaying the template object in a first region and displaying the explanation information in a second region.

12 . An electronic device, comprising:

a processor; and

a storage device storing a program,

wherein the program, when executed by the processor, causes the processor to perform the method for presenting comprising:

determining operation information of a user from a video, the operation information comprising a motion trajectory of a hand of the user in the video;

generating a profile of an object based on the motion trajectory;

searching for corresponding information that matches the profile in a database; and

presenting an image based on the corresponding information,

wherein the determining operation information of a user comprises:

collecting a plurality of frames of images in the video;

performing semantic partitioning on the plurality of frames of images to extract a hand region in the plurality of frames of images;

generating the motion trajectory based on the hand region in the plurality of frames of images;

performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and

correcting the motion trajectory based on the non-hand region in the plurality of frames of images.

13 . A non-transitory computer-readable storage medium having a computer program thereon, the computer program, when executed by a processor, causing the method for presenting comprising:

determining operation information of a user from a video, the operation information comprising a motion trajectory of a hand of the user in the video;

generating a profile of an object based on the motion trajectory;

searching for corresponding information that matches the profile in a database; and

presenting an image based on the corresponding information,

wherein the determining operation information of a user comprises:

collecting a plurality of frames of images in the video;

performing semantic partitioning on the plurality of frames of images to extract a hand region in the plurality of frames of images;

generating the motion trajectory based on the hand region in the plurality of frames of images;

performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and

correcting the motion trajectory based on the non-hand region in the plurality of frames of images.

14 . The electronic device of claim 12 , wherein searching for corresponding information that matches the profile in a predetermined database comprises:

determining a template object associated with the profile by a generative adversarial network (GAN); and

searching for corresponding information of the template object in the predetermined database.

15 . The electronic device of claim 12 , wherein the method further comprises:

performing semantic partitioning on the plurality of frames of images to extract a non-hand region in the plurality of frames of images; and

correcting the motion trajectory based on the non-hand region in the plurality of frames of images.

16 . The electronic device of claim 12 , wherein the method further comprises:

recognizing, based on the hand region in the plurality of frames of images, at least one of a hand posture and a hand-held object; and

determining a category of the object based on the at least one of the hand postures and the hand-held object.

17 . The electronic device of claim 12 , wherein the method further comprises, before searching for corresponding information that matches the profile in a predetermined database:

recognizing a keyword from a speech stream of the user by an automatic speech recognition (ASR) model; and

determining a category of the object based on the keyword.

18 . The electronic device of claim 16 , wherein the searching for corresponding information that matches the profile in a predetermined database comprises:

filtering corresponding information of a template object that is consistent with the category from the predetermined database; and

searching for the corresponding information that matches the profile from the corresponding information of the template object that is consistent with the category.

Priority Claims (1)
CN 202111217427.3 · Oct 19, 2021 · national
Continuity (1)
Related Publication 20240419723A1 · Dec 19, 2024
References Cited (37)
US 8854486B2 · Tian · 2014 [cited by examiner]
US 11640698B2 · Li · 2023 [cited by examiner]
US 20100199232A1 · Mistry · 2010 [cited by examiner]
US 20100235285A1 · Hoffberg · 2010 [cited by examiner]
US 20100317420A1 · Hoffberg · 2010 [cited by examiner]
US 20140177909A1 · Lin · 2014 [cited by examiner]
US 20140184496A1 · Gribetz · 2014 [cited by examiner]
US 20160012465A1 · Sharp · 2016 [cited by examiner]
US 20170206797A1 · Solomon · 2017 [cited by examiner]
US 20170336874A1 · Chun · 2017 [cited by examiner]
US 20190096035A1 · Li · 2019 [cited by examiner]
US 20190179526A1 · Yellen · 2019 [cited by examiner]
US 20200160616A1 · Li · 2020 [cited by examiner]
US 20200387698A1 · Yi · 2020 [cited by examiner]
US 20210072877A1 · Kim · 2021 [cited by examiner]
US 20210279950A1 · Phalak · 2021 [cited by examiner]
CN 103034323A · 2013 [cited by applicant]
CN 103823554A · 2014 [cited by applicant]
CN 104866216A · 2015 [cited by applicant]
CN 106503756A · 2017 [cited by applicant]
CN 106873893A · 2017 [cited by applicant]
CN 107944457A · 2018 [cited by applicant]
CN 108133197A · 2018 [cited by applicant]
CN 108961414A · 2018 [cited by applicant]
CN 109271023A · 2019 [cited by applicant]
CN 110333785A · 2019 [cited by applicant]
CN 110688965A · 2020 [cited by examiner]
CN 111064987A · 2020 [cited by applicant]
CN 111107264A · 2020 [cited by applicant]
CN 112599127A · 2021 [cited by applicant]
CN 112766231A · 2021 [cited by applicant]
CN 113325952A · 2021 [cited by applicant]
KR 101662022B1 · 2016 [cited by applicant]
Office action received from Chinese patent application No. 202111217427.3 mailed on Jun. 6, 2025, 15 pages (7 pages English Translation and 8 pages Original Copy). [cited by applicant]
Luo Na, Research of OpenCV-based Natural Gesture Recognition and Interactive System, Dissertation Submitted to Guangdong University of Technology, Jun. 2012, pp. [cited by applicant]
Notification of grant received for Chinese patent application No. 202111217427.3 mailed on Feb. 3, 2026, 7 pages (2 pages English Translation and 5 pages Original Copy). [cited by applicant]
Reed, B. S., Reconceptualizing mirroring: Sound imitation and rapport in naturally occurring interaction, Journal of Pragmatics, vol. 167, Oct. 2020, pp. 131-151. [cited by applicant]