IP Library › Granted Patent US 12,548,244
Granted Patent B2
US 12,548,244 · App. 18/139,929 · Granted Feb 10, 2026

Method and apparatus for processing action of virtual object, and storage medium

Inventors: Kai Tian (Beijing, CN); Wei Chen (Beijing, CN); Xuefeng Su (Beijing, CN)
Assignee: BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO., LTD.
G06T17/00G06V10/24G06V10/44G06V10/54G06V10/761H04N5/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,244
App. No.
18/139,929
Granted
Feb 10, 2026
Kind
B2
Abstract

A method and apparatus for processing an action of a virtual object, and a storage medium are provided. The method specifically includes: receiving an action instruction, the action instruction including: an action identifier and time-dependent information of performing an action associated with the action identifier; determining an action video frame sequence corresponding to the action identifier; determining, from the action video frame sequence, an action state image corresponding to a preset state image of the virtual object at a target time, the target time being determined according to the time-dependent information; generating a connection video frame sequence according to the action state image, the connection video frame sequence connecting the preset state image with the action video frame sequence; and splicing the connection video frame sequence with the action video frame sequence, to obtain an action video. Embodiments of this application can improve action processing efficiency of a virtual object.

Claims (65)

1 . A method for processing an action of a virtual object performed by a computer device, the method comprising:

receiving an action instruction, the action instruction comprising: an action identifier and time-dependent information of performing an action associated with the action identifier;

determining, among a plurality of pre-generated action video frame sequences, each action video frame sequence corresponding to an action performed by the virtual object and having a corresponding action identifier, an action video frame sequence corresponding to the action identifier;

determining, from the action video frame sequence, an action state image corresponding to a preset state image of the virtual object at a target time, the target time being determined according to the time-dependent information, further including:

comparing a visual feature corresponding to the preset state image with a visual feature corresponding to a candidate action state image in the action video frame sequence to obtain a match value between the preset state image and the candidate action state image; and

choosing the candidate action state image having a maximum match value with the preset state image as the action state image corresponding to the preset state image of the virtual object at the target time;

obtaining a pair of images including an aligned preset state image and an aligned action state image by performing pose information alignment on the preset state image and the action state image that further improves a matching degree between the virtual object in the preset state image and the virtual object in the action state image;

generating a connection video frame sequence according to the aligned preset state image and the aligned action state image, the connection video frame sequence representing a transition from a preset state corresponding to the preset state image to an action state corresponding to the action state image and connecting the preset state image with the action video frame sequence; and

splicing the connection video frame sequence with the action video frame sequence, to obtain an action video.

2 . The method according to claim 1 , wherein the generating a connection video frame sequence according to the action state image comprises:

determining optical flow features separately corresponding to the action state image; and

generating the connection video frame sequence according to the optical flow features.

3 . The method according to claim 1 , wherein the generating a connection video frame sequence according to the action state image comprises:

determining optical flow features and texture features and/or deep features separately corresponding to the action state image; and

generating the connection video frame sequence according to the optical flow features and the texture features and/or the deep features.

4 . The method according to claim 1 , wherein the pose information includes position information and posture information of the virtual object in the preset state image and the action state image.

5 . The method according to claim 1 , further comprising:

extracting a part preset state image from the preset state image, and determining, on the basis of three-dimensional reconstruction, a third visual feature corresponding to the part preset state image;

extracting a part action state image from the action state image, and determining, on the basis of three-dimensional reconstruction, a fourth visual feature corresponding to the part action state image;

generating a part connection video frame sequence according to the third visual feature and the fourth visual feature; and

adding the part connection video frame sequence to the connection video frame sequence.

6 . The method according to claim 1 , wherein the time-dependent information comprises: text information corresponding to the action identifier.

7 . A computer device, comprising a processor and a memory, the memory storing a program, the program, when executed by the processor, causing the computer device to perform a method for processing an action of a virtual object including:

receiving an action instruction, the action instruction comprising: an action identifier and time-dependent information of performing an action associated with the action identifier;

determining, among a plurality of pre-generated action video frame sequences, each action video frame sequence corresponding to an action performed by the virtual object and having a corresponding action identifier, an action video frame sequence corresponding to the action identifier;

determining, from the action video frame sequence, an action state image corresponding to a preset state image of the virtual object at a target time, the target time being determined according to the time-dependent information, further including:

comparing a visual feature corresponding to the preset state image with a visual feature corresponding to a candidate action state image in the action video frame sequence to obtain a match value between the preset state image and the candidate action state image; and

choosing the candidate action state image having a maximum match value with the preset state image as the action state image corresponding to the preset state image of the virtual object at the target time;

obtaining a pair of images including an aligned preset state image and an aligned action state image by performing pose information alignment on the preset state image and the action state image that further improves a matching degree between the virtual object in the preset state image and the virtual object in the action state image;

generating a connection video frame sequence according to the aligned preset state image and the aligned action state image, the connection video frame sequence representing a transition from a preset state corresponding to the preset state image to an action state corresponding to the action state image and connecting the preset state image with the action video frame sequence; and

splicing the connection video frame sequence with the action video frame sequence, to obtain an action video.

8 . The computer device according to claim 7 , wherein the generating a connection video frame sequence according to the action state image comprises:

determining optical flow features separately corresponding to the action state image; and

generating the connection video frame sequence according to the optical flow features.

9 . The computer device according to claim 7 , wherein the generating a connection video frame sequence according to the action state image comprises:

determining optical flow features and texture features and/or deep features separately corresponding to the action state image; and

generating the connection video frame sequence according to the optical flow features and the texture features and/or the deep features.

10 . The computer device according to claim 7 , wherein the pose information includes position information and posture information of the virtual object in the preset state image and the action state image.

11 . The computer device according to claim 7 , wherein the method further comprises:

extracting a part preset state image from the preset state image, and determining, on the basis of three-dimensional reconstruction, a third visual feature corresponding to the part preset state image;

extracting a part action state image from the action state image, and determining, on the basis of three-dimensional reconstruction, a fourth visual feature corresponding to the part action state image;

generating a part connection video frame sequence according to the third visual feature and the fourth visual feature; and

adding the part connection video frame sequence to the connection video frame sequence.

12 . The computer device according to claim 7 , wherein the time-dependent information comprises: text information corresponding to the action identifier.

13 . A non-transitory computer-readable storage medium, which stores a program, the program, when executed by one or more processors, causing the computer device to perform a method for processing an action of a virtual object including:

receiving an action instruction, the action instruction comprising: an action identifier and time-dependent information of performing an action associated with the action identifier;

determining, among a plurality of pre-generated action video frame sequences, each action video frame sequence corresponding to an action performed by the virtual object and having a corresponding action identifier, an action video frame sequence corresponding to the action identifier;

determining, from the action video frame sequence, an action state image corresponding to a preset state image of the virtual object at a target time, the target time being determined according to the time-dependent information, further including:

comparing a visual feature corresponding to the preset state image with a visual feature corresponding to a candidate action state image in the action video frame sequence to obtain a match value between the preset state image and the candidate action state image; and

choosing the candidate action state image having a maximum match value with the preset state image as the action state image corresponding to the preset state image of the virtual object at the target time;

obtaining a pair of images including an aligned preset state image and an aligned action state image by performing pose information alignment on the preset state image and the action state image that further improves a matching degree between the virtual object in the preset state image and the virtual object in the action state image;

generating a connection video frame sequence according to the aligned preset state image and the aligned action state image, the connection video frame sequence representing a transition from a preset state corresponding to the preset state image to an action state corresponding to the action state image and connecting the preset state image with the action video frame sequence; and

splicing the connection video frame sequence with the action video frame sequence, to obtain an action video.

14 . The non-transitory computer-readable storage medium according to claim 13 , wherein the generating a connection video frame sequence according to the action state image comprises:

determining optical flow features separately corresponding to the action state image; and

generating the connection video frame sequence according to the optical flow features.

15 . The non-transitory computer-readable storage medium according to claim 13 , wherein the generating a connection video frame sequence according to the action state image comprises:

determining optical flow features and texture features and/or deep features separately corresponding to the action state image; and

generating the connection video frame sequence according to the optical flow features and the texture features and/or the deep features.

16 . The non-transitory computer-readable storage medium according to claim 13 , wherein the pose information includes position information and posture information of the virtual object in the preset state image and the action state image.

17 . The non-transitory computer-readable storage medium according to claim 13 , wherein the method further comprises:

extracting a part preset state image from the preset state image, and determining, on the basis of three-dimensional reconstruction, a third visual feature corresponding to the part preset state image;

extracting a part action state image from the action state image, and determining, on the basis of three-dimensional reconstruction, a fourth visual feature corresponding to the part action state image;

generating a part connection video frame sequence according to the third visual feature and the fourth visual feature; and

adding the part connection video frame sequence to the connection video frame sequence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2023
From: SU, XUEFENG
To: BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO., LTD.
Reel/Frame 063566/0184 →
Priority Claims (1)
CN 202110770548.4 · Jul 7, 2021 · national
Continuity (2)
Continuation PCTCN2022100369 · Jun 22, 2022
Related Publication 20230368461A1 · Nov 16, 2023
References Cited (41)
US 9875242B2 · Poiesz · 2018 [cited by examiner]
US 10911775B1 · Zhu et al. · 2021 [cited by applicant]
US 20120180084A1 · Huang et al. · 2012 [cited by applicant]
US 20130060572A1 · Garland · 2013 [cited by examiner]
US 20130066816A1 · Shimomura et al. · 2013 [cited by applicant]
US 20140176603A1 · Kumar · 2014 [cited by examiner]
US 20140310595A1 · Acharya · 2014 [cited by examiner]
US 20160037067A1 · Lee et al. · 2016 [cited by applicant]
US 20180349108A1 · Brebner · 2018 [cited by examiner]
US 20200380769A1 · Liu et al. · 2020 [cited by applicant]
US 20210118151A1 · Abdelhak · 2021 [cited by examiner]
US 20220319169A1 · Sim · 2022 [cited by examiner]
CN 104038705A · 2014 [cited by applicant]
CN 107529091A · 2017 [cited by applicant]
CN 108304762A · 2018 [cited by applicant]
CN 108320021A · 2018 [cited by applicant]
CN 108665492A · 2018 [cited by applicant]
CN 109637518A · 2019 [cited by applicant]
CN 110148406A · 2019 [cited by applicant]
CN 110347867A · 2019 [cited by applicant]
CN 110378247A · 2019 [cited by applicant]
CN 111369687A · 2020 [cited by applicant]
CN 111508064A · 2020 [cited by applicant]
CN 111833439A · 2020 [cited by applicant]
CN 112040327A · 2020 [cited by applicant]
CN 112101196A · 2020 [cited by applicant]
CN 112233210A · 2021 [cited by applicant]
CN 113642394A · 2021 [cited by applicant]
WO WO2017124116A1 · 2017 [cited by examiner]
Stoll, C., Gall, J., De Aguiar, E., Thrun, S., & Theobalt, C. (2010). Video-based reconstruction of animatable human characters. ACM Transactions on Graphics (TOG), 29(6), 1-10. [cited by examiner]
Yao, B. Z., Nie, B. X., Liu, Z., & Zhu, S. C. (2013). Animated pose templates for modeling and detecting human actions. IEEE transactions on pattern analysis and machine intelligence, 36(3), 436-452. [cited by examiner]
Gleicher, M., Shin, H. J., Kovar, L., & Jepsen, A. (2008). Snap-together motion: assembling run-time animations. In ACM SIGGRAPH 2008 classes (pp. 1-9). [cited by examiner]
Tencent Technology, Extended European Search Report, EP Patent Application No. 22836714.0, Jul. 29, 2024, 8 pgs. [cited by applicant]
Ashish Verma et al., “Animating Expressive Faces Across Languages”, IEEE Transactions on Multimedia, vol. 6, No., 6, DOI: 10.1109/TMM.2004.837256, Dec. 2004, 10 pgs. [cited by applicant]
Feng Xu et al., “Video-Based Characters—Creating New Human Performances from a Multi-View Video Database”, Special Interest Group on Computer Graphics and Interactive Techniques Conference (SIGGRAPH), DOI: 10.1145/19649… [cited by applicant]
Masaki Oshita, “Generating Animation from Natural Language Texts and Framework of Motion Database”, IEEE, 2009 International Conference on CyberWorlds, DOI: 10.1109/CW.2009.46, Sep. 2009, 8 pgs. [cited by applicant]
Tencent Technology, ISR, PCT/CN2022/100369, Aug. 9, 2022, 3 pgs. [cited by applicant]
Adrien Gaidon et al., “Virtual Worlds as Proxy for Multi-Object Tracking Analysis”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), http://www.xrce xerox.com/Research-Development/Co… [cited by applicant]
Tencent Technology, WO, PCT/CN2022/100369, Aug. 9, 2022, 5 pgs. [cited by applicant]
Tencent Technology, IPRP, PCT/CN2022/100369, Dec. 14, 2023 6 pgs. [cited by applicant]
Beijing Sogou Technology Development Co., Ltd., Indian Office Action, IN Patent Application No. 202337060424, May 7, 2025, 8 pgs. [cited by applicant]