IP Library Granted Patent US 12682528
Granted Patent B2
US 12682528 · App. 18/251,182 · Granted Jul 14, 2026

Video processing method, video processing apparatus, and storage medium

Inventors: Zhili Chen (Los Angeles, CA); Linjie Luo (Los Angeles, CA); Jianchao Yang (Los Angeles, CA); Guohui Wang (Los Angeles, CA)
Assignee: BYTEDANCE INC.
G06T13/20G06T7/246G06T7/73G06V10/75G06V10/82G06V20/176G06V20/41G06V20/46G06V20/49G06V40/10G06T2207/10016G06T2207/20084G06T2207/30184G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682528
App. No.
18/251,182
Granted
Jul 14, 2026
Kind
B2
Abstract

The present disclosure provides a video processing method, a video processing apparatus, and a storage medium. A picture of the video includes a landmark building and a moving subject. The video processing method includes: identifying and tracking the landmark building in the video; extracting and tracking a key point of the moving subject in the video, and determining a posture of the moving subject based on information of the extracted key point of the moving subject; and making the key point of the moving subject correspond to the landmark building, and driving the landmark building to perform a corresponding action based on an action of the key point of the moving subject, so as to make a posture of the landmark building in the picture of the video correspond to the posture of the moving subject. The video processing method may enhance interactions between the user and the shot landmark.

Claims (47)

1 . A video processing method, wherein a picture of the video comprises a landmark building and a moving subject, the video processing method comprising:

identifying and tracking the landmark building in the video;

extracting and tracking a key point of the moving subject in the video, and determining a posture of the moving subject based on information of the extracted key point of the moving subject;

making the key point of the moving subject correspond to the landmark building, and driving the landmark building to perform a corresponding action based on an action of the key point of the moving subject, so as to make a posture of the landmark building in the picture of the video correspond to the posture of the moving subject, and

personifying the landmark building by means of a predefined animated material map, to make the landmark building have a feature of the moving subject.

2 . The video processing method of claim 1 , further comprising, prior to driving the landmark building to perform the corresponding action based on the action of the key point of the moving subject:

cutting out the landmark building from the picture of the video;

complementing, by a smooth interpolation algorithm, a background at the landmark building that has been cut out, based on pixels surrounding the landmark building in the picture of the video; and

restoring the landmark building to a position where the background has been complemented.

3 . The video processing method of claim 1 , wherein said making the key point of the moving subject correspond to the landmark building, and driving the landmark building to perform the corresponding action based on the action of the key point of the moving subject comprises:

mapping the key point of the moving subject onto the landmark building, to make the key point of the moving subject correspond to the landmark building, so that the landmark building follows the action of the key point of the moving subject to perform the corresponding action.

4 . The video processing method of claim 3 , wherein a spine line of the moving subject is mapped onto a central axis of the landmark building.

5 . The video processing method of claim 1 , wherein the key point of the moving subject in the picture of the video is extracted by a neural network model.

6 . The video processing method of claim 1 , wherein said identifying the landmark building in the picture of the video comprises:

extracting a feature point of the landmark building; and

matching the extracted feature point of the landmark building with a building feature point classification model to identify the landmark building.

7 . The video processing method of claim 1 , wherein said tracking the landmark building in the video and said tracking the key point of the moving subject in the video comprise:

detecting the landmark building and the key point of the moving subject in each frame of image in the video to track the landmark building and the key point of the moving subject.

8 . The video processing method of claim 1 , wherein in the same video, the key point of the moving subject is made to correspond to a plurality of landmark buildings, and the plurality of landmark buildings are driven based on the action of the key point of the moving subject to perform a corresponding action based on the action of the moving subject.

9 . The video processing method of claim 1 , wherein in the same video, key points of a plurality of moving subjects are made to respectively correspond to a plurality of landmark buildings, and the plurality of landmark buildings are driven respectively based on a plurality of actions of the key points of the plurality of moving subjects to perform corresponding actions respectively based on the actions of the moving subjects, so as to make postures of the plurality of landmark buildings correspond to postures of the plurality of moving subjects in a one-to-one correspondence.

10 . The video processing method of claim 1 , further comprising:

recording the video in real time by an image capturing apparatus, and processing each frame of image in the video in real time, so as to make the posture of the landmark building correspond to the posture of the moving subject.

11 . The video processing method of claim 1 , wherein the moving subject is a human body.

12 . A video processing apparatus, comprising:

a processor;

a memory; and

one or more computer program modules, wherein the one or more computer program modules are stored in the memory and configured to be executed by the processor, the one or more computer program modules comprising instructions to perform:

identifying and tracking a landmark building in the video, a picture of the video comprising the landmark building and a moving subject;

extracting and tracking a key point of the moving subject in the video, and determining a posture of the moving subject based on information of the extracted key point of the moving subject;

making the key point of the moving subject correspond to the landmark building, and driving the landmark building to perform a corresponding action based on an action of the key point of the moving subject, so as to make a posture of the landmark building in the picture of the video correspond to the posture of the moving subject, and

personifying the landmark building by means of a predefined animated material map, to make the landmark building have a feature of the moving subject.

13 . The video processing apparatus of claim 12 , the one or more computer program modules comprising instructions further to perform, prior to driving the landmark building to perform the corresponding action based on the action of the key point of the moving subject:

cutting out the landmark building from the picture of the video;

complementing, by a smooth interpolation algorithm, a background at the landmark building that has been cut out, based on pixels surrounding the landmark building in the picture of the video; and

restoring the landmark building to a position where the background has been complemented.

14 . The video processing apparatus of claim 12 , wherein said making the key point of the moving subject correspond to the landmark building, and driving the landmark building to perform the corresponding action based on the action of the key point of the moving subject comprises:

mapping the key point of the moving subject onto the landmark building, to make the key point of the moving subject correspond to the landmark building, so that the landmark building follows the action of the key point of the moving subject to perform the corresponding action.

15 . The video processing apparatus of claim 14 , wherein a spine line of the moving subject is mapped onto a central axis of the landmark building.

16 . The video processing apparatus of claim 12 , wherein the key point of the moving subject in the picture of the video is extracted by a neural network model.

17 . The video processing apparatus of claim 12 , wherein said identifying the landmark building in the picture of the video comprises:

extracting a feature point of the landmark building; and

matching the extracted feature point of the landmark building with a building feature point classification model to identify the landmark building.

18 . A storage medium non-transitorily storing computer-readable instructions, which, when executed by a computer, perform:

identifying and tracking a landmark building in the video, a picture of the video comprising the landmark building and a moving subject;

extracting and tracking a key point of the moving subject in the video, and determining a posture of the moving subject based on information of the extracted key point of the moving subject;

making the key point of the moving subject correspond to the landmark building, and driving the landmark building to perform a corresponding action based on an action of the key point of the moving subject, so as to make a posture of the landmark building in the picture of the video correspond to the posture of the moving subject, and

personifying the landmark building by means of a predefined animated material map, to make the landmark building have a feature of the moving subject.