IP Library Granted Patent US 12,675,982
Granted Patent B2
US 12,675,982 · App. 18/368,600 · Granted Jul 7, 2026

Video processing method using transfer learning and pre-training server

Inventors: Sa Im Shin (Seoul, KR); Jung Ho Kim (Seoul, KR); Bo Eun Kim (Seoul, KR)
Assignee: KOREA ELECTRONICS TECHNOLOGY INSTITUTE
G06V10/7747G06T7/251G06V10/255G06V10/82G06V10/95G06V20/46G06V40/23G06T2207/20081G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,982
App. No.
18/368,600
Filed
Sep 15, 2023
Granted
Jul 7, 2026
Kind
B2
Examiner
LU, ZHIYU
Art Unit
2665
USPC
382/103
Abstract

There is provided a video processing method performed by a computing device, the method including the steps of: collecting video from an external device; generating preprocessed data by extracting two-dimensional or three-dimensional skeleton information from the video; pre-training a first artificial intelligence model including N transformer blocks from the preprocessed data by applying an attention from a body of an object to a plurality of joints, an attention from each of the plurality of joints to the body, and an attention between persons; and learning, when parameters determined as a result of the pre-training of the first artificial intelligence model are transferred, a method of recognizing an action from the video received from the external device, using a second artificial intelligence model including the N transformer blocks on the basis of the parameters, wherein N is a natural number equal to or larger than 2.

Claims (30)

1 . A video processing method performed by a computing device, the method comprising the steps of:

collecting video from an external device;

generating preprocessed data by extracting two-dimensional or three-dimensional skeleton information from the video;

pre-training a first artificial intelligence model including N transformer blocks from the preprocessed data by applying an attention from a body of an object to a plurality of joints, an attention from each of the plurality of joints to the body, and an attention between persons; and

learning, when parameters determined as a result of the pre-training of the first artificial intelligence model are transferred, a method of recognizing an action from the video received from the external device, using a second artificial intelligence model including the N transformer blocks on the basis of the parameters, wherein

N is a natural number equal to or larger than 2.

2 . The method according to claim 1 , wherein the step of pre-training a first artificial intelligence model includes the steps of:

applying a slice of a positional embedding tensor corresponding to each frame of an output of a previous block, and performing first layer normalization, by an n-th transformer block among the N transformer blocks;

applying a spatial multi-head attention (MHA) to a result of the first layer normalization, by the n-th transformer block; and

deriving a first result by adding a result of applying the slice of a positional embedding tensor corresponding to each frame of an output of a previous block to a result of applying the spatial MHA, by the n-th transformer block, wherein

n is a natural number between 2 and N.

3 . The method according to claim 2 , further comprising the steps of:

applying a matrix with a changed dimension of the positional embedding tensor to a pose sequence matrix in which the first result corresponding to each frame is stacked, and performing second layer normalization;

applying a temporal MHA to a result of the second layer normalization; and

deriving a second result by adding a result of applying the matrix with a changed dimension of the positional embedding tensor to a pose sequence matrix to a result of applying the temporal MHA.

4 . The method according to claim 3 , further comprising the steps of:

performing third layer normalization on the second result;

applying a multi-layer perceptron (MLP) to a result of the third layer normalized; and

deriving a third result by adding the second result to a result of applying the MLP.

5 . The method according to claim 1 , further comprising the step of deriving a motion sequence representation for an input motion sequence according to the video, by an N-th block among the N transformer blocks.

6 . A pre-training server comprising:

a processor;

a memory; and

a computer program loaded on the memory and executed by the processor, wherein

the computer program includes:

an instruction for collecting video from an external device;

an instruction for generating preprocessed data by extracting two-dimensional or three-dimensional skeleton information from the video;

an instruction for pre-training a first artificial intelligence model including N transformer blocks from the preprocessed data by applying an attention from a body of an object to a plurality of joints, an attention from each of the plurality of joints to the body, and an attention between persons; and

an instruction for learning, when parameters determined as a result of the pre-training of the first artificial intelligence model are transferred, a method of recognizing an action from the video received from the external device, using a second artificial intelligence model including the N transformer blocks on the basis of the parameters, wherein

N is a natural number equal to or larger than 2.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: SHIN, SA IM; KIM, JUNG HO; KIM, BO EUN
To: KOREA ELECTRONICS TECHNOLOGY INSTITUTE
Reel/Frame 064915/0158 →
Priority Claims (1)
KR 10-2022-0118520 · Sep 20, 2022 · national
Continuity (1)
Related Publication 20240096071A1 · Mar 21, 2024
References Cited (20)
US 11270124B1 · Carvalho · 2022 [cited by examiner]
US 11482048B1 · Diaz-Arias · 2022 [cited by examiner]
US 20220051061A1 · Chi · 2022 [cited by examiner]
US 20220383639A1 · Javan Roshtkhari · 2022 [cited by examiner]
US 20230196841A1 · Min · 2023 [cited by examiner]
US 20230368578A1 · Diaz-Arias · 2023 [cited by examiner]
CN 111002292 · 2020 [cited by examiner]
CN 113408455 · 2021 [cited by examiner]
CN 113425290 · 2021 [cited by examiner]
CN 114386582 · 2022 [cited by examiner]
CN 114708649 · 2022 [cited by examiner]
CN 114818989 · 2022 [cited by examiner]
CN 112257534 · 2022 [cited by examiner]
CN 114998525 · 2022 [cited by examiner]
CN 115147676 · 2022 [cited by examiner]
CN 115359550 · 2022 [cited by examiner]
CN 113240714 · 2023 [cited by examiner]
CN 114821804 · 2025 [cited by examiner]
JP WO2023066536 · 2023 [cited by examiner]
JP 2023111554 · 2023 [cited by examiner]