IP Library › Granted Patent US 12,608,916
Granted Patent B2
US 12,608,916 · App. 17/740,211 · Granted Apr 21, 2026

Data augmentation based on attention

Inventors: Song Bai (Singapore, CN); Jieneng Chen (Beijing, CN); Shuyang Sun (Beijing, CN); Ju He (Beijing, CN); Bin Lu (Los Angeles, CA)
Assignee: LEMON INC.
G06V10/7715G06N20/00G06V10/26G06V10/48G06V10/774
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,916
App. No.
17/740,211
Granted
Apr 21, 2026
Kind
B2
Abstract

Implementations of the present disclosure relate to methods, devices, and computer program products for data augmentation. In the method, mixed data is generated from first data and second data, and the mixed data comprises a first portion from the first data and a second portion from the second data. An attention map is obtained for the mixed data based on distributions of the first and second portions in the mixed data, here the attention map describes contributions of the first and second data to the mixed data. A label is determined for the mixed data based on the attention map and a first label for the first data and a second label for the second data. With these implementations, the label is determined based on the contributions of the first and second images in an accurate and effective way, and thus has a value that is much closer to the ground true.

Claims (68)

1 . A method for data augmentation, comprising:

generating mixed data from first data and second data, the mixed data comprising a first portion from the first data and a second portion from the second data;

dividing the mixed data into a plurality of data blocks based on a mask describing a boundary between the first and second portions in the mixed data, wherein dividing the mixed data further comprises dividing the mixed data based on any of:

a size of any of the first and second data, and the mixed data;

a size of any of the first and second portions;

a predetermined size; and

a predetermined number;

obtaining an attention map for the plurality of data blocks in the mixed data based on distributions of the first and second portions in the mixed data, the attention map describing contributions of the first and second data to the mixed data; and

determining a label for the mixed data based on a first label for the first data, a second label for the second data, and the attention map.

2 . The method of claim 1 , wherein obtaining the attention map comprises:

determining a plurality of data tokens for the plurality of data blocks, respectively; and

obtaining the attention map based on a self-attention operation to a token sequence that comprises the plurality of data tokens and a class token associated with the attention map.

3 . The method of claim 2 , wherein obtaining the attention map comprises:

determining a query parameter and a key parameter for the self-attention operation based on the token sequence; and

obtaining the attention map based on the query and key parameters for the self-attention operation.

4 . The method of claim 3 , wherein obtaining the attention map based on the query and key parameters for the self-attention operation comprises:

obtaining a self-attention map for the token sequence during the self-attention operation; and

identifying a dimension, corresponding to the class token, in the self-attention map as the attention map.

5 . The method of claim 2 , wherein

the mixed data is generated by pasting the first portion into the second data based on the mask.

6 . The method of claim 2 , further comprising: generating a training sample based on the mixed data and the label, the training sample being for training a data model describing an association relationship between data and a label for the data.

7 . The method of claim 6 , wherein the first data is a first image and the second data is a second image, the self-attention operation is implemented by a vision transformer, and generating the training sample further comprises:

receiving a group of images that are used for training the data model; and

selecting the first and second images from the group of images for generating the training sample.

8 . The method of claim 1 , wherein determining the first weight comprises:

down-sampling the mask into a format corresponding to the number of the plurality of data blocks; and

determining the first weight based on the attention map and the down-sampled mask.

9 . The method of claim 1 , wherein determining the label for the mixed data comprises:

determining a first weight for the first label based on the attention map and the mask; and

determining the label based on the first and second labels and the first weight.

10 . An electronic device, comprising a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method for data augmentation, comprising:

generating mixed data from first data and second data, the mixed data comprising a first portion from the first data and a second portion from the second data;

dividing the mixed data into a plurality of data blocks based on a mask describing a boundary between the first and second portions in the mixed data, wherein dividing the mixed data further comprises dividing the mixed data based on any of:

a size of any of the first and second data, and the mixed data;

a size of any of the first and second portions;

a predetermined size; and

a predetermined number;

obtaining an attention map for the plurality of data blocks in the mixed data based on distributions of the first and second portions in the mixed data, the attention map describing contributions of the first and second data to the mixed data; and

determining a label for the mixed data based on a first label for the first data, a second label for the second data, and the attention map.

11 . The device of claim 10 , wherein obtaining the attention map comprises:

determining a plurality of data tokens for the plurality of data blocks, respectively; and

obtaining the attention map based on a self-attention operation to a token sequence that comprises the plurality of data tokens and a class token associated with the attention map.

12 . The device of claim 11 , wherein obtaining the attention map comprises:

determining a query parameter and a key parameter for the self-attention operation based on the token sequence; and

obtaining the attention map based on the query and key parameters for the self-attention operation.

13 . The device of claim 12 , wherein obtaining the attention map based on the query and key parameters for the self-attention operation comprises:

obtaining a self-attention map for the token sequence during the self-attention operation; and

identifying a dimension, corresponding to the class token, in the self-attention map as the attention map.

14 . The device of claim 11 , wherein

the mixed data is generated by pasting the first portion into the second data based on the mask.

15 . The device of claim 11 , wherein the first data is a first image and the second data is a second image, the self-attention operation is implemented by a vision transformer, and the method further comprising:

selecting the first and second images from a group of images that are used for training a data model, the data model describing an association relationship between data and a label for the data;

generating a training sample for training the data model based on the mixed data and the label.

16 . The device of claim 10 , wherein determining the first weight comprises:

down-sampling the mask into a format corresponding to the number of the plurality of data blocks; and

determining the first weight based on the attention map and the down-sampled mask.

17 . The device of claim 10 , wherein determining the label for the mixed data comprises:

determining a first weight for the first label based on the attention map and the mask; and

determining the label based on the first and second labels and the first weight.

18 . A computer program product, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform a method for data augmentation, the method comprising:

generating mixed data from first data and second data, the mixed data comprising a first portion from the first data and a second portion from the second data;

dividing the mixed data into a plurality of data blocks based on a mask describing a boundary between the first and second portions in the mixed data, wherein dividing the mixed data further comprises dividing the mixed data based on any of:

a size of any of the first and second data, and the mixed data;

a size of any of the first and second portions;

a predetermined size; and

a predetermined number;

obtaining an attention map for the plurality of data blocks in the mixed data based on distributions of the first and second portions in the mixed data, the attention map describing contributions of the first and second data to the mixed data; and

determining a label for the mixed data based on a first label for the first data, a second label for the second data, and the attention map.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2026
From: BYTEDANCE INC.; BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.; TIKTOK PTE. LTD.; SHANGHAI SUIXUNTONG ELECTRONIC TECHNOLOGY CO., LTD.
To: LEMON INC.
Reel/Frame 074349/0956 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2026
From: BAI, SONG; LU, BIN; HE, JU
To: TIKTOK PTE. LTD.; BYTEDANCE INC.; BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 074152/0892 →
Continuity (1)
Related Publication 20220270353A1 · Aug 25, 2022
References Cited (16)
US 11961281B1 · Kim · 2024 [cited by examiner]
US 20190065817A1 · Mesmakhosroshahi · 2019 [cited by examiner]
US 20190205761A1 · Wu · 2019 [cited by examiner]
US 20210027098A1 · Ge · 2021 [cited by examiner]
US 20210319255A1 · Pham · 2021 [cited by examiner]
US 20210357684A1 · Amirghodsi · 2021 [cited by examiner]
US 20220036194A1 · Sundaresan · 2022 [cited by examiner]
US 20220336058A1 · Merchant · 2022 [cited by examiner]
US 20230113643A1 · Mittal · 2023 [cited by examiner]
US 20230289967A1 · Jalal · 2023 [cited by examiner]
US 20240038252A1 · Fan · 2024 [cited by examiner]
CN 113887610A · 2022 [cited by examiner]
CN 114997395A · 2022 [cited by examiner]
Sangdoo et al., “Cutmix: Regularization strategy to train strong classifiers with localizable features,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6023-6032 (2019). [cited by applicant]
Zhang et al, “Mixup: Beyond empirical risk minimization,” ICLR (2018). [cited by applicant]
Vaswani et al., “Attention is all you need,” NeurIPS, pp. 5998-6008 (2017). [cited by applicant]