IP Library › Granted Patent US 12,450,841
Granted Patent B2
US 12,450,841 · App. 18/304,078 · Granted Oct 21, 2025

Embeddings representing visual augmentations

Inventors: Zhenpeng Zhou (Newark, CA); Patrick Poirson (Gilbert, AZ); Maksim Gusarov (Marina del Rey, CA); Chen Wang (Great Neck, NY); Oleg Tovstyi (Los Angeles, CA)
Assignee: Snap Inc.
G06T19/006G06T1/0021G06V10/761H04N5/2621
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,841
App. No.
18/304,078
Granted
Oct 21, 2025
Kind
B2
Abstract

An input video item that includes a target visual augmentation is accessed. A machine learning model uses the input video item to generate an embedding. The embedding may comprise a vector representation of a visual effect of the target visual augmentation. The machine learning model is trained, in an unsupervised training phase, to minimize loss between training video representations generated within each of a plurality of training sets. Each training set comprises a plurality of different training video items that each include a predefined visual augmentation. Based on the generation of the embedding of the input video item, the target visual augmentation is mapped to an augmentation identifier.

Claims (45)

1. A system comprising:

at least one processor;

at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

accessing an input video item that includes a target visual augmentation;

generating, by a machine learning model, an embedding of the input video item, the machine learning model being trained, in an unsupervised training phase, to minimize loss between training video representations generated within each of a plurality of training sets, each training set comprising a plurality of training video items having different video content and that each include the same predefined visual augmentation; and

mapping, based on the generation of the embedding of the input video item, the target visual augmentation to an augmentation identifier.

2. The system of claim 1 , wherein the embedding of the input video item comprises a vector representation of a visual effect of the target visual augmentation within the input video item.

3. The system of claim 1 , wherein the unsupervised training phase comprises positive-only, self-supervised contrastive learning.

4. The system of claim 1 , wherein each training set comprises a pair of training video items, and wherein the unsupervised training phase comprises:

accessing a first training video item of a first pair of training video items, the first training video item comprising first video content to which a first predefined visual augmentation is applied;

accessing a second training video item of the first pair of training video items, the second training video item comprising second video content to which the first predefined visual augmentation is applied;

generating a vector representation of the first training video item and a vector representation of the second training video item; and

generating a transformed representation based on the vector representation of the first training video item, wherein the minimizing of loss between training video representations comprises, for the first pair of training video items, utilizing a loss function to measure a similarity between the transformed representation and the vector representation of the second training video item.

5. The system of claim 4 , wherein the unsupervised training phase further comprises:

automatically updating parameters of the machine learning model to minimize the loss between training video representations within subsequent pairs of training video items based on the loss function.

6. The system of claim 4 , wherein the first video content is a first plate video and the second video content is a second plate video, the first plate video being different from the second plate video.

7. The system of claim 6 , the operations further comprising:

applying, by an augmentation rendering component, the first predefined visual augmentation to the first plate video; and

applying, by the augmentation rendering component, the first predefined visual augmentation to the second plate video.

8. The system of claim 4 , wherein the generating of the transformed representation comprises transforming, by a predictor component of the machine learning model, the vector representation of the first training video item into the transformed representation to create a final prediction target for the machine learning model.

9. The system of claim 1 , wherein the input video item is a first input video item and the embedding of the first input video item is a first embedding, the operations further comprising:

accessing a second input video item that includes the target visual augmentation;

generating, by the machine learning model, a second embedding of the second input video item; and

generating, based on the first embedding and the second embedding, an aggregated embedding, the aggregated embedding representing a visual effect of the target visual augmentation.

10. The system of claim 9 , wherein the generating of the aggregated embedding comprises determining a mean of the first embedding and the second embedding.

11. The system of claim 9 , wherein the augmentation identifier is the aggregated embedding.

12. The system of claim 1 , the operations further comprising:

automatically comparing the augmentation identifier with an embedding of a reference visual augmentation to determine a degree of similarity between the target visual augmentation and the reference visual augmentation.

13. The system of claim 1 , wherein the machine learning model is a first machine learning model, the operations further comprising:

transmitting the augmentation identifier to a second machine learning model, the second machine learning model comprising at least one of an augmentation ranking model, an augmentation retrieval model, an augmentation deduplication model, an augmentation tagging model, or an augmentation clustering model.

14. The system of claim 1 , the operations further comprising:

querying a database to identify the target visual augmentation based on the augmentation identifier.

15. The system of claim 1 , the operations further comprising:

receiving the input video item from a user device, the target visual augmentation having been applied to the input video item by a user of the user device.

16. The system of claim 1 , wherein the target visual augmentation is an augmented reality effect.

17. The system of claim 16 , wherein the augmented reality effect is a video filter.

18. The system of claim 1 , wherein the input video item is in a binary file format.

19. A method comprising:

accessing an input video item that includes a target visual augmentation;

generating, by a machine learning model, an embedding of the input video item, the machine learning model being trained, in an unsupervised training phase, to minimize loss between training video representations generated within each of a plurality of training sets, each training set comprising a plurality of training video items having different video content and that each include the same predefined visual augmentation; and

mapping, based on the generation of the embedding of the input video item, the target visual augmentation to an augmentation identifier.

20. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

accessing an input video item that includes a target visual augmentation;

generating, by a machine learning model, an embedding of the input video item, the machine learning model being trained, in an unsupervised training phase, to minimize loss between training video representations generated within each of a plurality of training sets, each training set comprising a plurality of training video items having different video content and that each include the same predefined visual augmentation; and

mapping, based on the generation of the embedding of the input video item, the target visual augmentation to an augmentation identifier.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2023
From: ZHOU, ZHENPENG; POIRSON, PATRICK; GUSAROV, MAKSIM; WANG, CHEN; TOVSTYI, OLEG
To: SNAP INC.
Reel/Frame 063392/0555 →
Continuity (1)
Related Publication 20240355063A1 · Oct 24, 2024
References Cited (32)
US 9239835B1 · Tiwari · 2016 [cited by examiner]
US 10334202B1 · Zhou · 2019 [cited by examiner]
US 11917266B1 · Pundi Ananth · 2024 [cited by examiner]
US 20140267405A1 · Mullins · 2014 [cited by examiner]
US 20140267406A1 · Mullins · 2014 [cited by examiner]
US 20150095804A1 · Grossman · 2015 [cited by examiner]
US 20170228598A1 · Eaton · 2017 [cited by examiner]
US 20180101540A1 · Stoop · 2018 [cited by examiner]
US 20180204068A1 · Eaton · 2018 [cited by examiner]
US 20190220525A1 · Song · 2019 [cited by examiner]
US 20190258722A1 · Guo · 2019 [cited by examiner]
US 20190258963A1 · Guo · 2019 [cited by examiner]
US 20190286943A1 · Leskovec · 2019 [cited by examiner]
US 20190392323A1 · Yan · 2019 [cited by examiner]
US 20200193164A1 · Katti · 2020 [cited by examiner]
US 20200302185A1 · Hussein · 2020 [cited by examiner]
US 20200355925A1 · Hu · 2020 [cited by examiner]
US 20210166066A1 · Ando · 2021 [cited by examiner]
US 20210334994A1 · Park · 2021 [cited by examiner]
US 20210397266A1 · Gupta · 2021 [cited by examiner]
US 20220245140A1 · Gylfason · 2022 [cited by examiner]
US 20230020218A1 · Berger et al. · 2023 [cited by applicant]
US 20230113643A1 · Mittal · 2023 [cited by examiner]
WO WO2024220547A1 · 2024 [cited by applicant]
Chen, Ting, et al., “Big Self-Supervised Models are Strong Semi-Supervised Learners”, arXiv:2006.10029v2 [cs.LG], (Oct. 26, 2020), 18 pgs. [cited by applicant]
Grill, Jean-Bastien, et al., “Bootstrap Your Own Latent a New Approach to Self-Supervised Learning”, arXiv:2006.07733v3 [cs.LG], (Sep. 10, 2020), 35 pgs. [cited by applicant]
Khosla, Prannay, et al., “Supervised Contrastive Learning”, arXiv:2004.11362v5 [cs.LG], (Mar. 10, 2021), 23 pgs. [cited by applicant]
Kondratyuk, Dan, et al., “MoViNets: Mobile Video Networks for Efficient Video Recognition”, arXiv:2103.11511v2 [cs.CV], (Apr. 18, 2021), 21 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/024998, International Search Report mailed Jul. 5, 2024”, 4 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/024998, Written Opinion mailed Jul. 5, 2024”, 8 pgs. [cited by applicant]
Chen, Ting, et al., “A Simple Framework for Contrastive Learning of Visual Representations”, arXiv:2002.05709v3 [cs.LG], (Jul. 1, 2020), 20 pgs. [cited by applicant]
Gowda, Shreyank N., et al., “Learn2Augment: Learning to Composite Videos for Data Augmentation in Action Recognition”, 17th European Conference, Tel Aviv, Israel, Oct. 23-27, 2022, Proceedings, Part XXXI, In: “European … [cited by applicant]