IP Library › Granted Patent US 12,681,956
Granted Patent B2
US 12,681,956 · App. 18/626,268 · Granted Jul 14, 2026

Input data item classification using memory data item embeddings

Inventors: Ahmet Iscen (Seyssinet-Pariset, FR); Alireza Fathi (Redwood City, CA); Cordelia Luise Schmid (Saint Ismier, FR)
Assignee: Google LLC
G06F16/285G06F16/2438
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,681,956
App. No.
18/626,268
Filed
Apr 3, 2024
Granted
Jul 14, 2026
Kind
B2
Art Unit
2156
USPC
707/737
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for performing a classification task on a data item. In particular, a system classifies an input data item using key and value embeddings of memory data items.

Claims (67)

1 . A method performed by one or more computers, the method comprising:

maintaining, for each of a plurality of memory data items, (i) a respective key embedding that has been generated by processing the memory data item using a first embedding neural network and (ii) a respective value embedding that has been generated by processing at least one of the memory data item or a corresponding data item associated with the memory data item using a second embedding neural network;

receiving a query data item; and

classifying the query data item using a semi-parametric model that uses the respective key and value embeddings for the plurality of memory data items, comprising:

processing the query data item using the first embedding neural network to generate a query embedding of the query data item;

searching through the respective key embeddings of the memory data items by using the query embedding as a query to identify a subset of memory data item;

generating, from the query embedding and the respective value embeddings for the identified subset of memory data items that were identified from searching through the respective key embeddings of the memory data items by using the query embedding as the query, an input to a classifier neural network, wherein the input comprises a set of one or more embeddings generated by fusing the query embedding and the respective value embeddings for the identified subset of memory data items, wherein the fusing comprises (i) generating the input based on the query embedding and a combined value embedding generated from the respective value embeddings, or (ii) processing the query embedding and the respective value embeddings through a sequence of one or more attention layers; and

processing the input using the classifier neural network to generate a classification output for the query data item that specifies one or more categories to which the query data item belongs.

2 . The method of claim 1 , wherein searching through the respective key embeddings of the memory data items by using the query embedding as a query to identify a subset of memory data items comprises:

performing a search of the key embeddings of the memory data items to identify, as the identified subset of memory data items, the memory data items that have the k most similar key embeddings to the query embedding according to a similarity measure, wherein k is a fixed integer that is less than a total number of memory data items.

3 . The method of claim 2 , wherein performing a search of the key embeddings comprises performing an approximate k-nearest neighbors search through the key embeddings.

4 . The method of claim 1 , wherein generating the input based on the query embedding and a combined value embedding generated from the respective value embeddings comprises:

summing the combined value embedding and the query embedding to generate the input to the classifier neural network.

5 . The method of claim 1 , wherein generating a combined value embedding from the respective value embeddings for the identified subset of memory data items comprises:

generating an initial combined value embedding from the respective value embeddings for the identified subset of memory data items.

6 . The method of claim 5 , wherein generating a combined value embedding from the respective value embeddings for the identified subset of memory data items comprises:

applying a dense neural network layer to the initial combined value embedding to generate the combined value embedding.

7 . The method of claim 6 , wherein the query embedding has a first dimensionality, each value embedding has a second, different dimensionality and wherein the dense neural network layer maps the initial combined value embedding from the second dimensionality to the first dimensionality.

8 . The method of claim 5 , wherein generating an initial combined value embedding from the respective value embeddings for the identified subset of memory data items comprises:

computing a mean of the respective value embeddings for the identified subset of memory data items.

9 . The method of claim 1 , wherein each attention layer in the sequence of one or more attention layers is configured to:

receive an input query embedding,

compute a respective attention weight for each identified subset of memory data item by, for each identified subset of memory data item, computing an attention weight between the input query embedding and the key embedding for the identified subset of memory data item,

compute an aggregated value embedding by computing a weighted sum of the value embeddings for the identified subset of memory data items in accordance with the respective attention weights, and

use the aggregated value embedding to update the input query embedding, and

the input query embedding for the first attention layer in the sequence is the query embedding.

10 . The method of claim 9 , wherein the updated query embedding generated by the last attention layer in the sequence is the input to the classifier neural network.

11 . The method of claim 9 , wherein the sequence comprises a plurality of attention layers, and wherein the input query embedding for each attention layer after the first attention layer in the sequence is the updated query embedding generated by a preceding attention layer in the sequence.

12 . The method of claim 9 , wherein using the aggregated value embedding to update the input query embedding comprises:

generating an initial updated query embedding from the aggregated value embedding; and

combining the initial updated query embedding and the query embedding to generate the updated query embedding.

13 . The method of claim 12 , wherein combining the initial updated query embedding and the query embedding to generate the updated query embedding comprises:

summing the initial updated query embedding and the query embedding to generate the updated query embedding.

14 . The method of claim 12 , wherein generating an initial updated query embedding from the aggregated value embedding comprises:

applying a dense neural network layer to the aggregated value embedding to generate the initial updated query embedding.

15 . The method of claim 14 , wherein the query embedding has a first dimensionality, each value embedding has a second, different dimensionality and wherein the dense neural network layer maps the aggregated value embedding from the second dimensionality to the first dimensionality.

16 . The method of claim 1 , wherein the query data item is an image.

17 . The method of claim 1 wherein the classifier neural network has been trained on labeled training data for a classification task.

18 . The method of claim 17 wherein the classifier neural network has been trained jointly with one or more attention layers on the labeled training data for the classification task.

19 . The method of claim 17 , wherein the first embedding neural network has been pre-trained prior to the training of the classifier neural network on the labeled training data for the classification task.

20 . The method of claim 19 , wherein the pre-trained first embedding neural network is held frozen during the training of the classifier neural network on the labeled training data for the classification task.

21 . The method of claim 19 , wherein the pre-trained first embedding neural network is fine-tuned during the training of the classifier neural network on the labeled training data for the classification task.

22 . The method of claim 17 , wherein the memory and key embeddings were generated prior to the training of the classifier neural network on the labeled training data for the classification task and are held frozen during the training of the classifier neural network on the labeled training data for the classification task.

23 . The method of claim 1 , wherein the key embeddings have been generated by processing the memory data items using the first embedding neural network after the first embedding neural network has been pre-trained and prior to training the classifier neural network.

24 . The method of claim 1 , wherein the second embedding neural network is different from the first embedding neural network.

25 . The method of claim 24 , wherein the second embedding neural network has (i) more parameters than the first embedding neural network, (ii) generates embeddings that have a higher dimensionality than the embeddings generated by the first embedding neural network, or (iii) both.

26 . The method of claim 1 , wherein the memory data item is of a first type and the corresponding data item is of a different, second type, and, for each memory data item, the value embedding has been generated by processing the corresponding data item using the second embedding neural network.

27 . The method of claim 26 , wherein the corresponding data item for each memory data item is text describing the memory data item.

28 . A system comprising:

one or more computers; and

one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

maintaining, for each of a plurality of memory data items, (i) a respective key embedding that has been generated by processing the memory data item using a first embedding neural network and (ii) a respective value embedding that has been generated by processing at least one of the memory data item or a corresponding data item associated with the memory data item using a second embedding neural network;

receiving a query data item; and

classifying the query data item using a semi-parametric model that uses the respective key and value embeddings for the plurality of memory data items, comprising:

processing the query data item using the first embedding neural network to generate a query embedding of the query data item;

searching through the respective key embeddings of the memory data items by using the query embedding as a query to identify a subset of memory data item;

generating, from the query embedding and the respective value embeddings for the identified subset of memory data items that were identified from searching through the respective key embeddings of the memory data items by using the query embedding as the query, an input to a classifier neural network, wherein the input comprises a set of one or more embeddings generated by fusing the query embedding and the respective value embeddings for the identified subset of memory data items, wherein the fusing comprises (i) generating the input based on the query embedding and a combined value embedding generated from the respective value embeddings, or (ii) processing the query embedding and the respective value embeddings through a sequence of one or more attention layers; and

processing the input using the classifier neural network to generate a classification output for the query data item that specifies one or more categories to which the query data item belongs.

29 . One or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

maintaining, for each of a plurality of memory data items, (i) a respective key embedding that has been generated by processing the memory data item using a first embedding neural network and (ii) a respective value embedding that has been generated by processing at least one of the memory data item or a corresponding data item associated with the memory data item using a second embedding neural network;

receiving a query data item; and

classifying the query data item using a semi-parametric model that uses the respective key and value embeddings for the plurality of memory data items, comprising:

processing the query data item using the first embedding neural network to generate a query embedding of the query data item;

searching through the respective key embeddings of the memory data items by using the query embedding as a query to identify a subset of memory data item;

generating, from the query embedding and the respective value embeddings for the identified subset of memory data items that were identified from searching through the respective key embeddings of the memory data items by using the query embedding as the query, an input to a classifier neural network, wherein the input comprises a set of one or more embeddings generated by fusing the query embedding and the respective value embeddings for the identified subset of memory data items, wherein the fusing comprises (i) generating the input based on the query embedding and a combined value embedding generated from the respective value embeddings, or (ii) processing the query embedding and the respective value embeddings through a sequence of one or more attention layers; and

processing the input using the classifier neural network to generate a classification output for the query data item that specifies one or more categories to which the query data item belongs.

30 . The method of claim 1 , wherein the respective key embeddings and the respective value embeddings of the plurality of memory data items are stored in a database, wherein the database is maintained separately from the classifier neural network, and wherein searching through the respective key embeddings of the memory data comprises searching the database.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2024
From: ISCEN, AHMET; FATHI, ALIREZA; SCHMID, CORDELIA LUISE
To: GOOGLE LLC
Reel/Frame 067649/0399 →
Continuity (2)
Provisional Application 63494215 · Apr 4, 2023
Related Publication 20240338387A1 · Oct 10, 2024
References Cited (67)
US 11947912B1 · Dong · 2024 [cited by examiner]
US 20200285940A1 · Sprechmann · 2020 [cited by examiner]
US 20210383226A1 · Doersch · 2021 [cited by examiner]
US 20220121702A1 · Kale · 2022 [cited by examiner]
US 20220277015A1 · Zhang · 2022 [cited by examiner]
US 20220414169A1 · Biswas · 2022 [cited by examiner]
US 20230113643A1 · Mittal · 2023 [cited by examiner]
US 20230252032A1 · Na · 2023 [cited by examiner]
US 20240202464A1 · Poirier · 2024 [cited by examiner]
US 20240232580A1 · Jaegle · 2024 [cited by examiner]
US 20240403339A1 · Azizi · 2024 [cited by examiner]
Alayrac et al., “Flamingo: a visual language model for few-shot learning,” CoRR, Submitted on Nov. 15, 2022, arXiv:2204.14198v2, pp. 1-54. [cited by applicant]
Basu et al., “Generalization properties of Retrieval-based models,” CoRR, Submitted on Oct. 6, 2022, arXiv:2210.02617v1, pp. 1-38. [cited by applicant]
Blattmann et al., “Semi-parametric neural image synthesis,” CoRR, Submitted on Oct. 24, 2022, arXiv:2204.11824v3, pp. 1-34. [cited by applicant]
Borgeaud et al., “Improving language models by retrieving from trillions of tokens,” Paper, Presented at Proceedings of the International Conference on Machine Learning, Baltimore, MD, Jul. 17-23, 2022; PMLR, 2022, 162:… [cited by applicant]
Brown et al., “Language models are few-shot learners,” Paper, Presented at the 34th Conference on Neural Information Processing Systems, Virtual Event, Dec. 6-12, 2020; Advances in Neural Information Processing Systems,… [cited by applicant]
Chen et al., “PaLI: A Jointly-Scaled Multilingual Language-Image Model,” CoRR, Submitted Sep. 16, 2022, arXiv:2209.06794v2, pp. 1-30. [cited by applicant]
Chen et al., “Re-Imagen: Retrieval-Augmented Text-to-Image Generator,” CoRR, Submitted on Nov. 22, 2022, arXiv:2209.14491v3, pp. 1-25. [cited by applicant]
Chowdhery et al., “PaLM: Scaling Language Modeling with Pathways,” CoRR, Submitted on Oct. 5, 2022, arXiv:2204.02311v5, pp. 1-87. [cited by applicant]
Collier et al., “Correlated input-dependent label noise in large-scale image classification,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, Nashville, TN, Jun. 20-2… [cited by applicant]
Cui et al., “Parametric contrastive learning,” Paper, Presented at Proceedings of the International Conference on Computer Vision, Montreal, Canada, Oct. 10-17, 2021, pp. 715-724. [cited by applicant]
Deng et al., “ImageNet: A Large-Scale Hierarchical Image Database,” Paper, Presented at Proceedings of the IEEE Conference: Computer Vision and Pattern Recognition, Miami, FL, Jun. 20-25, 2009, 8 pages. [cited by applicant]
Deng et al., “PML: Progressive Margin Loss for Long-tailed Age Classification,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, Nashville, TN, Jun. 20-25, 2021, pp. 1… [cited by applicant]
Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” Paper, Presented at Proceedings of the International Conference on Learning Representations, Virtual Event, May 3-7, 2021… [cited by applicant]
Graves et al., “Neural turing machines,” CoRR, Submitted on Dec. 10, 2014, arXiv:1410.5401v2, pp. 1-26. [cited by applicant]
Guo et al., “Accelerating Large-Scale Inference with Anisotropic Vector Quantization,” Paper, Presented at Proceedings of the International Conference on Machine Learning, Virtual Event, Jul. 13-18, 2020; PMLR, Mar. 202… [cited by applicant]
Guo et al., “CurriculumNet: Weakly supervised learning from large-scale web images,” Paper, Presented at Proceedings of the European Conference on Computer Vision, Germany, Munich, Sep. 8-14, 2018; LNCS, 2018, 11214:139… [cited by applicant]
Guo et al., “Long-tailed multi-label visual recognition by collaborative training on uniform and rebalanced samplings,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition… [cited by applicant]
Guu et al., “REALM: Retrieval-augmented language model pre-training,” CoRR, Submitted on Feb. 10, 2020, arXiv:2002.08909v1, 12 pages. [cited by applicant]
He et al., “Learning from imbalanced data,” Paper, IEEE Transactions on Knowledge and Data Engineering, Sep. 2009, 21(9):1263-1284. [cited by applicant]
Hong et al., “Disentangling label distribution for long-tailed visual recognition,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, Nashville, TN, Jun. 20-25, 2021, p… [cited by applicant]
Horn et al., “Benchmarking representation learning for natural world image collections,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, Nashville, TN, Jun. 20-25, 20… [cited by applicant]
Huang et al., “Learning deep representation for imbalanced classification,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, Las Vegas, NV, Jun. 27-30, 2016, pp. 5375-… [cited by applicant]
Iscen et al., “A memory transformer network for incremental learning,” CoRR, Submitted on Oct. 10, 2022, arXiv:2210.04485v1, pp. 1-12. [cited by applicant]
Iscen et al., “Learning with neighbor consistency for noisy labels,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, New Orleans, LA, Jun. 18-24, 2022, pp. 4672-4681. [cited by applicant]
Khandelwal et al., “Generalization through memorization: Nearest neighbor language models,” Paper, Presented at Proceedings of the International Conference on Learning Representations, Virtual Event, Apr. 26-May 1, 2020… [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization,” CoRR, Submitted on Jul. 20, 2015, arXiv:1412.6980v7, pp. 1-13. [cited by applicant]
Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Paper, Presented at the 34th Conference on Neural Information Processing Systems, Virtual Event, Dec. 6-12, 2020; Advances in Neural Info… [cited by applicant]
Li et al., “Web Vision Database: Visual Learning and Understanding from Web Data,” CoRR, Submitted on Aug. 9, 2017, arXiv:1708.02862v1, 9 pages. [cited by applicant]
Liu et al., “Large-scale long-tailed recognition in an open world,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, Long Beach, CA, Jun. 15-20, 2019, pp. 2537-2546. [cited by applicant]
Long et al., “Retrieval augmented classification for long-tail visual recognition,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, New Orleans, LA, Jun. 18-24, 2022,… [cited by applicant]
Loshchilov et al., “SGDR: Stochastic Gradient Descent with Warm Restarts,” Paper, Presented at Proceedings of the International Conference on Learning Representations, Toulon, France, Apr. 24-26, 2017, pp. 1-16. [cited by applicant]
Menon et al., “Long-tail learning via logit adjustment,” Paper, Presented at Proceedings of the International Conference on Learning Representations, Virtual Event, May 3-7, 2021, pp. 1-24. [cited by applicant]
Nakata et al., “Revisiting a kNN-Based Image Classification System with High-Capacity Storage,” Paper, Presented at Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, Oct. 23-27, 2022; LNCS, 20… [cited by applicant]
Park et al., “The Long Tail of Recommender Systems and How to Leverage It,” Paper, Presented at Proceedings of the RecSys Conference, Lausanne, Switzerland, Oct. 23-25, 2008, pp. 11-18. [cited by applicant]
Radford et al., “Learning transferable visual models from natural language supervision,” Paper, Presented at the International Conference on Machine Learning, Virtual Conference, Jul. 18-24, 2021; PMLR, Apr. 2021, 139:1… [cited by applicant]
Raffel et al., “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research, Jun. 2020, 21(140):1-67. [cited by applicant]
Rajeswar et al., “Multi-label iterated learning for image classification with label ambiguity,” CoRR, Submitted on Nov. 23, 2021, arXiv:2111.12172v1, pp. 1-14. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” Int J Comput Vis, Apr. 2015, 115(3):211-252. [cited by applicant]
Santoro et al., “Meta-learning with memory augmented neural networks,” Paper, Presented at Proceedings of the International Conference on Machine Learning, New York, NY, Jun. 19-24, 2016; PMLR, 2016, 48:9 pages. [cited by applicant]
Schuhmann et al., “LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs,” CoRR, Submitted on Nov. 3, 2021, arXiv:2111.02114v1, 5 pages. [cited by applicant]
Singh et al., “FLAVA: A foundational language and vision alignment model,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, New Orleans, LA, Jun. 18-24, 2022, pp. 1563… [cited by applicant]
Szegedy et al., “Rethinking the Inception Architecture for Computer Vision,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, Las Vegas, NV, Jun. 27-30, 2016, pp. 2818… [cited by applicant]
Thomee et al., “YFCC100M: The New Datain Multimedia Research,” Communications of the ACM, Feb. 2016, 59(2):64-73. [cited by applicant]
Tian et al., “VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition,” Paper, Presented at Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, Oct. 23-27… [cited by applicant]
Wang et al., “Image as a Foreign Language: BEIT Pretraining for All Vision and Vision-Language Tasks,” CoRR, Submitted on Aug. 31, 2022, arXiv:2208.10442v2, 18 pages. [cited by applicant]
Wang et al., “Long-tailed recognition by routing diverse distribution-aware experts,” Paper, Presented at Proceedings of the International Conference on Learning Representations, Virtual Event, May 3-7, 2021, pp. 1-15. [cited by applicant]
Wang et al., “Training data is more valuable than you think: A simple and effective method by retrieving from training data,” CoRR, Submitted on Mar. 16, 2022, arXiv:2203.08773v1, 11 pages. [cited by applicant]
Wu et al., “Memorizing transformers,” Paper, Presented at Proceedings of the International Conference on Learning Representations, Virtual Event, Apr. 25-29, 2022, pp. 1-19. [cited by applicant]
Xiang et al., “Learning from multiple experts: Self-paced knowledge distillation for long-tailed classification,” Paper, Presented at proceedings of the European Conference on Computer Vision, Glasgow, UK, Aug. 23-28, 2… [cited by applicant]
Yu et al., “CoCa: Contrastive Captioners are Image-Text Foundation Models,” CORR, Submitted on Jun. 14, 2022, arXiv:2205.01917v2, pp. 1-19. [cited by applicant]
Yuan et al., “Florence: A New Foundation Model for Computer Vision,” CoRR, Submitted on Nov. 22, 2021, arXiv:2111.11432v1, 17 pages. [cited by applicant]
Zhai et al., “LiT: Zero-Shot Transfer with Locked-image text Tuning, ” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, New Orleans, LA, Jun. 18-24, 2022, pp. 18123-18… [cited by applicant]
Zhai et al., “Scaling vision transformers,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, New Orleans, LA, Jun. 18-24, 2022, pp. 12104-12113. [cited by applicant]
Zhao et al., “Improving long-tailed classification from instance level,” CoRR, Submitted on Apr. 13, 2021, arXiv:2104.06094v1, 10 pages. [cited by applicant]
Zhou et al., “BBN: Bilateral-branch network with cumulative learning for long-tailed visual recognition,” Paper, Presented at Proceedings of the IEEE/CVF Conference: Computer Vision and Pattern Recognition, Seattle, WA,… [cited by applicant]
Zhou et al., “Places: A 10 Million Image Database for Scene Recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Jun. 2018, 40(6):1452-1464. [cited by applicant]