IP Library › Granted Patent US 12,400,078
Granted Patent B1
US 12,400,078 · App. 17/706,303 · Granted Aug 26, 2025

Interpretable embeddings

Inventors: Nisarg Vyas (Gujarat, IN); David Andre (San Francisco, CA)
Assignee: GOOGLE LLC
G06F40/284G06F18/2137G06F40/30G06N3/063G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,078
App. No.
17/706,303
Granted
Aug 26, 2025
Kind
B1
Abstract

This specification is generally directed to techniques for creating reduced-dimensionality embeddings (e.g., embedding layers of neural networks) with dimensions that are interpretable by and/or are meaningful to humans. In various implementations, a datum may sampled from a document. A dimensionality reduction process may be performed based on the sampled datum to generate a semantically-interpretable embedding having a number of individually-interpretable dimensions. The dimensionality reduction process may include: analyzing the sampled datum according to a number of distinct semantic queries to determine respective numeric solutions. The number of distinct semantic queries may correspond to the number of individually-interpretable dimensions. Each numeric solution may offer an inconclusive clue about the sampled datum. The dimensionality reduction process may also include populating the dimensions of the semantically-interpretable embedding with respective numeric solutions.

Claims (37)

1. A method for generating an explanation of a visual hierarchy of a digital image, the method implemented using one or more processors and comprising:

processing the digital image using at least part of a machine learning pipeline to yield a first latent embedding layer representation of the digital image;

using the first latent embedding layer, generating a first semantically-interpretable embedding having a number of individually-interpretable dimensions, including

populating the dimensions of the first semantically-interpretable embedding with respective probabilities of visual features of a first level of granularity being present in the digital image;

processing the first latent embedding layer using at least part of the machine learning pipeline to yield a second latent embedding layer representation of the digital image;

using the second latent embedding layer, generating a second semantically-interpretable embedding having a number of individually-interpretable dimensions, including populating the dimensions of the second semantically-interpretable embedding with respective probabilities of visual features of a second level of granularity being present in the digital image;

processing the second latent embedding layer using at least part of the machine learning pipeline to yield a third latent embedding layer representation of the digital image; and

using the third latent embedding layer, generating a third semantically-interpretable embedding having a number of individually-interpretable dimensions, including populating the dimensions of the third semantically-interpretable embedding with respective probabilities of visual features of a third level of granularity being present in the digital image.

2. The method of claim 1 , wherein the second level of granularity is greater than the first level of granularity.

3. The method of claim 2 , wherein the third level of granularity is greater than the second level of granularity.

4. The method of claim 3 , wherein the visual features of the first level of granularity include lines.

5. The method of claim 4 , wherein the visual features of the second or third level of granularity include one or more shapes selected from rectangles, squares, circles, and rhombuses.

6. The method of claim 4 , wherein the visual features of the second or third level of granularity include colors.

7. A system for generating an explanation of a visual hierarchy of a digital image, the system comprising one or more processors and memory storing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to:

process the digital image using at least part of a machine learning pipeline to yield a first latent embedding layer representation of the digital image;

perform a dimensionality reduction process using the digital image to generate the first latent embedding layer, generating a first semantically-interpretable embedding having a number of individually-interpretable dimensions, including populating the dimensions of the first semantically-interpretable embedding with respective probabilities of visual features of a first level of granularity being present in the digital image;

process the first latent embedding layer using at least part of the machine learning pipeline to yield a second latent embedding layer representation of the digital image;

using the second latent embedding layer, generate a second semantically-interpretable embedding having a number of individually-interpretable dimensions, including populating the dimensions of the second semantically-interpretable embedding with respective probabilities of visual features of a second level of granularity being present in the digital image;

process the second latent embedding layer using at least part of the machine learning pipeline to yield a third latent embedding layer representation of the digital image; and

using the third latent embedding layer, generate a third semantically-interpretable embedding having a number of individually-interpretable dimensions, including populating the dimensions of the third semantically-interpretable embedding with respective probabilities of visual features of a third level of granularity being present in the digital image.

8. The system of claim 7 , wherein the second level of granularity is greater than the first level of granularity.

9. The system of claim 8 , wherein the third level of granularity is greater than the second level of granularity.

10. The system of claim 9 , wherein the visual features of the first level of granularity include lines.

11. The system of claim 10 , wherein the visual features of the second or third level of granularity include one or more shapes selected from rectangles, squares, circles, and rhombuses.

12. The system of claim 10 , wherein the visual features of the second or third level of granularity include colors.

13. At least one non-transitory computer-readable medium for generating an explanation of a visual hierarchy of a digital image, the medium comprising instructions that, in response to execution of the instructions by one or more processors, cause the one or more processors to:

process the digital image using at least part of a machine learning pipeline to yield a first latent embedding layer representation of the digital image;

perform a dimensionality reduction process using the digital image to generate the first latent embedding layer, generating a first semantically-interpretable embedding having a number of individually-interpretable dimensions, including populating the dimensions of the first semantically-interpretable embedding with respective probabilities of visual features of a first level of granularity being present in the digital image;

process the first latent embedding layer using at least part of the machine learning pipeline to yield a second latent embedding layer representation of the digital image;

using the second latent embedding layer, generate a second semantically-interpretable embedding having a number of individually-interpretable dimensions, including populating the dimensions of the second semantically-interpretable embedding with respective probabilities of visual features of a second level of granularity being present in the digital image;

process the second latent embedding layer using at least part of the machine learning pipeline to yield a third latent embedding layer representation of the digital image; and

using the third latent embedding layer, generate a third semantically-interpretable embedding having a number of individually-interpretable dimensions, including populating the dimensions of the third semantically-interpretable embedding with respective probabilities of visual features of a third level of granularity being present in the digital image.

14. The at least one non-transitory computer-readable medium of claim 13 , wherein the second level of granularity is greater than the first level of granularity.

15. The at least one non-transitory computer-readable medium of claim 14 , wherein the third level of granularity is greater than the second level of granularity.

16. The at least one non-transitory computer-readable medium of claim 15 , wherein the visual features of the first level of granularity include lines.

17. The at least one non-transitory computer-readable medium of claim 16 , wherein the visual features of the second or third level of granularity include one or more shapes selected from rectangles, squares, circles, and rhombuses.

18. The at least one non-transitory computer-readable medium of claim 16 , wherein the visual features of the second or third level of granularity include colors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2022
From: VYAS, NISARG; ANDRE, DAVID
To: X DEVELOPMENT LLC
Reel/Frame 059604/0041 →
Continuity (1)
Provisional Application 63210717 · Jun 15, 2021
References Cited (22)
US 20140376804A1 · Akata · 2014 [cited by examiner]
US 20190266487A1 · Chollet · 2019 [cited by examiner]
US 20200175360A1 · Conti · 2020 [cited by examiner]
US 20230410471A1 · Highnam · 2023 [cited by examiner]
Deng, C., Lai, G. and Deng, H. (2020), Improving word vector model with part-of-speech and dependency grammar information. Nov. 2, 2020, CAAI Trans. Intell. Technol., 5: 276-282. https://doi.org/10.1049/trit.2020.0055 (… [cited by examiner]
Hwa Jong Kim, Seong Eun Hong, Kyung Jin Cha, seq2vec: Analyzing sequential data using multi-rank embedding vectors, Electronic Commerce Research and Applications, vol. 43, 2020, 101003, ISSN 1567-4223, https://doi.org/1… [cited by examiner]
Zeynep Akata, Florent Perronnin, Zaid Harchaoui, Cordelia Schmid, “Label-Embedding for Image Classification”, Oct. 1, 2015, arXiv:150308677 [cs.CV] (Year: 2015). [cited by examiner]
Lutfi Kerem Senel, Ihsan Utlu, Veysel Yucesoy, Aykut Koc, Tolga Cukur, “Semantic Structure and Interpretability of Word Embeddings”, May 16, 2018, arXiv:1711.00331 [cs.CL] (Year: 2018). [cited by examiner]
Viphavee Vongpumivitch, Ju-yu Huang, Yu-Chia Chang, “Frequency analysis of the words in the Academic Word List (AWL) and non-AWL content words in applied linguistics research papers”, English for Specific Purposes, vol.… [cited by examiner]
Altmann, E.G., Whichard, Z.L. & Motter, A.E. Identifying Trends in Word Frequency Dynamics. J Stat Phys 151, 277-288 (2013). https://doi.org/10.1007/s10955-013-0699-7 (Year: 2013). [cited by examiner]
Gerritsen, M., Jansen, F. (1980). Word Frequency and Lexical Diffusion in Dialect Borrowing and Phonological Change. In: Geerts, G., et al. Dutch Studies. Springer, Dordrecht. https://doi.org/10.1007/978-94-009-8855-2_3… [cited by examiner]
Liu, Pengfei et al. “Learning Context-Sensitive Word Embeddings with Neural Tensor Skip-Gram Model.” International Joint Conference on Artificial Intelligence (2015). (Year: 2015). [cited by examiner]
Senel et al., “Semantic Structure and Interpretability of Word Embeddings” arXiv:1711.00331v3 [cs.CL] dated May 16, 2018. 11 pages. [cited by applicant]
Subramanian et al., “SPINE: SParse Interpretable Neural Embeddings” 32nd AAAI Conference of Artificial Intelligence (AAAI-18). 8 pages. [cited by applicant]
Gupta et al., “SEMIE: SEMantically Infused Embeddings with Enhanced Interpretability for Domain-specific Small Corpus” arXiv:2103.11431v1 [cs.CL] dated May 21, 2021. 9 pages. [cited by applicant]
Panigrahi et al., “Word2Sense: Sparse Interpretable Word Embeddings” Microsoft Research India. 14 pages. [cited by applicant]
Templeton “Inherently Interpretable Sparse Word Embeddings through Sparse Coding” arXiv:2004.13847v1 [cs.CL] dated Apr. 8, 2020. 18 pages. [cited by applicant]
Zhang et al., “Word Embedding Visualization Via Dictionary Learning” arXiv:1910.03833v2 [cs.CL] dated Mar. 15, 2021. 14 pages. [cited by applicant]
Dufter et al., “Analytical Methods for Interpretable Ultradense Word Embeddings” arXiv:1904.08654v2 [cs.CL] dated Sep. 13, 2019. 10 pages. [cited by applicant]
Senel et al., “Imparting Interpretability to Word Embeddings while Preserving Semantic Structure” arXiv:1807.07279v4 [cs.CL] dated Jul. 2, 2020. 15 pages. [cited by applicant]
Qureshi et al., “EVE: Explainable Vector Based Embedding Technique Using Wikipedia” arXiv:1702.06891v1 [cs.CL] dated Feb. 22, 2017. 22 pages. [cited by applicant]
Jiang et al., “Language as an Abstraction for Hierarchical Deep Reinforcement Learning” arXiv:1906.07343v2 [cs.LG] dated Nov. 18, 2019. 25 pages. [cited by applicant]
Cited By (1)
US 12,749,494