IP Library Granted Patent US 11,151,203
Granted Patent B2
US 11,151,203 · App. 15/908,493 · Granted Oct 19, 2021

Interest embedding vectors

Inventor: Vishnu Priya Natchu (Mountain View, CA)
Assignee: APPLE INC.
G06F16/951H04L67/02G06F40/166H04L67/22H04L67/306
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,151,203
App. No.
15/908,493
Granted
Oct 19, 2021
Kind
B2
Abstract

Techniques for generating interest embedding vectors are disclosed. In some embodiments, a system/process/computer program product for generating interest embedding vectors includes aggregating a plurality of web documents associated with one or more entities, wherein the web documents are retrieved from a plurality of online content sources including one or more websites; selecting a plurality of tokens based on processing of the plurality of web documents; and generating embeddings of the selected tokens in an embedding space.

Claims (57)

1. A system, comprising:

a processor configured to:

aggregate a plurality of web documents associated with one or more entities, wherein the web documents are retrieved from a plurality of online content sources including one or more websites;

determine a score for each web document of the plurality of web documents;

determine a plurality of tokens from the plurality of web documents, each of the plurality of tokens corresponding to a respective entity;

generate a respective co-occurrence weight for each respective pair of tokens of the plurality of tokens using a square root of a product of the scores of the plurality of web documents that include the respective pair of tokens;

select tokens of the plurality of tokens based at least in part on the respective co-occurrence weights; and

generate embeddings of the selected tokens in an embedding space; and

a memory coupled with the processor, wherein the memory is configured to provide the processor with instructions.

2. The system of claim 1 , wherein each of the plurality of tokens can represent an entity, an n-gram, a search query, or a web document.

3. The system of claim 1 , wherein the embedding space includes embeddings for interests and web documents.

4. The system of claim 1 , wherein the embedding space includes embeddings for entities, n-grams, search queries, and web documents.

5. The system of claim 1 , wherein content from the plurality of online content sources includes text-based information, and wherein the processor is further configured to analyze the text-based information to determine a document score associated with each of the one or more entities.

6. The system of claim 1 , wherein the processor is further configured to:

determine a mapping from a web document to an interest;

determine a second mapping from the web document to another web document; and

determine a third mapping from the interest to another web document.

7. The system of claim 1 , wherein the processor is further configured to:

receive a user query; and

return a web document in response to the user query based on a distance between the web document and the user query in the embedding space.

8. The system of claim 1 , wherein the processor is further configured to:

generate a content feed that includes a first web document; and

send a second web document in an update to the content feed based on a distance between the first web document and the second web document in the embedding space.

9. A method, comprising:

aggregating a plurality of web documents associated with one or more entities, wherein the web documents are retrieved from a plurality of online content sources including one or more websites;

determining a plurality of tokens from the plurality of web documents, each of the plurality of tokens corresponding to a respective entity;

determining a score for each web document of the plurality of web documents;

generating a respective co-occurrence weight for each respective pair of tokens of the plurality of tokens using a square root of a product of the scores of the plurality of web documents that include the respective pair of tokens;

selecting tokens of the plurality of tokens based at least in part on the respective co-occurrence weights; and

generating embeddings of the selected tokens in an embedding space.

10. The method of claim 9 , wherein each of the plurality of tokens can represent an entity, an n-gram, a search query, or a web document.

11. The method of claim 9 , wherein the embedding space includes embeddings for interests and web documents.

12. The method of claim 9 , wherein the embedding space includes embeddings for entities, n-grams, search queries, and web documents.

13. The method of claim 9 , further comprising determining a mapping from a web document to an interest.

14. The method of claim 9 , further comprising determining a mapping from a first web document to a second web document.

15. The method of claim 9 , further comprising determining a mapping from an interest to a web document.

16. The method of claim 9 , further comprising:

receiving a user query; and

returning a web document in response to the user query based on a distance between the web document and the user query in the embedding space.

17. The method of claim 9 , further comprising:

generating a content feed that includes a first web document; and

sending a second web document in an update to the content feed based on a distance between the first web document and the second web document in the embedding space.

18. A computer program product, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

aggregating a plurality of web documents associated with one or more entities, wherein the web documents are retrieved from a plurality of online content sources including one or more websites;

determining a score for each web document of the plurality of web documents;

determining a plurality of tokens from the plurality of web documents, each of the plurality of tokens corresponding to a respective entity;

generating a respective co-occurrence weight for each respective pair of tokens of the plurality of tokens using a square root of a product of the scores of the plurality of web documents that include the respective pair of tokens;

selecting tokens of the plurality of tokens based at least in part on the respective co-occurrence weights; and

generating embeddings of the selected tokens in an embedding space.

19. The method of claim 9 , further comprising:

generating a co-occurrence matrix that indicates how many of the plurality of documents include each respective pair of tokens;

determining, based on the co-occurrence matrix, a pairwise mutual information metric that indicates a likelihood of a co-occurrence of each respective pair of tokens of the plurality of tokens, wherein the determining comprises factorizing a loss function that provides an estimate of a likelihood that the co-occurrence is a random co-occurrence; and

selecting the tokens of the plurality of tokens based at least in part on the pairwise mutual information metric determined for each respective pair of tokens of the plurality of tokens.

20. The system of claim 1 , wherein the processor is further configured to:

generate a co-occurrence matrix that indicates how many of the plurality of documents include each respective pair of tokens;

determine, based on the co-occurrence matrix, a pairwise mutual information metric that indicates a likelihood of a co-occurrence of each respective pair of tokens of the plurality of tokens, wherein the determining comprises factorizing a loss function that provides an estimate of a likelihood that the co-occurrence is a random co-occurrence; and

select the tokens of the plurality of tokens based at least in part on the pairwise mutual information metric determined for each respective pair of tokens of the plurality of tokens.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 2, 2021
From: LASERLIKE, INC.
To: APPLE INC.
Reel/Frame 057374/0539 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2018
From: NATCHU, VISHNU PRIYA
To: LASERLIKE INC.
Reel/Frame 045926/0181 →
Continuity (2)
Provisional Application 62465117 · Feb 28, 2017
Related Publication 20180253496A1 · Sep 6, 2018
Cited By (4)
US 12,266,006 US 12,411,982 US 12,572,814 US 12,671,667