IP Library Granted Patent US 10,803,248
Granted Patent B1
US 10,803,248 · App. 15/862,053 · Granted Oct 13, 2020

Consumer insights analysis using word embeddings

Inventors: Jonathan Michael Arfa (New York, NY); Nikhil Girish Nawathe (New York, NY); Bryan Kauder (New York, NY); Shriram Subramanian (New York, NY)
Assignee: Facebook, Inc.
G06F40/30G06F40/295G06N5/022H04L51/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,803,248
App. No.
15/862,053
Granted
Oct 13, 2020
Kind
B1
Abstract

In one embodiment, a method includes receiving a request to generate k keywords each of which is semantically related to a particular subject, where the request includes an input n-gram representing the particular subject, accessing a table of word vector relationships, where the table includes a plurality of unique n-grams and their corresponding word vectors, and wherein each of the word vectors represents a semantic context of a corresponding n-gram as a point in a d-dimensional embedding space, looking up, using the table, a first word vector corresponding to the input n-gram, selecting k word vectors closest to the first word vector in the embedding space using the table and based on a similarity metric, identifying, for each of the selected word vectors, a corresponding n-gram by looking up the selected word vector in the table, and sending a response message including the identified n-grams.

Claims (42)

1. A method comprising:

by a first computing device in an online social network, receiving, from a second computing device, a request to generate k keywords each of which is semantically related to a particular subject in public perceptions of a particular demographic, wherein the request comprises an input n-gram representing the particular subject, and wherein the request comprises one or more conditions characterizing the particular demographic;

by the first computing device, constructing a corpus of text by collecting text content from content objects created by users of the online social network who satisfy the one or more conditions;

by the first computing device, constructing a table of word vector relationships by training a word embedding mode using the constructed corpus of text as training data, wherein the table comprises a plurality of unique n-grams and their corresponding word vectors, and wherein each of the word vectors represents a semantic context of a corresponding n-gram as a point in a d-dimensional embedding space;

by the first computing device, looking up, using the table, a first word vector corresponding to the input n-gram;

by the first computing device, selecting, using the table and based on a similarity metric, k word vectors closest to the first word vector in the embedding space;

by the first computing device, identifying, for each of the selected word vectors, a corresponding n-gram by looking up the selected word vector in the table; and

by the first computing device, sending, to the second computing device, a response message comprising the identified n-grams.

2. The method of claim 1 , wherein the particular subject is a person, an object, or a concept.

3. The method of claim 1 , wherein the plurality of unique n-grams in the table are selected from the corpus of text.

4. The method of claim 1 , wherein the word embedding model is a word2vec model.

5. The method of claim 1 , wherein the looking up a word vector corresponding to the input n-gram comprises looking up the n-gram in the table.

6. The method of claim 1 , wherein the similarity metric is a cosign similarity, a Euclidean distance, or a Jaccard similarity coefficient.

7. The method of claim 1 , wherein the content objects were created within a pre-determined period of time.

8. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

receive, from a second computing device, a request to generate k keywords each of which is semantically related to a particular subject, wherein the request comprises an input n-gram representing the particular subject in public perceptions of a particular demographic, and wherein the request comprises one or more conditions characterizing the particular demographic;

construct a corpus of text by collecting text content from content objects created by users of the online social network who satisfy the one or more conditions;

construct a table of word vector relationships by training a word embedding model using the constructed corpus of text as training data, wherein the table comprises a plurality of unique n-grams and their corresponding word vectors, and wherein each of the word vectors represents a semantic context of a corresponding n-gram as a point in a d-dimensional embedding space;

look up, using the table, a first word vector corresponding to the input n-gram;

select, using the table and based on a similarity metric, k word vectors closest to the first word vector in the embedding space;

identify, for each of the selected word vectors, a corresponding n-gram by looking up the selected word vector in the table; and

send, to the second computing device, a response message comprising the identified n-grams.

9. The media of claim 8 , wherein the particular subject is a person, an object, or a concept.

10. The media of claim 8 , wherein the plurality of unique n-grams in the table are selected from the corpus of text.

11. The media of claim 8 , wherein the word embedding model is a word2vec model.

12. The media of claim 8 , wherein the looking up a word vector corresponding to the input n-gram comprises looking up the n-gram in the table.

13. The media of claim 8 , wherein the similarity metric is a cosign similarity, a Euclidean distance, or a Jaccard similarity coefficient.

14. A system comprising:

one or more processors; and

one or more computer-readable non-transitory storage media coupled to one or more of the processors and comprising instructions operable when executed by one or more of the processors to cause the system to:

receive, from a second computing device, a request to generate k keywords each of which is semantically related to a particular subject in public perceptions of a particular demographic, wherein the request comprises an input n-gram representing the particular subject, and wherein the request comprises one or more conditions characterizing the particular demographic;

construct a corpus of text by collecting text content from content objects created by users of the online social network who satisfy the one or more conditions;

construct a table of word vector relationships by training a word embedding model using the constructed corpus of text as training data, wherein the table comprises a plurality of unique n-grams and their corresponding word vectors, and wherein each of the word vectors represents a semantic context of a corresponding n-gram as a point in a d-dimensional embedding space;

look up, using the table, a first word vector corresponding to the input n-gram;

select, using the table and based on a similarity metric, k word vectors closest to the first word vector in the embedding space;

identify, for each of the selected word vectors, a corresponding n-gram by looking up the selected word vector in the table; and

send, to the second computing device, a response message comprising the identified n-grams.

15. The system of claim 14 , wherein the particular subject is a person, an object, or a concept.

16. The system of claim 14 , wherein the plurality of unique n-grams in the table are selected from the corpus of text.

17. The system of claim 14 , wherein the word embedding model is a word2vec model.

18. The system of claim 14 , wherein the looking up a word vector corresponding to the input n-gram comprises looking up the n-gram in the table.

19. The system of claim 14 , wherein the similarity metric is a cosign similarity, a Euclidean distance, or a Jaccard similarity coefficient.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2018
From: ARFA, JONATHAN MICHAEL; NAWATHE, NIKHIL GIRISH; KAUDER, BRYAN; SUBRAMANIAN, SHRIRAM
To: FACEBOOK, INC.
Reel/Frame 045464/0465 →