IP Library Granted Patent US 10,657,332
Granted Patent B2
US 10,657,332 · App. 15/850,382 · Granted May 19, 2020

Language-agnostic understanding

Inventors: Ying Zhang (Palo Alto, CA); Reshef Shilon (Palo Alto, CA); Jing Zheng (San Jose, CA)
Assignee: FACEBOOK, INC.
G06F40/49G06F16/35G06F40/216G06F40/284G06F40/30G06F40/44G06F40/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,657,332
App. No.
15/850,382
Granted
May 19, 2020
Kind
B2
Abstract

Exemplary embodiments relate to techniques to classify or detect the intent of content written in a language for which a classifier does not exist. These techniques involve building a code-switching corpus via machine translation, generating a universal embedding for words in the code-switching corpus, training a classifier on the universal embeddings to generate an embedding mapping/table; accessing new content written in a language for which a specific classifier may not exist, and mapping entries in the embedding mapping/table to the universal embeddings. Using these techniques, a classifier can be applied to the universal embedding without needing to be trained on a particular language. Exemplary embodiments may be applied to recognize similarities in two content items, make recommendations, find similar documents, perform deduplication, and perform topic tagging for stories in foreign languages.

Claims (29)

1. A method comprising:

accessing a code-switching corpus, wherein the code-switching corpus is built by machine-translating select words from a first language to a second language, the select words comprising words translatable with a confidence above a predetermined threshold;

generating a universal embedding for words in the code-switching corpus, the universal embedding mapping the words in the code-switching corpus to a corresponding language-agnostic semantic meaning of the words;

accessing content to be classified; and

classifying the content based on the universal embedding.

2. The method of claim 1 , wherein the universal embedding is a vector in an embedding space that uniquely a semantic meaning within the embedding space.

3. The method of claim 1 , further comprising applying the universal embedding to recognize similarities in two content items.

4. The method of claim 1 , further comprising applying the universal embedding to perform deduplication.

5. The method of claim 1 , further comprising applying the universal embedding to tag content in a target foreign language with a topic based on a corresponding tag applied to content in a source language.

6. The method of claim 1 , wherein the universal embedding is generating using a loss function, and further comprising refining the universal embedding for an identified task by using the loss function as a regularization term when tuning the universal embedding.

7. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:

access a code-switching corpus, wherein the code-switching corpus is built by machine-translating select words from a first language to a second language, the select words comprising words translatable with a confidence above a predetermined threshold;

generate a universal embedding for words in the code-switching corpus, the universal embedding mapping the words in the code-switching corpus to a corresponding language-agnostic semantic meaning of the words;

access content to be classified; and

classify the content based on the universal embedding.

8. The medium of claim 7 , wherein the universal embedding is a vector in an embedding space that uniquely a semantic meaning within the embedding space.

9. The medium of claim 7 , further storing instructions for applying the universal embedding to recognize similarities in two content items.

10. The medium of claim 7 , further storing instructions for applying the universal embedding to perform deduplication.

11. The medium of claim 7 , further storing instructions for applying the universal embedding to tag content in a target foreign language with a topic based on a corresponding tag applied to content in a source language.

12. The medium of claim 7 , wherein the universal embedding is generating using a loss function, and further storing instructions for refining the universal embedding for an identified task by using the loss function as a regularization term when tuning the universal embedding.

13. An apparatus comprising:

a non-transitory computer-readable medium configured to store a code-switching corpus, wherein the code-switching corpus is built by machine-translating select words from a first language to a second language, the select words comprising words translatable with a confidence above a predetermined threshold;

a hardware processor circuit;

embedding logic executable on the processor circuit to generate a universal embedding for words in the code-switching corpus, the universal embedding mapping the words in the code-switching corpus to a corresponding language-agnostic semantic meaning of the words;

a classifier configured to access content to be classified and classify the content based on the universal embedding.

14. The apparatus of claim 13 , wherein the universal embedding is a vector in an embedding space that uniquely a semantic meaning within the embedding space.

15. The apparatus of claim 13 , further comprising similarity logic for applying the universal embedding to recognize similarities in two content items.

16. The apparatus of claim 13 , further comprising deduplication logic for applying the universal embedding to perform deduplication.

17. The apparatus of claim 13 , further comprising tagging logic for applying the universal embedding to tag content in a target foreign language with a topic based on a corresponding tag applied to content in a source language.

Assignments (2)
CHANGE OF NAME Recorded May 5, 2022
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 059858/0387 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2018
From: ZHANG, YING; SHILON, RESHEF; ZHENG, JING
To: FACEBOOK, INC.
Reel/Frame 044621/0021 →
Continuity (1)
Related Publication 20190197119A1 · Jun 27, 2019
Cited By (2)
US 12,229,514 US 12,254,265