IP Library Granted Patent US 10,699,073
Granted Patent B2
US 10,699,073 · App. 16/210,405 · Granted Jun 30, 2020

Systems and methods for language detection

Inventors: Nikhil Bojja (Mountain View, CA); Pidong Wang (Cupertino, CA); Shiman Guo (Cupertino, CA)
Assignee: MZ IP Holdings, LLC
G06F40/263G06F40/232
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,699,073
App. No.
16/210,405
Granted
Jun 30, 2020
Kind
B2
Abstract

Implementations of the present disclosure are directed to a method, a system, and a computer program storage device for identifying a language in a message. Non-language characters are removed from a text message to generate a sanitized text message. An alphabet and/or a script are detected in the sanitized text message by performing at least one of (i) an alphabet-based language detection test to determine a first set of scores and (ii) a script-based language detection test to determine a second set of scores. Each score in the first set of scores represents a likelihood that the sanitized text message includes the alphabet for one of a plurality of different languages. Each score in the second set of scores represents a likelihood that the sanitized text message includes the script for one of the plurality of different languages. The language in the sanitized text message is identified based on at least one of the first set of scores, the second set of scores, and a combination of the first and second sets of scores.

Claims (49)

1. A method, comprising:

removing non-language characters from a text message to generate a sanitized text message;

performing a plurality of language detection tests on the sanitized text message,

wherein each language detection test determines a respective set of scores, and

wherein each score in the set of scores represents a likelihood that the sanitized text message is in a respective language of a plurality of different languages;

providing one or more combinations of the score sets as input to a plurality of classifiers,

wherein each classifier is trained using outputs from different combinations of the language detection tests;

obtaining as output from at least one of the plurality of classifiers a respective confidence score that the sanitized text message is in one of a plurality of different languages; and

identifying the language of the sanitized text message based on one of the confidence scores.

2. The method of claim 1 , wherein the non-language characters comprise at least one of an emoji, a punctuation mark, an extra space, a carriage return, and a numerical character.

3. The method of claim 1 , wherein each language detection test comprises one of a byte n-gram language detection test, a dictionary-based language detection test, an alphabet-based language detection test, a script-based language detection test, and a user language profile language detection test.

4. The method of claim 1 , wherein the plurality of language detection tests are performed substantially simultaneously.

5. The method of claim 1 , wherein the one or more combinations of the score sets comprise score sets from at least one of a script-based language detection test and an alphabet-based language detection test.

6. The method of claim 1 , wherein the one or more combinations of the score sets comprise score sets from a byte n-gram language detection test and a dictionary-based language detection test.

7. The method of claim 1 , wherein the score sets comprise at least one score from a user language profile language detection test that identifies a language preference from a user based on previous text messages authored by the user.

8. The method of claim 1 , wherein each classifier comprises one of a supervised learning model, a partially supervised learning model, an unsupervised learning model, and an interpolation.

9. The method of claim 1 , wherein identifying the language of the sanitized text message comprises:

selecting the confidence score based on an expected language detection accuracy.

10. The method of claim 1 , wherein identifying the language of the sanitized text message comprises:

selecting the confidence score based on a linguistic domain of the sanitized text message.

11. A system, comprising:

one or more computer processors programmed to perform operations to:

remove non-language characters from a text message to generate a sanitized text message;

perform a plurality of language detection tests on the sanitized text message,

wherein each language detection test determines a respective set of scores, and

wherein each score in the set of scores represents a likelihood that the sanitized text message is in a respective language of a plurality of different languages;

provide one or more combinations of the score sets as input to a plurality of classifiers,

wherein each classifier is trained using outputs from different combinations of the language detection tests;

obtain as output from at least one of the plurality of classifiers a respective confidence score that the sanitized text message is in one of a plurality of different languages; and

identify the language of the sanitized text message based on one of the confidence scores.

12. The system of claim 11 , wherein the non-language characters comprise at least one of an emoji, a punctuation mark, an extra space, a carriage return, and a numerical character.

13. The system of claim 11 , wherein each language detection test comprises one of a byte n-gram language detection test, a dictionary-based language detection test, an alphabet-based language detection test, a script-based language detection test, and a user language profile language detection test.

14. The system of claim 11 , wherein the one or more combinations of the score sets comprise score sets from at least one of a script-based language detection test and an alphabet-based language detection test.

15. The system of claim 11 , wherein the one or more combinations of the score sets comprise score sets from a byte n-gram language detection test and a dictionary-based language detection test.

16. The system of claim 11 , wherein the score sets comprise at least one score from a user language profile language detection test that identifies a language preference from a user based on previous text messages authored by the user.

17. The system of claim 11 , wherein each classifier comprises one of a supervised learning model, a partially supervised learning model, an unsupervised learning model, and an interpolation.

18. The system of claim 11 , wherein to identify the language of the sanitized text message the one or more computer processors are further to:

select the confidence score based on an expected language detection accuracy.

19. The system of claim 11 , wherein to identify the language of the sanitized text message the one or more computer processors are further to:

select the confidence score based on a linguistic domain of the sanitized text message.

20. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more computer processors, cause the one or more computer processors to:

remove non-language characters from a text message to generate a sanitized text message;

perform a plurality of language detection tests on the sanitized text message,

wherein each language detection test determines a respective set of scores, and

wherein each score in the set of scores represents a likelihood that the sanitized text message is in a respective language of a plurality of different languages;

provide one or more combinations of the score sets as input to a plurality of classifiers,

wherein each classifier is trained using outputs from different combinations of the language detection tests;

obtain as output from at least one of the plurality of classifiers a respective confidence score that the sanitized text message is in one of a plurality of different languages; and

identify the language of the sanitized text message based on one of the confidence scores.

Assignments (3)
NOTICE OF SECURITY INTEREST -- PATENTS Recorded Mar 19, 2019
From: MZ IP HOLDINGS, LLC
To: MGG INVESTMENT GROUP LP, AS COLLATERAL AGENT
Reel/Frame 048639/0751 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2019
From: BOJJA, NIKHIL; WANG, PIDONG; GUO, SHIMAN
To: MACHINE ZONE, INC.
Reel/Frame 048611/0812 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2019
From: MACHINE ZONE, INC.
To: MZ IP HOLDINGS, LLC
Reel/Frame 048611/0887 →
Continuity (4)
Continuation 15283646 · Oct 3, 2016
Continuation In Part 15161913 · May 23, 2016
Continuation 14517183 · Oct 17, 2014
Related Publication 20190108214A1 · Apr 11, 2019