IP Library Granted Patent US 8,073,680
Granted Patent B2
US 8,073,680 · App. 12/147,340 · Granted Dec 6, 2011

Language detection service

Assignee: Microsoft Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,073,680
App. No.
12/147,340
Granted
Dec 6, 2011
Kind
B2
Abstract

Language detection techniques are described. In implementation, a method comprises determining which human writing system is associated with text characters in a string based on values representing the text characters. When the values are associated with more than one human language, the string is compared with a targeted dictionary to identify a corresponding human language associated with the string.

Claims (27)

1. A method comprising:

determining, by one or more computing devices, which human writing system is associated with text characters in a string of one or more text characters based on values representing the text characters; and

when said values are associated with more than one human language, comparing, by one or more computing devices, the string with a targeted dictionary to identify a particular said human language associated with the string.

2. The method of claim 1 , further comprising ascertaining which said human language is associated with the string based on which said human language is associated with a substring in the string.

3. The method of claim 2 , wherein the substring is one or more of a prefix or a suffix.

4. The method of claim 1 , further comprising assigning a text including the string the particular said human language.

5. The method of claim 1 , wherein the targeted dictionary includes strings individually associated with a single said human language.

6. The method of claim 1 , wherein individual strings in the targeted dictionary are associated with a single human language and comprise one or more of a determiner, a preposition, a connector or a pronoun.

7. The method of claim 1 , further comprising selecting which of a plurality of strings in the targeted dictionary are to be compared based on said values representing the text characters in the string.

8. The method of claim 1 , wherein in said values are Unicode said values.

9. A system comprising:

at least a memory and a processor to implement a language detection service, the language detection service configured to:

identify which human writing system is associated with a string of text characters in a text based on values representing the text characters:

when the values are associate with more than one human language, identifying a particular said human language associated with the string by comparing the string with a targeted dictionary including a plurality of strings associated with the more than one said human language; and

select which of the plurality of strings in the targeted dictionary are to be compared based on the values representing the text characters in the string.

10. The system of claim 9 , wherein the language detection service is independent of an application providing the text.

11. The system of claim 9 , wherein the plurality of strings in the targeted dictionary are individually associated with a single said human language.

12. The system of claim 9 , wherein the plurality of strings includes individual said strings that are one or more of a determiner, a preposition, a connector or a pronoun.

13. The system of claim 9 , wherein the language detection service is further configured to ascertain which said human language is associated with the string from which said human language is associated with a substring in the string.

14. One or more tangible computer-readable media comprising instructions that are executable by a computer to:

determine which human writing system is associated with a string of text characters based on numerical values representing the text characters in the string; and

when the numerical values are associated with more than one human language, compare the string with a targeted dictionary, including a plurality of strings, in which individual said strings in the targeted dictionary are associated with a corresponding said human language, to identify the corresponding said human language associated with the string.

15. One or more tangible computer-readable media as described in claim 14 , wherein the instructions are further executable to ascertain which human language is associated with a substring in the string to identify which said human language is associated with the string.

16. One or more tangible computer-readable media as described in claim 14 , wherein the instructions are further executable to select which of the plurality of strings in the targeted dictionary are to be compared based on the numerical values.

17. One or more tangible computer-readable media as described in claim 14 wherein the targeted dictionary includes individual said strings that are associated with a single said human language.

18. One or more tangible computer-readable media as described in claim 14 , wherein the plurality of strings includes one of more of a determiner, a preposition, a connector or a pronoun.

19. One or more tangible computer-readable media as described in claim 14 , wherein each string in the targeted dictionary is associated with a single human language.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034564/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2009
From: GEORGIEV, DIMITER; YE, SHENGHUA; GUZMAN, GERARDO VILLARREAL; SNYDER, KIERAN; CAVALCANTE, RYAN M.; SAYED, TAREK M.; FEINBERG, YANIV; LIN, YUNG-SHIN
To: MICROSOFT CORPORATION
Reel/Frame 022066/0620 →
Continuity (1)
Related Publication 20090326918A1 · Dec 31, 2009