IP Library Granted Patent US 12,314,661
Granted Patent B2
US 12,314,661 · App. 18/082,919 · Granted May 27, 2025

Natural language detection

Inventor: Michael Zatsepin (Novokuznetsk, RU)
Assignee: ABBYY Development Inc.
G06F40/284G06F40/263G06V30/153
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,661
App. No.
18/082,919
Granted
May 27, 2025
Kind
B2
Abstract

An example method of language detection includes: identifying a document comprising a plurality of words in one or more natural languages; for each word of at least a subset of words of the document: generating a plurality of sets of tokens representing the word, wherein each set of tokens of the plurality of sets of tokens represents the word using a corresponding plurality of tokens defined for a corresponding natural language of a set of natural languages, and identifying, based on the plurality of sets of tokens, a primary natural language associated with the word; associating each natural language of the set of natural languages with a corresponding word count indicating a number of words of the subset of words for which the natural language has been identified as the primary natural language; identifying, among the set of natural languages, a natural language associated with a maximum word count; and associating the identified natural language with the document.

Claims (72)

1. A method, comprising:

identifying, by a processing device, a document comprising a plurality of words in one or more natural languages;

for each word of at least a subset of words of the document:

generating a plurality of sets of tokens representing the word, wherein each set of tokens of the plurality of sets of tokens represents the word using a corresponding plurality of tokens defined for a corresponding natural language of a set of natural languages, and

identifying, based on the plurality of sets of tokens, a primary natural language associated with the word;

associating each natural language of the set of natural languages with a corresponding word count indicating a number of words of the subset of words for which the natural language has been identified as the primary natural language;

identifying, among the set of natural languages, a natural language associated with a maximum word count; and

associating the identified natural language with the document.

2. The method of claim 1 , further comprising:

identifying, based on the plurality of sets of tokens, an alternative natural language associated with the word.

3. The method of claim 1 , further comprising:

iteratively performing the operations of:

updating the subset of words of the document by removing one or more words associated with the identified natural language,

associating each natural language of the set of natural languages with a corresponding word count indicating a number of words of the updated subset of words for which the natural language has been identified as the primary natural language,

identifying, among the set of natural languages, a natural language associated with a maximum word count, and

associating the identified natural language with the document.

4. The method of claim 3 , wherein the operations are iteratively performed until a number of words of the document that associated with at least one language reaches a certain high threshold.

5. The method of claim 3 , wherein the operations are iteratively performed until a number of words for which a language that has been identified as a primary or alternative language by a current iteration falls below a certain low threshold.

6. The method of claim 1 , further comprising:

performing, based on the identified natural language, a natural language processing task with respect to the document.

7. The method of claim 1 , wherein identifying the document further comprises:

performing optical character recognition (OCR) of an image.

8. The method of claim 1 , wherein the subset of words is identified by applying one or more filtering criteria to a plurality of words comprised by the document.

9. The method of claim 1 , wherein identifying the document further comprises:

splitting an input image into a plurality of portions;

removing one or more portions satisfying one or more geometric criteria;

performing optical character recognition (OCR) of at least a subset of remaining portions.

10. A system comprising:

a memory; and

a processing device operatively coupled to the memory, the processing device configured to:

identify a document comprising a plurality of words in one or more natural languages;

for each word of at least a subset of words of the document:

generate a plurality of sets of tokens representing the word, wherein each set of tokens of the plurality of sets of tokens represents the word using a corresponding plurality of tokens defined for a corresponding natural language of a set of natural languages, and

identify, based on the plurality of sets of tokens, a primary natural language associated with the word;

associate each natural language of the set of natural languages with a corresponding word count indicating a number of words of the subset of words for which the natural language has been identified as the primary natural language;

identify, among the set of natural languages, a natural language associated with a maximum word count; and

associate the identified natural language with the document.

11. The system of claim 10 , wherein the processing device is further configured to:

identify, based on the plurality of sets of tokens, an alternative natural language associated with the word.

12. The system of claim 10 , wherein the processing device is further configured to:

iteratively perform the operations of:

updating the subset of words of the document by removing one or more words associated with the identified natural language,

associating each natural language of the set of natural languages with a corresponding word count indicating a number of words of the updated subset of words for which the natural language has been identified as the primary natural language,

identifying, among the set of natural languages, a natural language associated with a maximum word count, and

associating the identified natural language with the document.

13. The system of claim 10 , wherein the processing device is further configured to:

perform, based on the identified natural language, a natural language processing task with respect to the document.

14. The system of claim 10 , wherein identifying the document further comprises:

performing optical character recognition (OCR) of an image.

15. The system of claim 10 , wherein the subset of words is identified by applying one or more filtering criteria to a plurality of words comprised by the document.

16. The system of claim 10 , wherein identifying the document further comprises:

splitting an input image into a plurality of portions;

removing one or more portions satisfying one or more geometric criteria;

performing optical character recognition (OCR) of at least a subset of remaining portions.

17. A non-transitory computer-readable storage medium including executable instructions that, when executed by a processing device, cause the processing device to:

identify a document comprising a plurality of words in one or more natural languages;

for each word of at least a subset of words of the document:

generate a plurality of sets of tokens representing the word, wherein each set of tokens of the plurality of sets of tokens represents the word using a corresponding plurality of tokens defined for a corresponding natural language of a set of natural languages, and

identify, based on the plurality of sets of tokens, a primary natural language associated with the word;

associate each natural language of the set of natural languages with a corresponding word count indicating a number of words of the subset of words for which the natural language has been identified as the primary natural language;

identify, among the set of natural languages, a natural language associated with a maximum word count; and

associate the identified natural language with the document.

18. The non-transitory computer-readable storage medium of claim 17 , further comprising executable instructions that, when executed by the processing device, cause the processing device to:

identify, based on the plurality of sets of tokens, an alternative natural language associated with the word.

19. The non-transitory computer-readable storage medium of claim 17 , further comprising executable instructions that, when executed by the processing device, cause the processing device to:

iteratively perform the operations of:

updating the subset of words of the document by removing one or more words associated with the identified natural language,

associating each natural language of the set of natural languages with a corresponding word count indicating a number of words of the updated subset of words for which the natural language has been identified as the primary natural language,

identifying, among the set of natural languages, a natural language associated with a maximum word count, and

associating the identified natural language with the document.

20. The non-transitory computer-readable storage medium of claim 17 , further comprising executable instructions that, when executed by the processing device, cause the processing device to:

perform, based on the identified natural language, a natural language processing task with respect to the document.

Assignments (2)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2022
From: ZATSEPIN, MICHAEL
To: ABBYY DEVELOPMENT INC.
Reel/Frame 062128/0668 →
Continuity (1)
Related Publication 20240202444A1 · Jun 20, 2024
References Cited (10)
US 8224642B2 · Goswami · 2012 [cited by applicant]
US 8635061B2 · Li · 2014 [cited by applicant]
US 8938384B2 · Goswami · 2015 [cited by applicant]
US 9330086B2 · Zhao · 2016 [cited by applicant]
US 10460192B2 · Gopalakrishnan · 2019 [cited by applicant]
RU 2500024C2 · 2013 [cited by applicant]
WO 2021161095A1 · 2021 [cited by applicant]
Lui, Marco et al, Department of Computing and Information Systems, The University of Melbourne, “Automatic Detection and Language Identification of Multilingual Documents”, 2014, 14 pages. [cited by applicant]
Jauhiainen, Tommi et al, Journal of Artificial Intelligence Research, “Automatic Language Identification in Texts: A Survey”, arXiv:1804.08186v2 [cs.CL] Nov. 21, 2018, 103 pages. [cited by applicant]
Barlas, P et al., HAL open science, “Language Identification in Document Images”, HAL Id: hal-01282930 https://hal.archives-ouvertes.fr/hal-01282930, Submitted on Mar. 8, 2016, 15 pages. [cited by applicant]