IP Library Granted Patent US 9,588,958
Granted Patent B2
US 9,588,958 · App. 13/535,638 · Granted Mar 7, 2017

Cross-language text classification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,588,958
App. No.
13/535,638
Granted
Mar 7, 2017
Kind
B2
Abstract

Methods are described for performing classification (categorization) of text documents written in various languages. Language-independent semantic structures are constructed before classifying documents. These structures reflect lexical, morphological, syntactic, and semantic properties of documents. The methods suggested are able to perform cross-language text classification which is based on document properties reflecting their meaning. The methods are applicable to genre classification, topic detection, news analysis, authorship analysis, etc.

Claims (45)

1. A method of performing text classification based on language-independent text features, the method comprising:

performing, by a processor, a first syntactic and semantic analysis of a training natural language text to produce a first plurality of language-independent semantic structures representing a plurality of sentences of the training natural language text;

producing, based on the first plurality of language-independent semantic structures, a text classifier model;

performing a second syntactic and semantic analysis of an input natural language text to produce a second plurality of language-independent semantic structures representing a plurality of sentences of the input natural language text;

extracting, using the second plurality of language-independent semantic structures, a set of features, wherein at least one feature references a semantic class of a language-independent semantic hierarchy comprising a plurality of semantic classes, in which the semantic class exhibits one or more properties inherited from its parent semantic class;

applying the text classifier model to the set of features to produce a classification spectrum comprising a plurality of weight values, wherein each weight value reflects a degree of association of the input natural language text with a particular category of natural language texts; and

associating the input natural language text with one or more categories using the classification spectrum.

2. The method of claim 1 , wherein the second syntactic and semantic analysis further includes determining a grammatical feature of the input natural language text.

3. The method of claim 1 , wherein the second syntactic and semantic analysis further includes determining a lexical feature of the input natural language text.

4. The method of claim 1 , wherein the second syntactic and semantic analysis further includes determining a syntactic feature of the input natural language text.

5. The method of claim 1 , wherein the second syntactic and semantic analysis further includes determining a semantic feature of the input natural language text.

6. The method of claim 1 , wherein the second syntactic and semantic analysis further includes generating a syntactic structure of a sentence of the input natural language text.

7. The method of claim 1 , wherein the categories are represented by language independent categories.

8. A non-transitory computer readable storage medium comprising executable instructions for causing a computing system to perform operations comprising:

performing a first syntactic and semantic analysis of a training natural language text to produce a first plurality of language-independent semantic structures representing a plurality of sentences of the training natural language text; producing, based on the first plurality of language-independent semantic structures, a text classifier model;

performing a second syntactic and semantic analysis of an input natural language text to produce a second plurality of language-independent semantic structures representing a plurality of sentences of the input natural language text;

extracting, using the second plurality of language-independent semantic structures, a set of features, wherein at least one feature references a semantic class of a language-independent semantic hierarchy comprising a plurality of semantic classes, in which the semantic class exhibits one or more properties inherited from its parent semantic class;

applying the text classifier model to the set of features to produce a classification spectrum comprising a plurality of weight values, wherein each weight value references a degree of association of the input natural language text with a particular category of natural language texts; and

associating the input natural language text with one or more categories using the classification spectrum.

9. The non-transitory computer readable storage medium of claim 8 , wherein the second syntactic and semantic analysis further includes determining a grammatical feature of the input natural language text.

10. The non-transitory computer readable medium of claim 8 , wherein the second syntactic and semantic analysis further includes determining a lexical feature of the input natural language text.

11. The non-transitory computer readable medium of claim 8 , wherein the second syntactic and semantic analysis further includes determining a syntactic feature of the input natural language text.

12. The non-transitory computer readable medium of claim 8 , wherein the second syntactic and semantic analysis further includes determining a semantic feature of the input natural language text.

13. The non-transitory computer readable medium of claim 8 , wherein the second syntactic and semantic analysis further includes generating a syntactic structure of a sentence of the input natural language text.

14. The non-transitory computer readable medium of claim 8 , wherein the categories are represented by language independent categories.

15. A computer system adapted to perform text classification based on language-independent text features, the computer system comprising:

a feature extractor adapted to perform operations comprising:

performing a first syntactic and semantic analysis of a training natural language text to produce a first plurality of language-independent semantic structures representing a plurality of sentences of the training natural language text;

producing, based on the first plurality of language-independent semantic structures, a text classifier model;

performing a second syntactic and semantic analysis of an input natural language text to produce a second plurality of language-independent semantic structures representing a plurality of sentences of the input natural language text;

extracting, using the second plurality of language-independent semantic structures, a set of features, wherein at least one feature references a semantic class of a language-independent semantic hierarchy comprising a plurality of semantic classes, in which the semantic class exhibits one or more properties inherited from its parent semantic class; and

a text classifier adapted to perform operations comprising:

applying the text classifier model to the set of features to generate a classification spectrum comprising a plurality of weight values, wherein each weight value references a degree of association of the input natural language text with a particular category of natural language texts; and

associating the input natural language text with one or more categories using the classification spectrum.

16. The computer system of claim 15 , wherein the feature extractor is further adapted to perform operations comprising:

determining a grammatical feature of the input natural language text.

17. The computer system of claim 15 , wherein the feature extractor is further adapted to perform operations comprising:

determining a lexical feature of the input natural language text.

18. The computer system of claim 15 , wherein the feature extractor is further adapted to perform operations comprising:

determining a syntactic feature of the input natural language text.

19. The computer system of claim 15 , wherein the feature extractor is further adapted to perform operations comprising:

determining a semantic feature of the input natural language text.

20. The computer system of claim 15 , wherein the feature extractor is further adapted to perform operations comprising:

generating a syntactic structure of a sentence of the input natural language text.

21. The computer system of claim 15 , wherein the categories are represented by language independent categories.

Assignments (5)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR DOC. DATE PREVIOUSLY RECORDED AT REEL: 042706 FRAME: 0279. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 25, 2017
From: ABBYY INFOPOISK LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 043676/0232 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2017
From: ABBYY INFOPOISK LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 042706/0279 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 16, 2012
From: DANIELYAN, TATIANA; ZUEV, KONSTANTIN; ANISIMOVICH, KONSTANTIN; SELEGEY, VLADIMIR
To: ABBYY INFOPOISK LLC
Reel/Frame 028558/0585 →