IP Library Granted Patent US 10,977,444
Granted Patent B2
US 10,977,444 · App. 16/233,674 · Granted Apr 13, 2021

Method and system for identifying key terms in digital document

Inventors: Gaurav Tripathi (Pune, IN); Vatsal Agarwal (Rampur, IN); Sudhanshu Shekhar (Patna, IN)
Assignee: Innoplexus AG
G06F40/30G06F40/279G06F40/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,977,444
App. No.
16/233,674
Granted
Apr 13, 2021
Kind
B2
Abstract

Disclosed is a method and a system for identifying key terms in a digital document. The method comprises providing the digital document and analysing the digital document to identify key terms in the digital document. The digital document includes a first text in a first language. Furthermore, analysing the digital document comprises translating the first text in the first language to obtain a second text in a second language, translating the first text in the first language to obtain a third text in a third language, translating the obtained second text in the second language to obtain a fourth text in the third language, comparing at least one pair of first text, second text, third text and fourth text to identify at least one set of similar text between the compared at least one pair, and processing the set of similar text to obtain key terms in the digital document.

Claims (35)

1. A method of identifying key terms in a digital document, wherein the method comprises:

providing the digital document, wherein the digital document includes a first text in a first language; and

analysing the digital document to identify the key terms in the digital document, wherein analysing comprises:

translating the first text in the first language to obtain a second text in a second language;

translating the first text in the first language to obtain a third text in a third language;

translating the obtained second text in the second language to obtain a fourth text in the third language;

comparing the third text and the fourth text to identify at least one set of similar text therebetween, wherein the at least one set of similar text is identified on the basis of known similarity methods, and wherein the at least one set of similar text is the number of recurring terms identified in the compared pair of text and comprises the terms which display little variation even after multiple language translations; and

processing the at least one set of similar text to obtain key terms in the digital document, wherein processing the at least one set of similar text comprises validating the at least one set of similar text based on an ontology.

2. The method of claim 1 , wherein the method further comprises developing the ontology using at least one curated database by:

applying conceptual indexing to plurality of terms stored in the at least one curated database;

identifying semantic associations, between the plurality of terms, established in the at least one curated database; and

identifying at least one class tagged with the plurality of terms in the at least one curated database.

3. The method of claim 1 , wherein the method further comprises classifying the identified key terms based on the ontology.

4. A system for identifying key terms in a digital document, wherein the system comprises:

a translation module operable to:

receive the digital document, wherein the digital document includes a first text in a first language;

translate the first text in the first language to obtain a second text in a second language;

translate the first text in the first language to obtain a third text in a third language; and

translate the obtained second text in the second language to obtain a fourth text in the third language;

a processing module communicably coupled to the translation module, the processing module operable to identify the key terms by:

comparing the third text and the fourth text to identify at least one set of similar text therebetween, wherein the at least one set of similar text is the number of recurring terms identified in the compared pair of text and comprises the terms which display little variation even after multiple language translations; and

processing the at least one set of similar text to obtain key terms in the digital document, wherein processing the at least one set of similar text comprises validating the at least one set of similar text based on an ontology.

5. The system of claim 4 , wherein the database arrangement is operable to store at least one curated database, wherein the ontology is developed using at least one curated database by:

applying conceptual indexing to plurality of terms stored in the at least one curated database;

identifying semantic associations, between the plurality of terms, established in the at least one curated database; and

identifying at least one class tagged with the plurality of terms in the at least one curated database.

6. The system of claim 4 , wherein the processing module is operable to further classify the identified key terms based on the ontology.

7. A non-transitory computer readable storage medium, containing program instructions for execution on a computer system, which when executed by a computer, cause the computer to perform method steps for identifying key terms in a digital document, the method comprising the steps of:

providing the digital document, wherein the digital document includes a first text in a first language; and

analysing the digital document to identify the key terms in the digital document, wherein analysing comprises:

translating the first text in the first language to obtain a second text in a second language;

translating the first text in the first language to obtain a third text in a third language;

translating the obtained second text in the second language to obtain a fourth text in the third language;

comparing the third text and the fourth text to identify at least one set of similar text therebetween, wherein the at least one set of similar text is the number of recurring terms identified in the compared pair of text and comprises the terms which display little variation even after multiple language translations; and

processing the at least one set of similar text to obtain key terms in the digital document, wherein processing the at least one set of similar text comprises validating the at least one set of similar text based on an ontology.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2019
From: INNOPLEXUS CONSULTING SERVICES PVT. LTD.
To: INNOPLEXUS AG
Reel/Frame 048730/0480 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2019
From: TRIPATHI, GAURAV; AGARWAL, VATSAL; SHEKHAR, SUDHANSHU
To: INNOPLEXUS CONSULTING SERVICES PVT. LTD.
Reel/Frame 048686/0426 →
Priority Claims (1)
GB 1722305 · Dec 30, 2017 · national
Continuity (1)
Related Publication 20190205389A1 · Jul 4, 2019