IP Library Granted Patent US 11,003,950
Granted Patent B2
US 11,003,950 · App. 16/369,358 · Granted May 11, 2021

System and method to identify entity of data

Inventors: Vatsal Agarwal (Rampur, IN); Vivek Verma (Pune, IN)
Assignee: Innoplexus AG
G06K9/6256G06F40/16G06F40/284G06K9/6264G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,003,950
App. No.
16/369,358
Granted
May 11, 2021
Kind
B2
Abstract

Disclosed is system comprising data processing arrangement including processors configured to receive sentences from unlabeled training data set; tokenize, using tokenizer module, sentences to obtain tokens; generate character level features for character of tokens of sentences; generate token level feature for each token of the sentences, wherein token level feature of token in sentence is identified using token coordinates of token and token coordinates of tokens neighboring token in sentence; train artificial neural network adapted to identify entities in sentences to determine first trend set, wherein training is based on received sentences, character level features for each character of tokens of sentences and token level feature for tokens of sentences; train the artificial neural network on set of labelled data to determine second trend set; identify, using identifier module, entity in text content, wherein identifier module uses first trend set and second trend set determined by artificial neural network.

Claims (37)

1. A system comprising:

a data processing arrangement including one or more processors configured to:

receive one or more sentences from unlabeled training data set;

tokenize, using a tokenizer module, the one or more sentences to obtain a plurality of tokens;

generate character level features for each character of the plurality of tokens of the one or more sentences, wherein a given token is divided into a plurality of characters, and wherein the character level features of each of the character is identified based on demographics of each character, wherein the character level features comprise a location of a character in a multi-dimensional hierarchical space;

generate a token level feature for each token of the plurality of tokens of the one or more sentences, wherein the token level feature of a given token in a given sentence is identified using token coordinates of the given token and token coordinates of tokens neighboring the given token in the given sentence, wherein token coordinates of a token represent a location of the token in the multi-dimensional hierarchical space;

train an artificial neural network configured to identify one or more entities in sentences to determine a first trend set, wherein the training is based on the received one or more sentences, the character level features for each character of the plurality of tokens of the one or more sentences and the token level feature for each token of the plurality of tokens of the one or more sentences;

train the artificial neural network on a set of labelled data to determine a second trend set;

identify, using an identifier module, an entity in a text content, wherein the identifier module uses the first trend set and the second trend set determined by the artificial neural network.

2. The system of claim 1 , wherein the system further includes a lexicon ontology represented into a multi-dimensional hierarchical space.

3. The system of claim 1 , wherein first trend set comprises one or more distributions of tokens in the one or more sentences of unlabeled training data set.

4. The system of claim 1 , wherein the second trend set comprises probability score for each of the plurality of tokens of the one or more sentences.

5. The system of claim 1 , wherein the training the artificial neural network involves semi-supervised training and transfer learning approach.

6. The system of claim 5 , wherein the artificial neural network is a recurrent neural network.

7. A method implemented via a system comprising:

a data processing arrangement including one or more processors configured to:

receive one or more sentences from unlabeled training data set;

tokenize, using a tokenizes module, the one or more sentences to obtain a plurality of tokens;

generate character level features for each character of the plurality of tokens of the one or more sentences, wherein a given token is divided into a plurality of characters, and wherein the character level features of each of the character is identified based on demographics of each character, wherein the character level features comprise a location of a character in a multi-dimensional hierarchical space;

generate a token level feature for each token of the plurality of tokens of the one or more sentences, wherein the token level feature of a given token in a given sentence is identified using token coordinates of the given token and token coordinates of tokens neighboring the given token in the given sentence, wherein token coordinated of a token represent a location of the token in the multi-dimensional hierarchical space;

train an artificial neural network adapted to identify one or more entities in sentences to determine a first trend set, wherein the training is based on the received one or more sentences, the character level features for each character of the plurality of tokens of the one or more sentences and the token level feature for each token of the plurality of tokens of the one or more sentences;

train the artificial neural network on a set of labelled data to determine a second trend set;

identify, using an identifier module, an entity in a text content, wherein the identifier module uses the first trend set and the second trend set determined by the artificial neural network.

8. The method of claim 7 , wherein the method further includes a lexicon ontology represented into a multi-dimensional hierarchical space.

9. The method of claim 7 , wherein first trend set comprises one or more distributions of tokens in the one or more sentences of unlabeled training data set.

10. The method of claim 7 , wherein the second trend set comprises probability score for each of the plurality of tokens of the one or more sentences.

11. The method of claim 7 , wherein the training the artificial neural network involves semi-supervised training and transfer learning approach.

12. The method of claim 11 , wherein the artificial neural network is a recurrent neural network.

13. A non-transitory computer readable storage medium containing program instructions for execution on a computer, which when executed by the computer, cause the computer to perform a method, wherein the method is implemented via a system comprising:

a data processing arrangement including one or more processors configured to:

receive one or more sentences from unlabeled training data set;

tokenize, using a tokenizer module, the one or more sentences to obtain a plurality of tokens;

generate character level features for each character of the plurality of tokens of the one or more sentences, wherein a given token is divided into a plurality of characters, and wherein the character level features of each of the character is identified based on demographics of each character, wherein the character level features comprise a location of a character in a multi-dimensional hierarchical space;

generate a token level feature for each token of the plurality of tokens of the one or more sentences, wherein the token level feature of a given token in a given sentence is identified using token coordinates of the given token and token coordinates of tokens neighboring the given token in the given sentence, wherein token coordinated of a token represent a location of the token in the multi-dimensional hierarchical space;

train an artificial neural network adapted to identify one or more entities in sentences to determine a first trend set, wherein the training is based on the received one or more sentences, the character level features for each character of the plurality of tokens of the one or more sentences and the token level feature for each token of the plurality of tokens of the one or more sentences;

train the artificial neural network on a set of labelled data to determine a second trend set;

identify, using an identifier module, an entity in a text content, wherein the identifier module uses the first trend set and the second trend set determined by the artificial neural network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2019
From: INNOPLEXUS CONSULTING SERVICES PVT. LTD.
To: INNOPLEXUS AG
Reel/Frame 051004/0565 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2019
From: AGARWAL, VATSAL; VERMA, VIVEK
To: INNOPLEXUS CONSULTING SERVICES PVT. LTD.
Reel/Frame 048739/0577 →
Continuity (1)
Related Publication 20200311473A1 · Oct 1, 2020