IP Library › Granted Patent US 11,086,913
Granted Patent B2
US 11,086,913 · App. 15/922,779 · Granted Aug 10, 2021

Named entity recognition from short unstructured text

Inventors: Navaneethan Santhanam (Chennai, IN); Saurabh Arora (Chennai, IN); Satyam Saxena (Chennai, IN); Anuj Gupta (Chennai, IN)
G06F16/3346G06F16/335G06F16/3347G06F16/353G06F40/216G06F40/295
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,086,913
App. No.
15/922,779
Granted
Aug 10, 2021
Kind
B2
Abstract

A process for extracting and recognizing named entities from a short unstructured chat-style text input. The process may tokenize an inbound electronic message, and use a combination of entity specific classifiers and databases comprising known named entities such as gazetteer(s) to identify one or more named entities within the inbound electronic message. The identified named entities are then compiled as response message and transmitted to the user.

Claims (60)

1. A computer-implemented method for recognizing named entities in a message, comprising:

receiving an electronic message;

tokenizing the electronic message;

probabilistically identifying, using either a short-phrase classifier or a regular text classifier, whether each token constitutes a named entity, thereby identifying one or more named entities within the electronic message, wherein each token is a continuous sequence of characters grouped together, and wherein the probabilistically identifying comprises accessing the short-term classifier or the regular text classifier depending on whether the number of tokens being less than or equal to a predefined number or a threshold number;

comparing each token with one or more databases comprising known named entities, thereby identifying the one or more name entities within the electronic message, wherein

the probabilistically identifying and the comparing of each token is performed simultaneously; and

returning a response message to the user when the one or more named entities are found in the electronic message, wherein the response message identifies the one or more named entities from the comparison, or

returning a null response message when the comparison fails to identify named entities in the electronic message.

2. The computer-implemented method of claim 1 , wherein the named entities comprise a name, a location, and/or an organization.

3. The computer-implemented method of claim 1 , wherein the tokenizing of the electronic message comprises converting each sub-word, word, or punctuation within the electronic message into tokens for recognizing the one or more named entities.

4. The computer-implemented method of claim 1 , wherein the probabilistically identifying comprises

accessing a short-phrase classifier to probabilistically identify whether one or more tokens or a sequence of tokens constitutes the one or more named entities, when a number of tokens is less than or equal to a threshold.

5. The computer-implemented method of claim 1 , wherein a process for probabilistically identifying comprises

projecting the one or more tokens or the sequence of tokens in a d-dimension vector space, where d is 100, 200, 300, or 600; and

further projecting a token vector for each of the one or more tokens or the sequence of tokens to a dimensional space where linear separation is possible, providing a maximum likelihood of distinguishing the one or more tokens or the sequences of tokens representing the one or more named entities from other tokens or sequences that do not represent named entities.

6. The computer-implemented method of claim 1 , further comprising

accessing a regular text classifier to probabilistically identify whether one or more tokens or a sequence of tokens constitutes the one or more named entities, when a number of tokens is greater than a threshold.

7. The computer-implemented method of claim 1 , further comprising

accessing one or more gazetteer lookup databases to scan for a recognized named entity; and

filtering each token that fails to match with a named entity to quickly identify a token that contains the named entity.

8. The computer-implemented method of claim 1 , further comprising:

removing duplicates for the one or more named entities by comparing a result from a short-phrase classifier, a regular text classifier, a gazetteer lookup database, or any combination thereof.

9. The computer-implemented method of claim 1 , further comprising:

combining one or more tokens to form the one or more named entities, wherein the one or more tokens are results from a short-phrase classifier, a regular text classifier, a gazetteer lookup database, or any combination thereof.

10. The computer-implemented method of claim 1 , further comprising:

selecting a superset of the one or more named entities when a short-phrase classifier, a regular text classifier, a gazetteer lookup database, or any combination thereof returns the superset of the one or more named entities and a subset of the superset of the one or more named entities.

11. An apparatus for recognizing named entities in an electronic communication, comprising:

at least one processor; and

memory comprising a set of instructions, wherein

the set of instructions are configured to cause the at least one processor to

receive an electronic message;

tokenize the electronic message;

probabilistically identify, using either a short-phrase classifier or a regular text classifier, whether each token constitutes a named entity, thereby identifying one or more named entities within the electronic message, wherein each token is a continuous sequence of characters grouped together, and wherein the probabilistically identifying comprises access the short-term classifier or the regular text classifier depending on whether the number of tokens being less than or equal to a predefined number or a threshold number;

compare each token, with one or more databases comprising known named entities, thereby identified the one or more name entities within the electronic message, wherein

the probabilistically identifying and the comparing of each token is performed simultaneously; and

return a response message to the user when the one or more named entities are found in the electronic message, wherein the response message identifies the one or more named entities from the comparison, or

return a null response message when the comparison fails to identify named entities in the electronic message.

12. The apparatus of claim 11 , wherein the named entities comprise a name, a location, and/or an organization.

13. The apparatus of claim 11 , wherein the set of instructions are further configured to cause the at least one processor to convert each sub-word, word, or punctuation within the electronic message into tokens for recognizing the one or more named entities.

14. The apparatus of claim 11 , wherein the set of instructions are further configured to cause the at least one processor to

access a short-phrase classifier to probabilistically identify whether one or more tokens or a sequence of tokens constitutes the one or more named entities when a number of tokens is less than or equal to a threshold.

15. The apparatus of claim 11 , wherein the set of instructions are further configured to cause the at least one processor to

project the one or more tokens or the sequence of tokens in a d-dimension vector space, where d is 100, 200, 300, or 600; and

further project a token vector for each of the one or more tokens or the sequence of tokens to a dimensional space where linear separation is possible, providing a maximum likelihood of distinguishing the one or more tokens or the sequences of tokens representing the one or more named entities from other tokens or sequences that do not represent named entities.

16. The apparatus of claim 12 , wherein the set of instructions are further configured to cause the at least one processor to

access a regular text classifier to probabilistically identify whether one or more tokens or a sequence of tokens constitutes the one or more named entities when a number of tokens is greater than a threshold.

17. The apparatus of claim 11 , wherein the set of instructions are further configured to cause the at least one processor to

access one or more gazetteer lookup databases to scan for a recognized named entity; and

filter each token that fails to match with a named entity to quickly identify a token that contains the named entity.

18. The apparatus of claim 11 , wherein the set of instructions are further configured to cause the at least one processor to

remove duplicates for the one or more named entities by comparing a result from a short-phrase classifier, a regular text classifier, a gazetteer lockup database, or any combination thereof.

19. The apparatus of claim 11 , wherein the set of instructions are further configured to cause the at least one processor to

combine one or more tokens to form the one or more named entities, the one or more tokens are results from a short-phrase classifier, a regular text classifier, a gazetteer lockup database, or any combination thereof.

20. The apparatus of claim 11 , wherein the set of instructions are further configured to cause the at least one processor to

select a superset of the one or more named entities when a short-phrase classifier, a regular text classifier, a gazetteer lockup database, or any combination thereof returns the superset of the one or more named entities and a subset of the superset of the one or more named entities.

21. A computer-implemented process, comprising:

tokenizing an electronic message received from another computing device;

probabilistically identifying, using either a short-phrase classifier or a regular text classifier, whether each token constitutes a named entity, thereby identifying one or more named entities within the electronic message, wherein each token is a continuous sequence of characters grouped together and wherein the probabilistically identifying comprises accessing the short-term classifier or the regular text classifier depending on whether the number of tokens being less than or equal to a predefined number or a threshold number; and

returning a response message to the user when the one or more named entities are found in the electronic message, or

returning a null response message when the comparison fails to identify named entities in the electronic message.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2023
From: FRESHWORKS TECHNOLOGIES PRIVATE LIMITED
To: FRESHWORKS INC.
Reel/Frame 064206/0935 →
CORRECTIVE ASSIGNMENT TO CORRECT THE TITLE INSIDE THE ASSIGNMENT DOCUMENT PREVIOUSLY RECORDED AT REEL: 045243 FRAME: 0317. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 27, 2018
From: SANTHANAM, NAVANEETHAN; ARORA, SAURABH; SAXENA, SATYAM; GUPTA, ANUJ
To: FRESHWORKS TECHNOLOGIES PRIVATE LIMITED
Reel/Frame 046481/0476 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2018
From: SANTHANAM, NAVANEETHAN; ARORA, SAURABH; SAXENA, SATYAM; GUPTA, ANUJ
To: FRESHWORKS TECHNOLOGIES PRIVATE LIMITED
Reel/Frame 045243/0317 →
Priority Claims (1)
IN 201841000125 · Jan 2, 2018 · national
Continuity (1)
Related Publication 20190205463A1 · Jul 4, 2019