IP Library Granted Patent US 9,965,458
Granted Patent B2
US 9,965,458 · App. 14/964,512 · Granted May 8, 2018

Intelligent system that dynamically improves its knowledge and code-base for natural language understanding

Inventors: Robert J. Munro (San Francisco, CA); Rob Voigt (Palo Alto, CA); Schuyler D. Erle (San Francisco, CA); Brendan D. Callahan (Philadelphia, PA); Gary C. King (Los Altos, CA); Jessica D. Long (San Francisco, CA); Jason Brenier (Oakland, CA); Tripti Saxena (Cupertino, CA); Stefan Krawczyk (Menlo Park, CA)
Assignee: Sansa AI Inc.
G06F17/277G06F17/2715G06F17/2785
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,965,458
App. No.
14/964,512
Granted
May 8, 2018
Kind
B2
Abstract

Systems, methods, and apparatuses are presented for a novel natural language tokenizer and tagger. In some embodiments, a method for tokenizing text for natural language processing comprises: generating from a pool of documents, a set of statistical models comprising one or more entries each indicating a likelihood of appearance of a character/letter sequence in the pool of documents; receiving a set of rules comprising rules that identify character/letter sequences as valid tokens; transforming one or more entries in the statistical models into new rules that are added to the set of rules when the entries indicate a high likelihood; receiving a document to be processed; dividing the document to be processed into tokens based on the set of statistical models and the set of rules, wherein the statistical models are applied where the rules fail to unambiguously tokenize the document; and outputting the divided tokens for natural language processing.

Claims (53)

1. A method for tokenizing text for natural language processing, the method comprising:

generating, by one or more processors in a natural language processing platform, and from a pool of documents, a set of statistical models comprising one or more entries each indicating a likelihood of appearance of a character/letter sequence in the pool of documents;

receiving, by the one or more processors, a set of rules comprising rules that identify character/letter sequences as valid tokens;

transforming, by the one or more processors, one or more entries in the statistical models into new rules that are added to the set of rules when the entries indicate a high likelihood;

receiving, by the one or more processors, a document to be processed;

dividing, by the one or more processors, the document to be processed into tokens based on the set of statistical models and the set of rules, wherein the statistical models are applied where the rules fail to unambiguously tokenize the document; and

outputting, by the one or more processors, the divided tokens for natural language processing.

2. The method of claim 1 , wherein the set of statistical models further comprises statistical models based on human annotation, and the method further comprises:

generating, by the one or more processors, one or more human readable prompts configured to elicit annotations of one or more documents in the pool of documents, wherein the annotations comprise identification of one or more character/letter sequences in the documents as valid tokens;

receiving, by the one or more processors, one or more annotations elicited by the human readable prompts; and

generating, by the one or more processors, statistical models based on the received annotations, wherein the statistical models comprise one or more entries each indicating a likelihood of appearance of a character/letter sequence in the character/letter sequences annotated as valid tokens.

3. The method of claim 1 , wherein:

the document to be processed is in one or more languages; and

the divided tokens are outputted in a language agnostic format.

4. The method of claim 3 , wherein:

the document to be processed is in more than one language;

the set of rules further comprises a rule that divides portions of the document in different languages into different segments; and

the segments of the document in different languages are divided into tokens based on a different combination of rules and statistical models.

5. The method of claim 1 , wherein the set of rules further comprises a rule that triggers the application of rules and/or statistical models for further tokenization.

6. The method of claim 1 , wherein at least one of the divided tokens contains a morpheme.

7. The method of claim 1 , wherein at least one of the divided tokens contains a group of words.

8. The method of claim 7 , wherein at least one of the divided tokens contains a turn in a conversation.

9. The method of claim 1 , wherein dividing the document to be processed into tokens based on the set of statistical models comprises comparing statistical likelihood of more than one candidate set of tokens.

10. The method of claim 9 , wherein the candidate set of tokens that contains tokens with smallest sizes is preferred.

11. The method of claim 9 , wherein more than one candidate set of tokens is outputted for natural language processing.

12. The method of claim 1 , wherein:

the set of statistical models further comprises one or more statistical models for normalizing variants of a token into a single token and/or the set of rules further comprises one or more rules for normalizing variants of a token into a single token; and

the method further comprises normalizing variants of a token into a single token based on the statistical models and/or the rules.

13. The method of claim 1 , wherein:

the set of statistical models further comprises one or more statistical models for adding tags to the tokens and/or the set of rules further comprises one or more rules for adding tags to the tokens; and

the method further comprises adding tags to the tokens based on the statistical models and/or the rules.

14. The method of claim 13 , wherein the tags are based on semantic information and/or structural information.

15. The method of claim 1 , wherein the set of rules further comprises one or more rules that identify markup language content, an Internet address, a hashtag, or an emoji/emoticon.

16. The method of claim 1 , wherein the set of statistical models and/or the set of rules are adjusted based at least in part on an author of the document.

17. The method of claim 1 , wherein the set of statistical models and/or the set of rules are based at least in part on intra-document information.

18. An apparatus for tokenizing text for natural language processing, the apparatus comprising one or more processors configured to:

generate from a pool of documents a set of statistical models comprising one or more entries each indicating a likelihood of appearance of a character/letter sequence in the pool of documents;

receive a set of rules comprising rules that identify character/letter sequences as valid tokens;

transform one or more entries in the statistical models into new rules that are added to the set of rules when the entries indicate a high likelihood;

receive a document to be processed;

divide the document to be processed into tokens based on the set of statistical models and the set of rules, wherein the statistical models are applied where the rules fail to unambiguously tokenize the document; and

output the divided tokens for natural language processing.

19. The apparatus of claim 18 , wherein the set of statistical models further comprises statistical models based on human annotation, and the one or more processors are further configured to:

generate one or more human readable prompts configured to elicit annotations of one or more documents in the pool of documents, wherein the annotations comprise identification of one or more character/letter sequences in the documents as valid tokens;

receive one or more annotations elicited by the human readable prompts; and

generate statistical models based on the received annotations, wherein the statistical models comprise one or more entries each indicating a likelihood of appearance of a character/letter sequence in the character/letter sequences annotated as valid tokens.

20. A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to:

generate from a pool of documents a set of statistical models comprising one or more entries each indicating a likelihood of appearance of a character/letter sequence in the pool of documents;

receive a set of rules comprising rules that identify character/letter sequences as valid tokens;

transform one or more entries in the statistical models into new rules that are added to the set of rules when the entries indicate a high likelihood;

receive a document to be processed;

divide the document to be processed into tokens based on the set of statistical models and the set of rules, wherein the statistical models are applied where the rules fail to unambiguously tokenize the document; and

output the divided tokens for natural language processing.

Assignments (13)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2024
From: 100.CO GLOBAL HOLDINGS, LLC
To: AI IP INVESTMENTS LTD.
Reel/Frame 066636/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2023
From: DAASH INTELLIGENCE, INC.
To: 100.CO GLOBAL HOLDINGS, LLC
Reel/Frame 064420/0108 →
CHANGE OF NAME Recorded Mar 7, 2023
From: 100.CO TECHNOLOGIES, INC.
To: DAASH INTELLIGENCE, INC.
Reel/Frame 062992/0333 →
NUNC PRO TUNC ASSIGNMENT Recorded Dec 16, 2022
From: 100.CO, LLC
To: 100.CO TECHNOLOGIES, INC.
Reel/Frame 062131/0714 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE CITY PREVIOUSLY RECORDED AT REEL: 055929 FRAME: 0975. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 5, 2021
From: AI IP INVESTMENTS LTD.
To: 100.CO, LLC
Reel/Frame 056151/0150 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2021
From: AI IP INVESTMENTS LTD.
To: 100.CO, LLC
Reel/Frame 055929/0975 →
CORRECTIVE ASSIGNMENT TO CORRECT THE COVENANT INFORMATION TO BE UPDATED FROM AIRPARC HOLDING PTE. LTD. AND REPLACED WITH TREVOR HEALY (SEE MARKED ASSIGNMENT) PREVIOUSLY RECORDED ON REEL 047110 FRAME 0510. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 24, 2021
From: HEALY, TREVOR
To: AIPARC HOLDINGS PTE. LTD.
Reel/Frame 055404/0561 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: AIPARC HOLDINGS PTE. LTD.
To: AI IP INVESTMENTS LTD
Reel/Frame 055377/0995 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2018
From: IDIBON, INC.
To: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
Reel/Frame 047110/0178 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2018
From: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
To: HEALY, TREVOR
Reel/Frame 047110/0449 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2018
From: HEALY, TREVOR
To: AIPARC HOLDINGS PTE. LTD.
Reel/Frame 047110/0510 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2016
From: VOIGT, ROB
To: IDIBON, INC.
Reel/Frame 038709/0030 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2016
From: MUNRO, ROBERT J.; ERLE, SCHUYLER D.; CALLAHAN, BRENDAN D.; KING, GARY C.; LONG, JESSICA D.; BRENIER, JASON; SAXENA, TRIPTI; KRAWCZYK, STEFAN
To: IDIBON, INC.
Reel/Frame 038648/0564 →
Continuity (7)
Provisional Application 62089736 · Dec 9, 2014
Provisional Application 62089742 · Dec 9, 2014
Provisional Application 62089745 · Dec 9, 2014
Provisional Application 62089747 · Dec 9, 2014
Provisional Application 62254090 · Nov 11, 2015
Provisional Application 62254095 · Nov 11, 2015
Related Publication 20160162466A1 · Jun 9, 2016