IP Library Granted Patent US 11,675,977
Granted Patent B2
US 11,675,977 · App. 16/832,632 · Granted Jun 13, 2023

Intelligent system that dynamically improves its knowledge and code-base for natural language understanding

Inventors: Robert J. Munro (San Francisco, CA); Rob Voigt (Palo Alto, CA); Schuyler D. Erle (San Francisco, CA); Brendan D. Callahan (Philadelphia, PA); Gary C. King (Los Altos, CA); Jessica D. Long (San Francisco, CA); Jason Brenier (Oakland, CA); Tripti Saxena (Cupertino, CA); Stefan Krawczyk (Menlo Park, CA)
Assignee: Daash Intelligence, Inc.
G06F40/284G06F40/216G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,675,977
App. No.
16/832,632
Granted
Jun 13, 2023
Kind
B2
Abstract

Systems, methods, and apparatuses are presented for a novel natural language tokenizer and tagger. In some embodiments, a method for tokenizing text for natural language processing comprises: generating from a pool of documents, a set of statistical models comprising one or more entries each indicating a likelihood of appearance of a character/letter sequence in the pool of documents; receiving a set of rules comprising rules that identify character/letter sequences as valid tokens; transforming one or more entries in the statistical models into new rules that are added to the set of rules when the entries indicate a high likelihood; receiving a document to be processed; dividing the document to be processed into tokens based on the set of statistical models and the set of rules, wherein the statistical models are applied where the rules fail to unambiguously tokenize the document; and outputting the divided tokens for natural language processing.

Claims (53)

1. A method for tokenizing text for natural language processing, the method comprising:

generating, by one or more processors in a natural language processing platform, and from a pool of documents, a set of statistical models comprising one or more entries each indicating a likelihood of appearance of a character/letter sequence in the pool of documents;

receiving, by the one or more processors, a set of rules comprising rules that identify character/letter sequences as valid tokens;

transforming, by the one or more processors, one or more entries in the statistical models into new rules that are added to the set of rules when the entries indicate a high likelihood;

receiving, by the one or more processors, a document to be processed;

dividing, by the one or more processors, the document to be processed into tokens based on the set of statistical models and the set of rules; and

outputting, by the one or more processors, the tokens for natural language processing.

2. The method of claim 1 , wherein the set of statistical models further comprises statistical models based on human annotation, and the method further comprises:

generating, by the one or more processors, one or more human readable prompts configured to elicit annotations of one or more documents in the pool of documents, wherein the annotations comprise identification of one or more character/letter sequences in the documents as valid tokens;

receiving, by the one or more processors, one or more annotations elicited by the human readable prompts; and

generating, by the one or more processors, statistical models based on the received annotations, wherein the statistical models comprise one or more entries each indicating a likelihood of appearance of a character/letter sequence in the character/letter sequences annotated as valid tokens.

3. The method of claim 1 , wherein:

the document to be processed is in one or more languages; and

the tokens are outputted in a language agnostic format.

4. The method of claim 3 , wherein:

the document to be processed is in more than one language;

the set of rules further comprises a rule that divides portions of the document in different languages into different segments; and

the segments of the document in different languages are divided into tokens based on a different combination of rules and statistical models.

5. The method of claim 1 , wherein the set of rules further comprises a rule that triggers the application of rules and/or statistical models for further tokenization.

6. The method of claim 1 , wherein at least one of the tokens contains a morpheme.

7. The method of claim 1 , wherein at least one of the tokens contains a group of words.

8. The method of claim 7 , wherein at least one of the tokens contains a turn in a conversation.

9. The method of claim 1 , wherein dividing the document to be processed into tokens based on the set of statistical models comprises comparing statistical likelihood of more than one candidate set of tokens.

10. The method of claim 9 , wherein the candidate set of tokens that contains tokens with smallest sizes is preferred.

11. The method of claim 9 , wherein more than one candidate set of tokens is outputted for natural language processing.

12. The method of claim 1 , wherein:

the set of statistical models further comprises one or more statistical models for normalizing variants of a token into a single token and/or the set of rules further comprises one or more rules for normalizing variants of a token into a single token; and

the method further comprises normalizing variants of a token into a single token based on the statistical models and/or the rules.

13. The method of claim 1 , wherein:

the set of statistical models further comprises one or more statistical models for adding tags to the tokens and/or the set of rules further comprises one or more rules for adding tags to the tokens; and

the method further comprises adding tags to the tokens based on the statistical models and/or the rules.

14. The method of claim 13 , wherein the tags are based on semantic information and/or structural information.

15. The method of claim 1 , wherein the set of rules further comprises one or more rules that identify markup language content, an Internet address, a hashtag, or an emoji/emoticon.

16. The method of claim 1 , wherein the set of statistical models and/or the set of rules are adjusted based at least in part on an author of the document.

17. The method of claim 1 , wherein the set of statistical models and/or the set of rules are based at least in part on intra-document information.

18. An apparatus for tokenizing text for natural language processing, the apparatus comprising one or more processors configured to:

generate from a pool of documents a set of statistical models comprising one or more entries each indicating a likelihood of appearance of a character/letter sequence in the pool of documents;

receive a set of rules comprising rules that identify character/letter sequences as valid tokens;

transform one or more entries in the statistical models into new rules that are added to the set of rules when the entries indicate a high likelihood;

receive a document to be processed;

divide the document to be processed into tokens based on the set of statistical models and the set of rules; and

output the tokens for natural language processing.

19. The apparatus of claim 18 , wherein the set of statistical models further comprises statistical models based on human annotation, and the one or more processors are further configured to:

generate one or more human readable prompts configured to elicit annotations of one or more documents in the pool of documents, wherein the annotations comprise identification of one or more character/letter sequences in the documents as valid tokens;

receive one or more annotations elicited by the human readable prompts; and

generate statistical models based on the received annotations, wherein the statistical models comprise one or more entries each indicating a likelihood of appearance of a character/letter sequence in the character/letter sequences annotated as valid tokens.

20. A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to:

generate from a pool of documents a set of statistical models comprising one or more entries each indicating a likelihood of appearance of a character/letter sequence in the pool of documents;

receive a set of rules comprising rules that identify character/letter sequences as valid tokens;

transform one or more entries in the statistical models into new rules that are added to the set of rules when the entries indicate a high likelihood;

receive a document to be processed;

divide the document to be processed into tokens based on the set of statistical models and the set of rules; and

output the tokens for natural language processing.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2024
From: 100.CO GLOBAL HOLDINGS, LLC
To: AI IP INVESTMENTS LTD.
Reel/Frame 066636/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2023
From: DAASH INTELLIGENCE, INC.
To: 100.CO GLOBAL HOLDINGS, LLC
Reel/Frame 064420/0108 →
CHANGE OF NAME Recorded Mar 7, 2023
From: 100.CO TECHNOLOGIES, INC.
To: DAASH INTELLIGENCE, INC.
Reel/Frame 062992/0333 →
NUNC PRO TUNC ASSIGNMENT Recorded Dec 16, 2022
From: 100.CO, LLC
To: 100.CO TECHNOLOGIES, INC.
Reel/Frame 062131/0714 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2021
From: AI IP INVESTMENTS LTD.
To: 100.CO, LLC
Reel/Frame 056145/0509 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2021
From: AIPARC HOLDINGS PTE. LTD.
To: AI IP INVESTMENTS LTD
Reel/Frame 056096/0278 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2021
From: HEALY, TREVOR
To: AIPARC HOLDINGS PTE. LTD.
Reel/Frame 056083/0123 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2021
From: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
To: HEALY, TREVOR
Reel/Frame 056057/0325 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2021
From: IDIBON, INC.
To: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
Reel/Frame 055978/0362 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2021
From: MUNRO, ROBERT J.; VOIGT, ROB; ERLE, SCHUYLER D.; CALLAHAN, BRENDAN D.; KING, GARY C.; LONG, JESSICA D.; BRENIER, JASON; SAXENA, TRIPTI; KRAWCZYK, STEFAN
To: IDIBON, INC.
Reel/Frame 055947/0571 →
Continuity (10)
Continuation 16056263 · Aug 6, 2018
Continuation 15596855 · May 16, 2017
Continuation 14964512 · Dec 9, 2015
Provisional Application 62254090 · Nov 11, 2015
Provisional Application 62254095 · Nov 11, 2015
Provisional Application 62089736 · Dec 9, 2014
Provisional Application 62089747 · Dec 9, 2014
Provisional Application 62089745 · Dec 9, 2014
Provisional Application 62089742 · Dec 9, 2014
Related Publication 20210157984A1 · May 27, 2021