IP Library Granted Patent US 11,288,444
Granted Patent B2
US 11,288,444 · App. 17/119,902 · Granted Mar 29, 2022

Optimization techniques for artificial intelligence

Inventors: Robert J. Munro (San Francisco, CA); Schuyler D. Erle (San Francisco, CA); Jason Brenier (Oakland, CA); Paul A. Tepper (San Francisco, CA); Tripti Saxena (Cupertino, CA); Gary C. King (Los Altos, CA); Jessica D. Long (San Francisco, CA); Brendan D. Callahan (Philadelphia, PA); Tyler J. Schnoebelen (San Francisco, CA); Stefan Krawczyk (Menlo Park, CA); Veena Basavaraj (San Francisco, CA)
Assignee: 100.co, LLC
G06F40/169G06F3/0482G06F16/243G06F16/24532G06F16/285G06F16/288G06F16/3329G06F16/35G06F16/367G06F16/93G06F16/951G06F40/137G06F40/221G06F40/30G06F40/40G06F40/42G06N20/00G06Q50/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,288,444
App. No.
17/119,902
Granted
Mar 29, 2022
Kind
B2
Abstract

Methods, apparatuses and computer readable medium are presented for generating a natural language model. A method for generating a natural language model comprises: selecting from a pool of documents, a first set of documents to be annotated; receiving annotations of the first set of documents elicited by first human readable prompts; training a natural language model using the annotated first set of documents; determining documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations; selecting from the pool of documents, a second set of documents to be annotated comprising documents having uncertain natural language processing results; receiving annotations of the second set of documents elicited by second human readable prompts; and retraining a natural language model using the annotated second set of documents.

Claims (45)

1. A method for generating a natural language model, the method comprising:

selecting by one or more processors in a natural language platform, from a pool of documents, a first set of documents to be annotated;

for each document in the first set of documents, generating, by the one or more processors, a first human readable prompt configured to elicit an annotation of said document;

receiving annotations of the first set of documents by a plurality of annotators, the annotations being elicited by the first human readable prompts;

training, by the one or more processors, a natural language model using the annotated first set of documents;

determining, by the one or more processors, documents in the pool having uncertain natural language processing results based on a level of disagreement among more than one annotator of the plurality of annotators;

selecting by the one or more processors, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results;

for each document in the second set of documents, generating, by the one or more processors, a second human readable prompt configured to elicit an annotation of said document;

receiving annotations of the second set of documents elicited by the second human readable prompts; and

retraining, by the one or more processors, a natural language model using the annotated second set of documents.

2. The method of claim 1 , wherein the annotations of the first and second sets of documents comprise classification of the documents into one or more categories among a plurality of categories.

3. The method of claim 1 , wherein the annotations of the first and second sets of documents comprise selection of one or more portions of the documents relevant to one or more topics.

4. The method of claim 1 , wherein the steps of determining documents having uncertain natural language processing results; selecting a second set of documents to be annotated; generating a second human readable prompt; receiving annotations of the second set of documents; and retraining a natural language model using the annotated second set of documents, are repeated until the trained model has reached a predetermined performance level.

5. The method of claim 1 , wherein selecting the first set of documents comprises selecting documents that are evenly distributed among different document types.

6. The method of claim 1 , wherein selecting the first set of documents comprises selecting at least one document within each of a plurality of machine-discovered topics.

7. The method of claim 1 , wherein selecting the first set of documents comprises selecting documents based on a keyword search.

8. The method of claim 1 , wherein selecting the first set of documents comprises selecting documents based on confidence levels generated by analysis of the documents by one or more pre-existing natural language models.

9. The method of claim 1 , wherein selecting the first set of documents comprises a manual selection.

10. The method of claim 1 , wherein selecting the first set of documents comprises removing exact duplicates and/or near duplicates from the first set of documents.

11. The method of claim 1 , wherein selecting the first set of documents comprises selecting documents based on features contained therein.

12. The method of claim 1 , wherein determining documents having uncertain natural language processing results is further based on confidence levels generated by analysis of the documents by the trained model.

13. The method of claim 1 , further comprising training, by the one or more processors, a plurality of additional natural language models using the annotated first set of documents, wherein determining documents having uncertain natural language processing results is further based on a level of disagreement among the plurality of additional natural language models.

14. The method of claim 13 , wherein the level of disagreement among the plurality of additional natural language models is determined by assigning more weight to models with better known performance levels than models with worse known performance levels.

15. The method of claim 1 , wherein the level of disagreement among the more than one annotator is determined by assigning more weight to annotators with better known performance levels than annotators with worse known performance levels.

16. The method of claim 1 , wherein selecting the second set of documents comprises selecting documents similar to documents that have a high level of disagreement among the more than one annotator.

17. The method of claim 1 , wherein the second human readable prompt is configured to elicit a true-or-false answer aimed at resolving uncertainty in the natural language processing results.

18. The method of claim 16 , wherein the selection of documents similar to documents that have a high level of disagreement is based on text or metadata similarity.

19. An apparatus for generating a natural language model, the apparatus comprising one or more processors configured to:

select, from a pool of documents, a first set of documents to be annotated;

for each document in the first set of documents, generate a first human readable prompt configured to elicit an annotation of said document;

receive annotations of the first set of documents by a plurality of annotators, the annotators being elicited by the first human readable prompts;

train a natural language model using the annotated first set of documents;

determine documents in the pool having uncertain natural language processing results based on a level of disagreement among more than one annotator of the plurality of annotators;

select, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results;

for each document in the second set of documents, generate a second human readable prompt configured to elicit an annotation of said document;

receive annotations of the second set of documents elicited by the second human readable prompts; and

retrain a natural language model using the annotated second set of documents.

20. A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to:

select, from a pool of documents, a first set of documents to be annotated;

for each document in the first set of documents, generate a first human readable prompt configured to receive annotations of the first set of documents by a plurality of annotators;

train a natural language model using the annotated first set of documents;

determine documents in the pool having uncertain natural language processing results based on a level of disagreement among more than one annotator of the plurality of annotators;

select, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results;

for each document in the second set of documents, generate a second human readable prompt configured to receive annotations of the second set of documents; and

retrain a natural language model using the annotated second set of documents.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2024
From: 100.CO GLOBAL HOLDINGS, LLC
To: AI IP INVESTMENTS LTD.
Reel/Frame 066636/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2023
From: DAASH INTELLIGENCE, INC.
To: 100.CO GLOBAL HOLDINGS, LLC
Reel/Frame 064420/0108 →
CHANGE OF NAME Recorded Mar 7, 2023
From: 100.CO TECHNOLOGIES, INC.
To: DAASH INTELLIGENCE, INC.
Reel/Frame 062992/0333 →
NUNC PRO TUNC ASSIGNMENT Recorded Dec 16, 2022
From: 100.CO, LLC
To: 100.CO TECHNOLOGIES, INC.
Reel/Frame 062131/0714 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2021
From: AI IP INVESTMENTS LTD.
To: 100.CO, LLC
Reel/Frame 056145/0509 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2021
From: AIPARC HOLDINGS PTE. LTD.
To: AI IP INVESTMENTS LTD
Reel/Frame 056096/0278 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2021
From: HEALY, TREVOR
To: AIPARC HOLDINGS PTE. LTD.
Reel/Frame 056083/0123 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2021
From: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
To: HEALY, TREVOR
Reel/Frame 056057/0325 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2021
From: IDIBON, INC.
To: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
Reel/Frame 055978/0362 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2021
From: MUNRO, ROBERT J.; ERLE, SCHUYLER D.; BRENIER, JASON; TEPPER, PAUL A.; SAXENA, TRIPTI; KING, GARY C.; LONG, JESSICA D.; CALLAHAN, BRENDAN D.; SCHNOEBELEN, TYLER J.; KRAWCZYK, STEFAN; BASAVARAJ, VEENA
To: IDIBON, INC.
Reel/Frame 055945/0838 →
Continuity (7)
Continuation 16198453 · Nov 21, 2018
Continuation 14964520 · Dec 9, 2015
Provisional Application 62089745 · Dec 9, 2014
Provisional Application 62089747 · Dec 9, 2014
Provisional Application 62089736 · Dec 9, 2014
Provisional Application 62089742 · Dec 9, 2014
Related Publication 20210232760A1 · Jul 29, 2021
Cited By (2)
US 12,229,522 US 12,675,512