IP Library Patent Application 14964520
Patent Application
App. No. 14/964,520

OPTIMIZATION TECHNIQUES FOR ARTIFICIAL INTELLIGENCE

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/964,520
Abstract

Methods, apparatuses and computer readable medium are presented for generating a natural language model. A method for generating a natural language model comprises: selecting from a pool of documents, a first set of documents to be annotated; receiving annotations of the first set of documents elicited by first human readable prompts; training a natural language model using the annotated first set of documents; determining documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations; selecting from the pool of documents, a second set of documents to be annotated comprising documents having uncertain natural language processing results; receiving annotations of the second set of documents elicited by second human readable prompts; and retraining a natural language model using the annotated second set of documents.

Claims (45)

1 . A method for generating a natural language model, the method comprising:

selecting by one or more processors in a natural language platform, from a pool of documents, a first set of documents to be annotated;

for each document in the first set of documents, generating, by the one or more processors, a first human readable prompt configured to elicit an annotation of said document;

receiving annotations of the first set of documents elicited by the first human readable prompts;

training, by the one or more processors, a natural language model using the annotated first set of documents;

determining, by the one or more processors, documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations;

selecting by the one or more processors, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results;

for each document in the second set of documents, generating, by the one or more processors, a second human readable prompt configured to elicit an annotation of said document;

receiving annotations of the second set of documents elicited by the second human readable prompts; and

retraining, by the one or more processors, a natural language model using the annotated second set of documents.

2 . The method of claim 1 , wherein the annotations of the first and second sets of documents comprise classification of the documents into one or more categories among a plurality of categories.

3 . The method of claim 1 , wherein the annotations of the first and second sets of documents comprise selection of one or more portions of the documents relevant to one or more topics.

4 . The method of claim 1 , wherein the steps of determining documents having uncertain natural language processing results; selecting a second set of documents to be annotated; generating a second human readable prompt; receiving annotations of the second set of documents; and retraining a natural language model using the annotated second set of documents, are repeated until the trained model has reached a predetermined performance level.

5 . The method of claim 1 , wherein selecting the first set of documents comprises selecting documents that are evenly distributed among different document types.

6 . The method of claim 1 , wherein selecting the first set of documents comprises selecting at least one document within each of a plurality of machine-discovered topics.

7 . The method of claim 1 , wherein selecting the first set of documents comprises selecting documents based on a keyword search.

8 . The method of claim 1 , wherein selecting the first set of documents comprises selecting documents based on confidence levels generated by analysis of the documents by one or more pre-existing natural language models.

9 . The method of claim 1 , wherein selecting the first set of documents comprises a manual selection.

10 . The method of claim 1 , wherein selecting the first set of documents comprises removing exact duplicates and/or near duplicates from the first set of documents.

11 . The method of claim 1 , wherein selecting the first set of documents comprises selecting documents based on features contained therein.

12 . The method of claim 1 , wherein determining documents having uncertain natural language processing results is based on confidence levels generated by analysis of the documents by the trained model.

13 . The method of claim 1 , further comprising training, by the one or more processors, a plurality of additional natural language models using the annotated first set of documents, wherein determining documents having uncertain natural language processing results is based on a level of disagreement among the plurality of additional natural language models.

14 . The method of claim 13 , wherein the level of disagreement is determined by assigning more weight to models with better known performance levels than models with worse known performance levels.

15 . The method of claim 1 , wherein determining documents having uncertain natural language processing results is based on a level of disagreement among more than one annotator.

16 . The method of claim 15 , wherein the level of disagreement is determined by assigning more weight to annotators with better known performance levels than annotators with worse known performance levels.

17 . The method of claim 15 , wherein selecting the second set of documents comprises selecting documents similar to documents that have a high level of disagreement among more than one annotator.

18 . The method of claim 1 , wherein the second human readable prompt is configured to elicit a true-or-false answer aimed at resolving uncertainty in the natural language processing results.

19 . An apparatus for generating a natural language model, the apparatus comprising one or more processors configured to:

select, from a pool of documents, a first set of documents to be annotated;

for each document in the first set of documents, generate a first human readable prompt configured to elicit an annotation of said document;

receive annotations of the first set of documents elicited by the first human readable prompts;

train a natural language model using the annotated first set of documents;

determine documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations;

select, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results;

for each document in the second set of documents, generate a second human readable prompt configured to elicit an annotation of said document;

receive annotations of the second set of documents elicited by the second human readable prompts; and

retrain a natural language model using the annotated second set of documents.

20 . A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to:

select, from a pool of documents, a first set of documents to be annotated;

for each document in the first set of documents, generate a first human readable prompt configured to receive annotations of the first set of documents;

train a natural language model using the annotated first set of documents;

determine documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations;

select, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results;

for each document in the second set of documents, generate a second human readable prompt configured to receive annotations of the second set of documents; and

retrain a natural language model using the annotated second set of documents.

Assignments (12)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2024
From: 100.CO GLOBAL HOLDINGS, LLC
To: AI IP INVESTMENTS LTD.
Reel/Frame 066636/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2023
From: DAASH INTELLIGENCE, INC.
To: 100.CO GLOBAL HOLDINGS, LLC
Reel/Frame 064420/0108 →
CHANGE OF NAME Recorded Jul 19, 2023
From: 100.CO TECHNOLOGIES, INC.
To: DAASH INTELLIGENCE, INC.
Reel/Frame 064347/0117 →
NUNC PRO TUNC ASSIGNMENT Recorded Dec 16, 2022
From: 100.CO, LLC
To: 100.CO TECHNOLOGIES, INC.
Reel/Frame 062131/0714 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE CITY PREVIOUSLY RECORDED AT REEL: 055929 FRAME: 0975. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 5, 2021
From: AI IP INVESTMENTS LTD.
To: 100.CO, LLC
Reel/Frame 056151/0150 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2021
From: AI IP INVESTMENTS LTD.
To: 100.CO, LLC
Reel/Frame 055929/0975 →
CORRECTIVE ASSIGNMENT TO CORRECT THE COVENANT INFORMATION TO BE UPDATED FROM AIRPARC HOLDING PTE. LTD. AND REPLACED WITH TREVOR HEALY (SEE MARKED ASSIGNMENT) PREVIOUSLY RECORDED ON REEL 047110 FRAME 0510. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 24, 2021
From: HEALY, TREVOR
To: AIPARC HOLDINGS PTE. LTD.
Reel/Frame 055404/0561 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: AIPARC HOLDINGS PTE. LTD.
To: AI IP INVESTMENTS LTD
Reel/Frame 055377/0995 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2018
From: IDIBON, INC.
To: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
Reel/Frame 047110/0178 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2018
From: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
To: HEALY, TREVOR
Reel/Frame 047110/0449 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2018
From: HEALY, TREVOR
To: AIPARC HOLDINGS PTE. LTD.
Reel/Frame 047110/0510 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2016
From: MUNRO, ROBERT J.; ERLE, SCHUYLER D.; BRENIER, JASON; TEPPER, PAUL A.; SAXENA, TRIPTI; KING, GARY C.; LONG, JESSICA D.; CALLAHAN, BRENDAN D.; SCHNOEBELAN, TYLER J.; KRAWCZYK, STEFAN; BASAVARAJ, VEENA
To: IDIBON, INC.
Reel/Frame 038618/0726 →