IP Library Patent Application 14964528
Patent Application
App. No. 14/964,528

TECHNIQUES FOR COMBINING HUMAN AND MACHINE LEARNING IN NATURAL LANGUAGE PROCESSING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/964,528
Abstract

Methods, apparatuses and computer readable medium are presented for generating a natural language model. A method for generating a natural language model comprises: receiving more than one annotation of a document; calculating a level of agreement among the received annotations; determining that a criterion among a first criterion, a second criterion, and a third criterion is satisfied based at least in part on the level of agreement; determining an aggregated annotation representing an aggregation of information in the received annotations and training a natural language model using the aggregated annotation, when the first criterion is satisfied; generating at least one human readable prompt configured to receive additional annotations of the document, when the second criterion is satisfied; and discarding the received annotations from use in training the natural language model, when the third criterion is satisfied.

Claims (47)

1 . A method for generating a natural language model, the method comprising:

receiving more than one annotation of a document;

calculating a level of agreement among the received annotations;

determining that a criterion among a first criterion, a second criterion, and a third criterion is satisfied based at least in part on the level of agreement;

determining an aggregated annotation representing an aggregation of information in the received annotations and training a natural language model using the aggregated annotation, when the first criterion is satisfied;

generating at least one human readable prompt configured to receive additional annotations of the document, when the second criterion is satisfied; and

discarding the received annotations from use in training the natural language model, when the third criterion is satisfied.

2 . The method of claim 1 , wherein the second criterion is satisfied when the number of annotations received is less than a minimum number.

3 . The method of claim 1 , wherein the annotations of the document comprise selection of one or more portions of the document relevant to one or more topics.

4 . The method of claim 1 , wherein the annotations of the document comprise selection of one or more categories among a plurality of categories.

5 . The method of claim 4 , wherein the level of agreement is determined for each category based on a percentage of annotations that select said category.

6 . The method of claim 5 , wherein:

the first criterion is satisfied when the number of annotations received is at least a minimum number and the level of agreement for a category is at least a threshold level; and

the aggregated annotation is determined as selecting or not selecting said category.

7 . The method of claim 5 , wherein the second criterion is satisfied when the number of annotations received is less than a maximum number and the level of agreement is less than a threshold level.

8 . The method of claim 5 , wherein the third criterion is satisfied when the number of annotations received is at least a maximum number and the level of agreement is less than a threshold level.

9 . The method of claim 4 , wherein a numerical value is assigned to each of the plurality of categories.

10 . The method of claim 9 , wherein:

the level of agreement comprises a difference between the highest numerical value and the lowest numerical value among the selected categories;

the first criterion is satisfied when the difference is no more than a threshold value; and

the third criterion is satisfied when the difference is more than the threshold value.

11 . The method of claim 10 , wherein the aggregated annotation is determined as selection of a category with the numerical value closest to a mean of the numerical values of all received annotations.

12 . The method of claim 10 , wherein the aggregated annotation is determined as selection of a category with the numerical value closest to a median of the numerical values of all received annotations.

13 . The method of claim 1 , wherein determining that the criterion among the first criterion, the second criterion, and the third criterion is satisfied is further based on a result of an analysis of the document by one or more pre-existing natural language models.

14 . The method of claim 1 , wherein determining that the criterion among the first criterion, the second criterion, and the third criterion is satisfied is further based on known performance levels of annotators.

15 . The method of claim 1 , wherein at least one of the annotations received comprises prediction by a pre-existing natural language model.

16 . An apparatus for generating a natural language model, the apparatus comprising one or more processors configured to:

receive more than one annotation of a document;

calculate a level of agreement among the received annotations;

determine that a criterion among a first criterion, a second criterion, and a third criterion is satisfied based at least in part on the level of agreement;

determine an aggregated annotation representing an aggregation of information in the received annotations and train a natural language model using the aggregated annotation, when the first criterion is satisfied;

generate at least one human readable prompt configured to receive additional annotations of the document, when a second criterion is satisfied; and

discard the received annotations from use in training the natural language model, when the third criterion is satisfied.

17 . The apparatus of claim 16 , wherein the annotations of the document comprise selection of one or more categories among a plurality of categories.

18 . The apparatus of claim 17 , wherein the level of agreement is determined for each category based on a percentage of annotations that select said category.

19 . The apparatus of claim 17 , wherein

a numerical value is assigned to each of the plurality of categories;

the level of agreement comprises a difference between the highest numerical value and the lowest numerical value among the selected categories;

the first criterion is satisfied when the difference is no more than a threshold value; and

the third criterion is satisfied when the difference is more than a threshold value.

20 . A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to:

receive more than one annotation of a document;

calculate a level of agreement among the received annotations;

determine that a criterion among a first criterion, a second criterion, and a third criterion is satisfied based at least in part on the level of agreement;

determine an aggregated annotation representing an aggregation of information in the received annotations and train a natural language model using the aggregated annotation, when the first criterion is satisfied;

generate at least one human readable prompt configured to receive additional annotations of the document, when a second criterion is satisfied; and

discard the received annotations from use in training the natural language model, when the third criterion is satisfied.

Assignments (12)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2024
From: 100.CO GLOBAL HOLDINGS, LLC
To: AI IP INVESTMENTS LTD.
Reel/Frame 066636/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2023
From: DAASH INTELLIGENCE, INC.
To: 100.CO GLOBAL HOLDINGS, LLC
Reel/Frame 064420/0108 →
CHANGE OF NAME Recorded Jul 19, 2023
From: 100.CO TECHNOLOGIES, INC.
To: DAASH INTELLIGENCE, INC.
Reel/Frame 064347/0117 →
NUNC PRO TUNC ASSIGNMENT Recorded Dec 16, 2022
From: 100.CO, LLC
To: 100.CO TECHNOLOGIES, INC.
Reel/Frame 062131/0714 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE CITY PREVIOUSLY RECORDED AT REEL: 055929 FRAME: 0975. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 5, 2021
From: AI IP INVESTMENTS LTD.
To: 100.CO, LLC
Reel/Frame 056151/0150 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2021
From: AI IP INVESTMENTS LTD.
To: 100.CO, LLC
Reel/Frame 055929/0975 →
CORRECTIVE ASSIGNMENT TO CORRECT THE COVENANT INFORMATION TO BE UPDATED FROM AIRPARC HOLDING PTE. LTD. AND REPLACED WITH TREVOR HEALY (SEE MARKED ASSIGNMENT) PREVIOUSLY RECORDED ON REEL 047110 FRAME 0510. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 24, 2021
From: HEALY, TREVOR
To: AIPARC HOLDINGS PTE. LTD.
Reel/Frame 055404/0561 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: AIPARC HOLDINGS PTE. LTD.
To: AI IP INVESTMENTS LTD
Reel/Frame 055377/0995 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2018
From: IDIBON, INC.
To: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
Reel/Frame 047110/0178 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2018
From: IDIBON (ASSIGNMENT FOR THE BENEFIT OF CREDITORS), LLC
To: HEALY, TREVOR
Reel/Frame 047110/0449 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2018
From: HEALY, TREVOR
To: AIPARC HOLDINGS PTE. LTD.
Reel/Frame 047110/0510 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2016
From: MUNRO, ROBERT J.; WALKER, CHRISTOPHER; LUGER, SARAH K.; CALLAHAN, BRENDAN D.; KING, GARY C.; TEPPER, PAUL A.; THOMPSON, JANA N.; SCHNOEBELEN, TYLER J.; BRENIER, JASON; LONG, JESSICA D.
To: IDIBON, INC.
Reel/Frame 038651/0408 →