IP Library › Granted Patent US 12,555,003
Granted Patent B2
US 12,555,003 · App. 17/112,396 · Granted Feb 17, 2026

Self improving annotation quality on diverse and multi-disciplinary content

Inventors: Pierpaolo Tommasi (Dublin, IE); Charles Arthur Jochim (Dublin, IE); Joao H Bettencourt-Silva (Dublin, IE); Alessandra Pascale (Phoenix Park Racecourse, IE)
Assignee: International Business Machines Corporation
G06N5/04G06N20/00G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,003
App. No.
17/112,396
Granted
Feb 17, 2026
Kind
B2
Abstract

Content can be dynamically associated to an annotator based on machine learning of annotator expertise and dynamically maintaining a set of labels and relationships among the labels. A probable group of labels associated with the content can be determined by a first machine learning model. Using a set of labels and relationships among the set of labels maintained dynamically, a second machine learning model can select an annotator having subject matter expertise associated with the probable group of labels. The first machine learning model and the second machine learning model can be retrained based on annotations performed on the content by the annotator as feedback. A third machine learning model can dynamically maintain the set of labels and the relationships among the set of labels.

Claims (39)

1 . A system comprising:

a processor;

a memory device coupled with the processor;

the processor configured to dynamically and automatically associate content to an annotator based on machine learning of annotator expertise and dynamically maintain a set of labels and relationships among the labels;

a first machine learning model trained to output a probable set of labels associated with the content;

a second machine learning model configured to determine an annotator having expertise in an area associated with the probable set of labels, the second machine learning model configured to learn annotator expertise at least based on annotator behavior, the annotator behavior including at least an amount of time a given annotator takes in labeling a given content, the second machine learning model further configured to discover and learn that the annotator has expertise in other fields than a field associated with the probable set of labels, and update dynamically the annotator's area of expertise; and

a third machine learning model configured to maintain the set of labels and relationships among the labels, and further configured to learn new relationships among the labels and new labels responsive to receiving annotator tagged new labels associated the content,

wherein the processor is configured to run the first machine learning model using the content as input to determine the probable set of labels associated with the content, automatically run the second machine learning model using the probable set of labels as input to automatically determine the annotator, automatically queue the content in a queue on the memory device associated with the annotator, run the third machine learning model and fetch from the third machine learning model additional labels associated with the probable set of labels and also associated with a field of expertise of the annotator, and provide the additional labels with the content to the queue associated with the annotator, the content that is labeled by the annotator being used as part of a training set for supervised machine learning, the content that is labeled by the annotator being usable for multiple domains without having to repeat a labeling process of the content for every domain of the multiple domains by using, for the supervised machine learning in different domains of the multiple domains, aggregated labels having hierarchical relationships with annotator provided labels of the content.

2 . The system of claim 1 , wherein the content includes at least one of text, images, audio, video and media data.

3 . The system of claim 1 , wherein the first machine learning model trains and uses annotations performed by the annotator as feedback to retrain itself.

4 . The system of claim 1 , wherein the second machine learning model learns annotator expertise based on at least one of: collected annotations of the annotator with associated metadata, collected annotations from another annotator and an expertise of a similar annotator determined to be within a threshold similarity with the annotator.

5 . The system of claim 4 , wherein the second machine learning model learns the annotator expertise based on a consensus validating the annotator's performance on the content.

6 . The system of claim 4 , wherein the second machine learning model learns the annotator expertise based on scoring the annotator's performance on the content.

7 . The system of claim 4 , wherein the second machine learning model learns the annotator expertise using ontology associated with the set of labels and the relationships among the set of labels maintained by the third machine learning model.

8 . The system of claim 1 , wherein the third machine learning model learns new relationships among labels and new user-provided labels.

9 . The system of claim 1 , wherein the content includes unlabeled content.

10 . The system of claim 1 , wherein the content includes partially labeled content.

11 . The system of claim 1 , wherein the processor is further configured to collect annotations from the annotator.

12 . A computer-implemented method comprising:

determining, by at least one processor, a probable group of labels associated with content by running a first machine learning model using the content as input;

using a set of labels and relationships among the set of labels maintained dynamically, selecting automatically, by the at least one processor, by running a second machine learning model based on the probable group of labels automatically determined by the first machine learning model and received as input, an annotator having subject matter expertise associated with the probable group of labels, the second machine learning model configured to learn annotator expertise at least based on annotator behavior, the annotator behavior including at least an amount of time a given annotator takes in labeling a given content, the second machine learning model further configured to discover and learn that the annotator has expertise in other fields than a field associated with the probable set of labels, and update dynamically the annotator's area of expertise;

dynamically maintaining, by the at least one processor, by running a third machine learning model the set of labels and the relationships among the set of labels, the third machine learning model further configured to learn new relationships among the labels and new labels responsive to receiving annotator tagged new labels associated the content;

automatically queueing, by the at least one processor, the content in an in-memory queue associated with the annotator selected automatically by running the second machine learning model;

fetching, by the at least one processor, from the third machine learning model additional labels associated with the probable group of labels and also associated with a field of expertise of the annotator;

providing, by the at least one processor, the additional labels with the content to the in-memory queue associated with the annotator; and

retraining, by the at least one processor, the first machine learning model and the second machine learning model based on annotations performed on the content by the annotator as feedback,

wherein at least some of the set of labels are provided to the annotator to label the content with the at least some of the set of labels,

the content that is labeled by the annotator being used as part of a training set for supervised machine learning, the content that is labeled by the annotator being usable for multiple domains without having to repeat a labeling process of the content for every domain of the multiple domains by using, for the supervised machine learning in different domains of the multiple domains, aggregated labels having hierarchical relationships with annotator provided labels of the content.

13 . The method of claim 12 , wherein the first machine learning model is trained to output labels for given content.

14 . The method of claim 12 , wherein the second machine learning model is trained to associate labels to annotators with a given confidence.

15 . The method of claim 12 , wherein the second machine learning model is trained to associate labels to annotators with a given confidence based on at least one of consensus validating the annotations performed on the content by the annotator, scoring of reliability associated with the annotator based on the annotations performed on the content by the annotator, and ontology associated with the set of labels and the relationships among the set of labels maintained dynamically.

16 . The method of claim 12 , wherein the content includes at least one of digital data, text, images, audio, video and media data.

17 . A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable by a device to cause the device to perform operations comprising:

dynamically and automatically associate content to an annotator based on machine learning of annotator expertise and dynamically maintain a set of labels and relationships among the labels, at least some of the set of labels being provided to the annotator to label the content with the at least some of the set of labels, the device further caused to run a first machine learning model using the content as input to determine a probable set of labels associated with the content;

automatically run a second machine learning model using the probable set of labels as input to automatically determine the annotator, the second machine learning model configured to learn annotator expertise at least based on annotator behavior, the annotator behavior including at least an amount of time a given annotator takes in labeling a given content, the second machine learning model further configured to discover and learn that the annotator has expertise in other fields than a field associated with the probable set of labels, and update dynamically the annotator's area of expertise;

automatically queue the content in an in-memory queue associated with the annotator, run a third machine learning model configured to maintain the set of labels and relationships among the labels, the third machine learning model further configured to learn new relationships among the labels and new labels responsive to receiving annotator tagged new labels associated the content;

fetch from the third machine learning model additional labels associated with the probable set of labels and also associated with a field of expertise of the annotator; and

provide the additional labels with the content to the queue associated with the annotator,

the content that is labeled by the annotator being used as part of a training set for supervised machine learning, the content that is labeled by the annotator being usable for multiple domains without having to repeat a labeling process of the content for every domain of the multiple domains by using, for the supervised machine learning in different domains of the multiple domains, aggregated labels having hierarchical relationships with annotator provided labels of the content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2020
From: TOMMASI, PIERPAOLO; JOCHIM, CHARLES ARTHUR; BETTENCOURT-SILVA, JOAO H; PASCALE, ALESSANDRA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054550/0817 →
Continuity (1)
Related Publication 20220180224A1 · Jun 9, 2022
References Cited (30)
US 9311599B1 · Attenberg et al. · 2016 [cited by applicant]
US 9928278B2 · Welinder et al. · 2018 [cited by applicant]
US 10496369B2 · Guttmann · 2019 [cited by applicant]
US 10977518B1 · Sharma · 2021 [cited by examiner]
US 11263391B2 · Potts · 2022 [cited by examiner]
US 11893772B1 · Gokalp · 2024 [cited by examiner]
US 20050027664A1 · Johnson et al. · 2005 [cited by applicant]
US 20110071967A1 · Fung · 2011 [cited by examiner]
US 20130346356A1 · Welinder et al. · 2013 [cited by applicant]
US 20140344191A1 · Lebow · 2014 [cited by applicant]
US 20180060307A1 · Misra · 2018 [cited by examiner]
US 20190050428A1 · Hou · 2019 [cited by examiner]
US 20190384807A1 · Dernoncourt et al. · 2019 [cited by applicant]
US 20200152316A1 · Lee et al. · 2020 [cited by applicant]
US 20200160231A1 · Asthana · 2020 [cited by examiner]
US 20200372338A1 · Woods, Jr. · 2020 [cited by examiner]
US 20220019729A1 · Sharma · 2022 [cited by examiner]
KR 102081037 · 2020 [cited by applicant]
Synced, “Data Annotation: The Billion Dollar Business Behind AI Breakthroughs”, https://medium.com/syncedreview/data-annotation-the-billion-dollar-business-behind-ai-breakthroughs-d929b0a50d23, Aug. 28, 2019, Printed on… [cited by applicant]
Hovy, D., et al., “Learning Whom to Trust with MACE”, Proceedings of NAACL-HLT 2013, Jun. 9-14, 2013, pp. 1120-1130. [cited by applicant]
Dipper, S., et al., “Towards User-Adaptive Annotation Guidelines”, https://www.aclweb.org/anthology/W04-1904.pdf, Printed on Aug. 14, 2020, 7 pages. [cited by applicant]
Appen, “Figure Eight”, https://appen.com/tag/figure-eight/, Printed on Aug. 17, 2020, 8 pages. [cited by applicant]
Amazon Mechanical Turk, Inc., Amazon Mechanical Turk, https://www.mturk.com/, Printed on Nov. 20, 2020, 5 pages. [cited by applicant]
Amazon Mechanical Turk, Inc., “Mechanical Turk—Features”, https://www.mturk.com/product-details, Printed on Nov. 20, 2020, 3 pages. [cited by applicant]
Appen, “Annotation Capabilities”, https://appen.com/solutions/annotation-capabilities/, Printed on Nov. 20, 2020, 9 pages. [cited by applicant]
Wikipedia, “Figure Eight Inc.”, https://en.wikipedia.org/wiki/Figure_Eight_Inc., Last edited on Oct. 24, 2020, Printed on Nov. 20, 2020, 5 pages. [cited by applicant]
Amazon Web Services (AWS), “Figure Eight Data Labeling Platform”, https://aws.amazon.com/solutionspace/financial-services/solutions/figure-eight/, Printed on Nov. 20, 2020, 6 pages. [cited by applicant]
NIST, “Nist Cloud Computing Program”, http://csrc.nist.gov/groups/SNS/cloud-computing/index.html, Created Dec. 1, 2016, Updated Oct. 6, 2017, 9 pages. [cited by applicant]
Youtube, Screen capture from YouTube video clip entitled “Mturk easy work”, https://www.youtube.com/watch?v=UumoaPFPyMg, uploaded on Mar. 7, 2019 by user “Mturk Worker”, 1 page. [cited by applicant]
Youtube, Screen capture from YouTube video clip entitled “Amazon MTurk Made Easy,” https://www.youtube.com/watch?v=CwT0JTxy2YU, uploaded on Apr. 21, 2016 by user “Fun and Budget with Tinesha Davis”, 1 page. [cited by applicant]