IP Library Granted Patent US 12,373,690
Granted Patent B2
US 12,373,690 · App. 18/183,111 · Granted Jul 29, 2025

Targeted crowd sourcing for metadata management across data sets

Inventors: Robert Woods, Jr. (Plano, TX); Mark D. Austin (Allen, TX)
Assignee: AT&T Intellectual Property I, L.P.
G06N3/08G06F16/2365G06Q10/063112
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,690
App. No.
18/183,111
Granted
Jul 29, 2025
Kind
B2
Abstract

A system includes: a memory operable to store a predictive model; a first processor communicatively coupled to the memory, the first processor operable to execute the predictive model to perform operations including generating knowledge score metrics based on a set of attributes for individuals included in a specified population, where the knowledge score metrics quantify a prediction of a capability of an individual for performing metadata labeling; a second processor communicatively coupled to the memory and the first processor, the second processor is operable to perform operations including comparing the knowledge score metrics to a specified threshold, and identifying attributes of individuals from a specified population having knowledge score metrics exceeding the specified threshold as attributes of individuals capable of performing metadata labeling.

Claims (60)

1. A method comprising:

forming a machine learning model operable on a neural network processor in a computing environment, the machine learning model configured to analyze input data comprising an initial set of attributes to predict analytical capabilities of a set of individuals for performing a specific task;

operating the neural network processor to execute the machine learning model, wherein the machine learning model generates knowledge score metrics for the set of individuals based at least in part on the initial set of attributes, and wherein the knowledge score metrics quantify a prediction of an analytic capability of each individual in the set of individuals for performing the specific task;

comparing, by a second processor in the computing environment, the knowledge score metrics to a specified threshold;

identifying, by the second processor, an additional set of attributes of individuals from the set of individuals having knowledge score metrics exceeding the specified threshold, wherein the additional set of attributes specifically identifies attributes of individuals capable of performing the specific task, wherein at least one attribute of the additional set of attributes is not included in the initial set of attributes used to generate the knowledge score metrics; and

providing, as additional training input to the machine learning model, the additional set of attributes.

2. The method of claim 1 , wherein the specific task comprises:

analyzing a data set; and

based on analyzing the data set, generating metadata labels for data in the data set.

3. The method of claim 1 , wherein the initial set of attributes comprises one or more of: demographic data, education data, or data regarding a role within a company, and wherein the initial set of attributes indicates a degree of familiarity with a given data set.

4. The method of claim 1 , wherein the set of individuals comprises people within a company that meet a specified minimum criterion.

5. The method of claim 1 , wherein the specified threshold for the knowledge score metrics is based on the knowledge score metrics of a seed set of experts.

6. The method of claim 1 , further comprising:

requesting individuals having the additional set of attributes to perform metadata labeling of metadata for data in a data set;

assessing an accuracy of the metadata labeling performed by the individuals; and

updating the metadata when the accuracy of the metadata labeling is assessed as accurate.

7. The method of claim 1 , further comprising:

requesting individuals having the additional set of attributes to perform metadata labeling;

assessing an accuracy of the metadata labeling performed by the individuals; and

updating the machine learning model for the individuals having the additional set of attributes, based on the accuracy of the metadata labeling.

8. The method of claim 7 , further comprising:

updating the machine learning model by adjusting weights of attributes of at least one of: the initial set of attributes or the additional set of attributes which are determined to be beneficial or less relevant for performing metadata labeling.

9. A non-transitory computer readable medium storing instructions which, when executed by at least one processor, cause the at least one processor to perform operations, the operations comprising:

forming a machine learning model in a computing environment, the machine learning model configured to analyze input data comprising an initial set of attributes to predict analytical capabilities of a set of individuals for performing a specific task;

executing the machine learning model, wherein the machine learning model generates knowledge score metrics for the set of individuals based at least in part on the initial set of attributes, and wherein the knowledge score metrics quantify a prediction of an analytic capability of each individual in the set of individuals for performing the specific task;

comparing, in the computing environment, the knowledge score metrics to a specified threshold;

identifying an additional set of attributes of individuals from the set of individuals having knowledge score metrics exceeding the specified threshold, wherein the additional set of attributes specifically identifies attributes of individuals capable of performing the specific task, wherein at least one attribute of the additional set of attributes is not included in the initial set of attributes used to generate the knowledge score metrics; and

providing, as additional training input to the machine learning model, the additional set of attributes.

10. The non-transitory computer readable medium of claim 9 , wherein the specific task comprises:

analyzing a data set; and

based on analyzing the data set, generating metadata labels for data in the data set.

11. The non-transitory computer readable medium of claim 9 , wherein the initial set of attributes comprises one or more of: demographic data, education data, or data regarding a role within a company, and wherein the initial set of attributes indicates a degree of familiarity with a given data set.

12. The non-transitory computer readable medium of claim 9 , wherein the set of individuals comprises people within a company that meet a specified minimum criterion.

13. The non-transitory computer readable medium of claim 9 , wherein the specified threshold for the knowledge score metrics is based on the knowledge score metrics of a seed set of experts.

14. The non-transitory computer readable medium of claim 9 , the operations further comprising:

requesting individuals having the additional set of attributes to perform metadata labeling of metadata for data in a data set;

assessing an accuracy of the metadata labeling performed by the individuals; and

updating the metadata when the accuracy of the metadata labeling is assessed as accurate.

15. The non-transitory computer readable medium of claim 9 , the operations further comprising:

requesting individuals having the additional set of attributes to perform metadata labeling;

assessing an accuracy of the metadata labeling performed by the individuals; and

updating the machine learning model for the individuals having the additional set of attributes based on the accuracy of the metadata labeling.

16. A system comprising:

a memory operable to store a predictive model;

a first processor communicatively coupled to the memory, wherein the first processor is operable to execute the predictive model to perform operations, the operations comprising:

generating knowledge score metrics based on an initial set of attributes for individuals included in a specified population, wherein the knowledge score metrics quantify a prediction of a capability of an individual for performing metadata labeling; and

a second processor communicatively coupled to the memory and the first processor, wherein the second processor is operable to perform second operations, the second operations comprising:

comparing the knowledge score metrics to a specified threshold; and

identifying an additional set of attributes of individuals from a specified population having knowledge score metrics exceeding the specified threshold, wherein the additional set of attributes specifically identifies attributes of individuals capable of performing metadata labeling, wherein at least one attribute of the additional set of attributes is not included in the initial set of attributes on which the knowledge score metrics are based,

wherein the first processor is further for updating the predictive model using the additional set of attributes as new training input.

17. The system of claim 16 , wherein the initial set of attributes comprises one or more of: demographic data, education data, or data regarding a role within a company, and wherein the initial set of attributes indicates a degree of familiarity with a given data set.

18. The system of claim 16 , wherein the specified threshold for the knowledge score metrics is based on knowledge score metrics of a seed set of experts.

19. The system of claim 16 ,

wherein the second processor is further operable to perform second operations comprising:

generating requests to individuals having the additional set of attributes to perform metadata labeling; and

wherein the first processor is further operable to perform operations comprising:

receiving as input an assessment of an accuracy of the metadata labeling performed by the individuals; and

updating the predictive model for the individuals having the additional set of attributes based on the accuracy of the metadata labeling.

20. The system of claim 19 , wherein the second processor is further operable to perform second operations comprising:

updating the predictive model by adjusting weights of attributes of at least one of: the initial set of attributes or the additional set of attributes which are determined by the predictive model to be beneficial or less relevant for performing metadata labeling.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2023
From: WOODS, ROBERT, JR.; AUSTIN, MARK D.
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 063563/0544 →
Continuity (2)
Continuation 16419651 · May 22, 2019
Related Publication 20230222341A1 · Jul 13, 2023
References Cited (71)
US 6879967B1 · Stork · 2005 [cited by applicant]
US 7945470B1 · Cohen · 2011 [cited by examiner]
US 8027939B2 · Fung et al. · 2011 [cited by applicant]
US 8265977B2 · Scarborough · 2012 [cited by examiner]
US 8527432B1 · Guo et al. · 2013 [cited by applicant]
US 8805756B2 · Boss et al. · 2014 [cited by applicant]
US 8825561B2 · Bellamy et al. · 2014 [cited by applicant]
US 9070085B2 · Arseneault et al. · 2015 [cited by applicant]
US 9135573B1 · Rodriguez et al. · 2015 [cited by applicant]
US 9342796B1 · McClintock et al. · 2016 [cited by applicant]
US 9396439B2 · Ebadollahi et al. · 2016 [cited by applicant]
US 9461876B2 · Van Dusen et al. · 2016 [cited by applicant]
US 9626628B2 · Dasgupta et al. · 2017 [cited by applicant]
US 9703823B2 · Daly et al. · 2017 [cited by applicant]
US 10002330B2 · Kulkarni et al. · 2018 [cited by applicant]
US 10540607B1 · Oldridge · 2020 [cited by examiner]
US 10592806B1 · Kannan · 2020 [cited by examiner]
US 10733556B2 · Bencke · 2020 [cited by examiner]
US 11107292B1 · Little · 2021 [cited by examiner]
US 11120480B2 · Agost et al. · 2021 [cited by applicant]
US 11308364B1 · Herman et al. · 2022 [cited by applicant]
US 11386442B2 · Lahav · 2022 [cited by examiner]
US 12106239B1 · Perry · 2024 [cited by examiner]
US 20020174008A1 · Noteboom · 2002 [cited by examiner]
US 20100049574A1 · Paul · 2010 [cited by examiner]
US 20120110071A1 · Zhou et al. · 2012 [cited by applicant]
US 20130152091A1 · Liu · 2013 [cited by examiner]
US 20130290207A1 · Bonmassar · 2013 [cited by examiner]
US 20140149507A1 · Redfern · 2014 [cited by examiner]
US 20150317564A1 · Chen et al. · 2015 [cited by applicant]
US 20160055010A1 · Baird · 2016 [cited by examiner]
US 20170220952A1 · Baughman et al. · 2017 [cited by applicant]
US 20170293859A1 · Gusev et al. · 2017 [cited by applicant]
US 20170316514A1 · Huang et al. · 2017 [cited by applicant]
US 20170323233A1 · Bencke · 2017 [cited by examiner]
US 20170344899A1 · Zimmerman et al. · 2017 [cited by applicant]
US 20180001206A1 · Osman et al. · 2018 [cited by applicant]
US 20180039910A1 · Hari Haran et al. · 2018 [cited by applicant]
US 20180046938A1 · Allen et al. · 2018 [cited by applicant]
US 20180082211A1 · Allen et al. · 2018 [cited by applicant]
US 20180253664A1 · Rabenoro et al. · 2018 [cited by applicant]
US 20180268319A1 · Guo et al. · 2018 [cited by applicant]
US 20180365592A1 · Gu et al. · 2018 [cited by applicant]
US 20190034883A1 · Liang et al. · 2019 [cited by applicant]
US 20190102681A1 · Roberts · 2019 [cited by examiner]
US 20190102693A1 · Yates et al. · 2019 [cited by applicant]
US 20190251457A1 · Byrnes · 2019 [cited by examiner]
US 20200005214A1 · Kulkarni · 2020 [cited by examiner]
US 20200104777A1 · Bouhini · 2020 [cited by examiner]
US 20200151641A1 · Avegliano · 2020 [cited by examiner]
US 20200211041A1 · Raudies · 2020 [cited by examiner]
US 20200302368A1 · Mathiesen · 2020 [cited by examiner]
US 20200302398A1 · Wali et al. · 2020 [cited by applicant]
US 20200342409A1 · Jang · 2020 [cited by applicant]
US 20200372090A1 · Markman · 2020 [cited by examiner]
US 20210073737A1 · Flynn · 2021 [cited by examiner]
CN 104778173A · 2015 [cited by examiner]
CN 108829763A · 2018 [cited by applicant]
CN 109388674A · 2019 [cited by applicant]
EP 2863351A1 · 2015 [cited by examiner]
RU 2632143C1 · 2017 [cited by applicant]
WO WO9505649A1 · 1995 [cited by examiner]
WO 2010080722A2 · 2010 [cited by applicant]
WO 2012166660A2 · 2012 [cited by applicant]
WO 2014136316A1 · 2014 [cited by applicant]
WO 2014144587A2 · 2014 [cited by applicant]
WO 2017052942A1 · 2017 [cited by applicant]
Roli, “Semi-Supervised Multiple Classifier Systems: Background and Research Directions”, Dept. of Electrical and Electronic Engineering, University of Cagliari, 2014, 10 pages. [cited by applicant]
Pal et al., “Evolution of Experts in Question Answering Communities”, Proceedings of the Sixth International AAAI Conference on Weblogs and Social Media, Department of Computer Science, University of Minnesota, 2012, 8 … [cited by applicant]
Mahapatra et al., “Combining Multiple Expert Annotations Using Semi-Supervised Learning and Graph Cuts For Crohn's Disease Segmentation”, International MICCAI Workshop on Computational and Clinical Challenges in Abdomin… [cited by applicant]
Barbuzzi et al., “About restraining rule in multi-expert intelligent system for semi-supervised learning using SVM classifiers”, International Journal of Signal and Imaging Systems Engineering, vol. 7, No. 4, Dec. 2014,… [cited by applicant]