IP Library Granted Patent US 12,724,637
Granted Patent B2
US 12,724,637 · App. 18/343,941 · Granted Sep 1, 2026

Prioritizing curation targets for data curation based on downstream impact

Inventors: Ofir Ezrielev (Beer Sheva, IL); Hanna Yehuda (Acton, MA); Kristen Jeanne Walsh (Austin, TX)
Assignee: Dell Products L.P.
G06F9/5005G06F9/5038G06F16/215G06F18/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,637
App. No.
18/343,941
Granted
Sep 1, 2026
Kind
B2
Abstract

Methods and systems for curating data by a data manager are disclosed. Data may be curated from various data sources before being provided to downstream consumers that may rely on the trustworthiness of the curated data in order to provide desired computer-implemented services. During the data curation process, data curation resources are used to improve the trustworthiness and/or value of the collected data. However, data curation resources (e.g., data curators, computing resources) may be limited and/or insufficient to perform the data curation process as desired, which may result in unusable and/or uncurated (e.g., untrustworthy) data. Thus, portions of the data (e.g., curation targets) may be prioritized (e.g., relative to other curation targets). The curation targets may be curated with the available data curation resources based on their relative priority in order to reduce the likelihood of providing untrustworthy data to entities (e.g., downstream consumers) that facilitate the computer-implemented services.

Claims (65)

1 . A method for curating data by a data manager, comprising:

obtaining at least a portion of the data from a data source;

identifying curation targets of the data;

making a determination regarding whether sufficient data curation resources are available to perform a data curation process for the curation targets within a target period of time;

in an instance of the determination where there are insufficient data curation resources available:

identifying a data curation resource of the data curation resources that has available curation bandwidth;

obtaining, based on scoring criteria, impact scores for the curation targets, wherein each impact score of the impact scores is based on at least one of:

a frequency of use of a corresponding curation target by an inference model that ingests at least a second portion of the data to generate an inference;

a measure of relative contribution of the corresponding curation target to the inference;

a measure of confidence in the inference; and

a measure of importance of the corresponding curation target or the inference to a downstream consumer;

obtaining a rank for at least one curation target based on the impact score, the rank being usable to order the at least one curation target;

selecting a curation target of the curation targets for the data curation resource based on a rank ordering of the curation targets, the rank ordering being based on the impact scores for the curation targets;

assigning the curation target to the data curation resource in order to complete the data curation process for a portion of the curation targets within the target period of time; and

curating the curation target using the data curation resource to obtain at least partially curated data.

2 . The method of claim 1 , wherein each impact score is based, at least in part, on a number of occurrences of the curation target in downstream use of the data.

3 . The method of claim 2 , wherein each impact score is further based, at least in part, on an attribution score for the curation target, the attribution score indicating a relative level of contribution to a future outcome in which the curation target is usable in the downstream use of the data.

4 . The method of claim 3 , wherein each impact score is further based, at least in part, on a level of confidence in predicting the future outcome through the downstream use of the data.

5 . The method of claim 4 , wherein each impact score is further based, at least in part, on the measure of importance of the curation target to a downstream consumer.

6 . The method of claim 5 , wherein each impact score is further based, at least in part, on a measure of dependence that the downstream consumer has on the predicting of the future outcome.

7 . The method of claim 6 , wherein the portion of the curation targets excludes at least one of the curation targets.

8 . The method of claim 1 ,

wherein the at least partially curated data complying with a schema for downstream use of the curation target.

9 . A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for curating data by a data manager, the operations comprising:

obtaining at least a portion of the data from a data source;

identifying curation targets of the data;

making a determination regarding whether sufficient data curation resources are available to perform a data curation process for the curation targets within a target period of time;

in an instance of the determination where there are insufficient data curation resources available:

identifying a data curation resource of the data curation resources that has available curation bandwidth;

obtaining, based on scoring criteria, impact scores for the curation targets, wherein each impact score of the impact scores is based on at least one of:

a frequency of use of a corresponding curation target by an inference model that ingests at least a second portion of the data to generate an inference;

a measure of relative contribution of the corresponding curation target to the inference;

a measure of confidence in the inference; and

a measure of importance of the corresponding curation target or the inference to a downstream consumer;

obtaining a rank for at least one curation target based on the impact score, the rank being usable to order the at least one curation target;

selecting a curation target of the curation targets for the data curation resource based on a rank ordering of the curation targets, the rank ordering being based on the impact scores for the curation targets;

assigning the curation target to the data curation resource in order to complete the data curation process for a portion of the curation targets within the target period of time; and

curating the curation target using the data curation resource to obtain at least partially curated data.

10 . The non-transitory machine-readable medium of claim 9 , wherein each impact score is based at least in part, on a number of occurrences of the curation target in downstream use of the data.

11 . The non-transitory machine-readable medium of claim 10 , wherein each impact score is further based, at least in part, on an attribution score for the curation target, the attribution score indicating a relative level of contribution to a future outcome in which the curation target is usable in the downstream use of the data.

12 . The non-transitory machine-readable medium of claim 9 , wherein the at least partially curated data complying with a schema for downstream use of the curation target.

13 . A data processing system, comprising:

a processor; and

a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for curating data by a data manager, the operations comprising:

obtaining at least a portion of the data from a data source;

identifying curation targets of the data;

making a determination regarding whether sufficient data curation resources are available to perform a data curation process for the curation targets within a target period of time;

in an instance of the determination where there are insufficient data curation resources available:

identifying a data curation resource of the data curation resources that has available curation bandwidth;

obtaining, based on scoring criteria, impact scores for the curation targets, wherein each impact score of the impact scores is based on at least one of:

a frequency of use of a corresponding curation target by an inference model that ingests at least a second portion of the data to generate an inference;

a measure of relative contribution of the corresponding curation target to the inference;

a measure of confidence in the inference; and

a measure of importance of the corresponding curation target or the inference to a downstream consumer;

obtaining a rank for at least one curation target based on the impact score, the rank being usable to order the at least one curation target;

selecting a curation target of the curation targets for the data curation resource based on a rank ordering of the curation targets, the rank ordering being based on the impact scores for the curation targets;

assigning the curation target to the data curation resource in order to complete the data curation process for a portion of the curation targets within the target period of time; and

curating the curation target using the data curation resource to obtain at least partially curated data.

14 . The data processing system of claim 13 , wherein each impact score is based at least in part, on a number of occurrences of the curation target in downstream use of the data.

15 . The data processing system of claim 14 , wherein each impact score is further based, at least in part, on an attribution score for the curation target, the attribution score indicating a relative level of contribution to a future outcome in which the curation target is usable in the downstream use of the data.

16 . The data processing system of claim 15 , wherein each impact score is further based, at least in part, on a level of confidence in predicting the future outcome through the downstream use of the data.

17 . The data processing system of claim 16 , wherein each impact score is further based, at least in part, on the measure of importance of the curation target to a downstream consumer.

18 . The data processing system of claim 17 , wherein each impact score is further based, at least in part, on a measure of dependence that the downstream consumer has on the predicting of the future outcome.

19 . The data processing system of claim 18 , wherein the portion of the curation targets excludes at least one of the curation targets.

20 . The data processing system of claim 13 , wherein the at least partially curated data complying with a schema for downstream use of the curation target.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2023
From: EZRIELEV, OFIR; YEHUDA, HANNA; WALSH, KRISTEN JEANNE
To: DELL PRODUCTS L.P.
Reel/Frame 064153/0745 →
Continuity (1)
Related Publication 20250004855A1 · Jan 2, 2025
References Cited (86)
US 7315805B2 · Slater · 2008 [cited by applicant]
US 9990383B2 · Brinnand · 2018 [cited by applicant]
US 10168691B2 · Zornio et al. · 2019 [cited by applicant]
US 10339470B1 · Dutta · 2019 [cited by applicant]
US 10936479B2 · Maag et al. · 2021 [cited by applicant]
US 11101037B2 · Allen · 2021 [cited by applicant]
US 11221270B2 · Evans · 2022 [cited by applicant]
US 11341605B1 · Singh · 2022 [cited by applicant]
US 11853853B1 · Beauchesne et al. · 2023 [cited by applicant]
US 12008046B1 · Curtis et al. · 2024 [cited by applicant]
US 12216651B2 · Krishnan · 2025 [cited by applicant]
US 12242892B1 · Burnett · 2025 [cited by applicant]
US 20040064750A1 · Conway · 2004 [cited by applicant]
US 20060009881A1 · Ferber et al. · 2006 [cited by applicant]
US 20130205285A1 · Pizlo · 2013 [cited by applicant]
US 20130226838A1 · Chu · 2013 [cited by applicant]
US 20130227573A1 · Morsi · 2013 [cited by applicant]
US 20140037161A1 · Rucker · 2014 [cited by applicant]
US 20140136184A1 · Hatsek · 2014 [cited by applicant]
US 20160098037A1 · Zornio · 2016 [cited by applicant]
US 20180046926A1 · Achin · 2018 [cited by applicant]
US 20180081871A1 · Williams · 2018 [cited by applicant]
US 20180137431A1 · Goldfarb · 2018 [cited by applicant]
US 20190034430A1 · Das · 2019 [cited by applicant]
US 20190236204A1 · Canim · 2019 [cited by applicant]
US 20190251479A1 · Anderson et al. · 2019 [cited by applicant]
US 20190370263A1 · Nucci · 2019 [cited by applicant]
US 20200166558A1 · Weis · 2020 [cited by applicant]
US 20200167224A1 · Abali · 2020 [cited by applicant]
US 20200202478A1 · Thumpudi et al. · 2020 [cited by applicant]
US 20200293684A1 · Harris · 2020 [cited by applicant]
US 20200344325A1 · Sarisky · 2020 [cited by applicant]
US 20200387836A1 · Nasr-Azadani · 2020 [cited by applicant]
US 20210027771A1 · Hall · 2021 [cited by applicant]
US 20210098133A1 · Chowdhry · 2021 [cited by applicant]
US 20210116505A1 · Shu · 2021 [cited by applicant]
US 20210117548A1 · Gokhman · 2021 [cited by applicant]
US 20210374143A1 · Neill · 2021 [cited by applicant]
US 20210377286A1 · Shukla et al. · 2021 [cited by applicant]
US 20210406110A1 · Vaid et al. · 2021 [cited by applicant]
US 20220092234A1 · Karri · 2022 [cited by applicant]
US 20220215142A1 · Gutierrez · 2022 [cited by applicant]
US 20220215243A1 · Narayanaswami · 2022 [cited by applicant]
US 20220301027A1 · Basta · 2022 [cited by applicant]
US 20220309411A1 · Ramaswamy · 2022 [cited by examiner]
US 20220310276A1 · Wilkinson · 2022 [cited by applicant]
US 20220374399A1 · Kementsietsidis · 2022 [cited by applicant]
US 20230014438A1 · Jones · 2023 [cited by applicant]
US 20230040284A1 · Ali-Tolppa · 2023 [cited by applicant]
US 20230040834A1 · Haile · 2023 [cited by applicant]
US 20230126260A1 · Elsakhawy et al. · 2023 [cited by applicant]
US 20230153095A1 · Rahill-Marier · 2023 [cited by applicant]
US 20230161596A1 · Vadapandeshwara · 2023 [cited by applicant]
US 20230196096A1 · Milne · 2023 [cited by applicant]
US 20230213930A1 · Rakshit · 2023 [cited by applicant]
US 20230293907A1 · Shade · 2023 [cited by applicant]
US 20230315078A1 · Sepulveda et al. · 2023 [cited by applicant]
US 20230342281A1 · Haile · 2023 [cited by applicant]
US 20230418280A1 · Emery · 2023 [cited by applicant]
US 20240095576A1 · Pinho · 2024 [cited by applicant]
US 20240119364A1 · Jain · 2024 [cited by applicant]
US 20240126888A1 · Kalou et al. · 2024 [cited by applicant]
US 20240235952A9 · Hicks · 2024 [cited by applicant]
US 20240281419A1 · Alfaras · 2024 [cited by applicant]
US 20240281522A1 · Kuo · 2024 [cited by applicant]
US 20240320252A1 · U · 2024 [cited by applicant]
US 20240330136A1 · Furlong · 2024 [cited by applicant]
US 20240370749A1 · Binkley · 2024 [cited by applicant]
US 20240412104A1 · Zhang · 2024 [cited by applicant]
Bosch et al., “Towards Automated Detection of Data Pipeline Faults”, 2020 27th Asia-Pacific Software Engineering Conference (APSEC). IEEE, pp. 346-355 (Year: 2020). [cited by applicant]
Grafberger et al., Towards Interactively Improving ML Data Preparation Code via “Shadow Pipelines”, DEEM '24: Proceedings of the Eighth Workshop on Data Management for End-to-End Machine Learning, published on Jun. 9, 2… [cited by applicant]
Chowdhury et al., “An Approach for Data Pipeline with Distributed Query Engine for Industrial Applications”, published in 2020 25th IEEE International Conference on Emerging Technologies and Factory Automation (ETFA), r… [cited by applicant]
Wang, Haozhe, et al., “A graph neural network-based digital twin for network slicing management,” IEEE Transactions on Industrial Informatics 18.2 (2020): 1367-1376 (11 Pages). [cited by applicant]
Almasan, Paul, et al., “Digital Twin Network: Opportunities and challenges,” arXiv preprint arXiv:2201.01144 (2022) (7 Pages). [cited by applicant]
Hu, Weifei, et al., “Digital twin: A state-of-the-art review of its enabling technologies, applications and challenges,” Journal of Intelligent Manufacturing and Special Equipment 2.1 (2021): 1-34 (34 Pages). [cited by applicant]
Khan, Latif U., et al., “Digital-Twin-Enabled 6G: Vision, Architectural Trends, and Future Directions,” IEEE Communications Magazine 60.1 (2022): 74-80 (7 Pages). [cited by applicant]
Nguyen, Huan X., et al., “Digital Twin for 5G and Beyond,” IEEE Communications Magazine 59.2 (2021): 10-15. (12 Pages). [cited by applicant]
Wang, Danshi, et al., “The Role of Digital Twin in Optical Communication: Fault Management, Hardware Configuration, and Transmission Simulation,” IEEE Communications Magazine 59.1 (2021): 133-139 (6 Pages). [cited by applicant]
Pang, Toh Yen, et al., “Developing a digital twin and digital thread framework for an ‘Industry 4.0’Shipyard,” Applied Sciences 11.3 (2021): 1097 (22 Pages). [cited by applicant]
Isto, Pekka, et al., “5G based machine remote operation development utilizing digital twin,” Open Engineering 10.1 (2020): 265-272 (8 Pages). [cited by applicant]
Redick, William, “What is Outcome-Based Selling?” Global Performance, Web Page <https://globalperformancegroup.com/what-is-outcome-based-selling/> accessed on Feb. 14, 2023 (8 Pages). [cited by applicant]
“The Best Data Curation Tools for Computer Vision in 2022,” Web Page <https://www.lightly.ai/post/data-curation-tools-2022> accessed on Feb. 14, 2023 (9 Pages). [cited by applicant]
Bebee, Troy et al., “How to detect machine-learned anomalies in real-time foreign exchange data,” Google Cloud, Jun. 10, 2021, Web Page <https://cloud.google.com/blog/topics/financial-services/detect-anomalies-in-real-t… [cited by applicant]
Wang, Haozhe, et al., “A graph neural network-based digital twin for network slicing management,” IEEE Transactions on Industrial Informatics 18.2 (2020): 1367-1376 (10 Pages). [cited by applicant]
Stonebraker et al., “Data Curation at Scale: The Data Tamer System”, 6th Biennial Conference on Innovative Data Systems Research (CIDR '13), Jan. 6-9, 2013, 10 pps. (Year: 2013). [cited by applicant]
T. Leemann et al., “I Prefer not to say: Protecting User Consent in Models with Optional Personal Data”, retrieved from <https://arxiv.org/abs/2210.13954v4> Version 4, Jun. 6, 2023, pp. 1-32 (32 pages). [cited by applicant]