IP Library Granted Patent US 12,468,541
Granted Patent B2
US 12,468,541 · App. 18/343,930 · Granted Nov 11, 2025

Curating anomalous data for use in a data pipeline through interaction with a data source

Inventors: Ofir Ezrielev (Beer Sheva, IL); Hanna Yehuda (Acton, MA); Kristen Jeanne Walsh (Austin, TX)
Assignee: Dell Products L.P.
G06F9/3867G06F9/3895
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,541
App. No.
18/343,930
Granted
Nov 11, 2025
Kind
B2
Abstract

Methods and systems for curating data by a data manager are disclosed. Data may be curated from various data sources before being provided to downstream consumers that may rely on the trustworthiness of the curated data in order to provide desired computer-implemented services. During the data curation process, data curation resources are used to improve the trustworthiness and/or value of the collected data. However, data curation resources (e.g., data curators, computing resources) may be limited and/or insufficient to perform the data curation process as desired, which may result in unusable and/or uncurated (e.g., untrustworthy) data. Thus, the data may be screened for anomalous data points. Features of the anomalous data points that meet importance criteria may be presented to the data source and the data source may indicate whether the features are expected. If the features are expected, the data pipeline may be populated with the anomalous data points.

Claims (52)

1 . A method of curating data by a data manager, the method comprising:

making an identification that data obtained from a data source associated with a data pipeline comprises an anomalous data point;

identifying a feature of a set of features associated with the anomalous data point that meets importance criteria;

obtaining a final data point based on the feature and through an interaction with the data source; and

populating the data pipeline with the final data point to provide the final data point to a downstream consumer of the data pipeline using one or more application programming interfaces (APIs) associated with the data pipeline, wherein populating the data pipeline with the final data point comprises storing the final data point in a data repository associated with the data pipeline, and

wherein the final data point is populated into the data pipeline to prevent the anomalous data point from being provided to a data consumer and negatively impacting operations and functionalities of a data processing system that uses the data in the data pipeline in one or more processes executed by the data processing system for providing computer-implemented services to the downstream consumer.

2 . The method of claim 1 , wherein the anomalous data point is subject to multiple valid but contrasting interpretations.

3 . The method of claim 2 , wherein the final data point is based on one of the multiple valid but contrasting interpretations as selected by the data source through the interaction.

4 . The method of claim 1 , wherein the interaction comprises presenting information regarding the feature to the data source, the information being usable to confirm whether the feature is expected by the data source.

5 . The method of claim 4 , wherein obtaining the final data point comprises:

obtaining a response from the data source that is responsive to the information;

in a first instance of the obtaining where the response confirms that the anomalous data point is expected:

using the anomalous data point as the final data point; and

in a second instance of the obtaining where the response rejects the anomalous data point as being expected:

using a data source supplied data point from the response as the final data point.

6 . The method of claim 1 , further comprising:

removing the anomalous data point from the data pipeline.

7 . The method of claim 1 , further comprising:

providing computer-implemented services using the final data point from the populated data pipeline.

8 . A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for curating data by a data manager, the operations comprising:

making an identification that data obtained from a data source associated with a data pipeline comprises an anomalous data point;

identifying a feature of a set of features associated with the anomalous data point that meets importance criteria;

obtaining a final data point based on the feature and through an interaction with the data source; and

populating the data pipeline with the final data point to provide the final data point to a downstream consumer of the data pipeline using one or more application programming interfaces (APIs) associated with the data pipeline, wherein populating the data pipeline with the final data point comprises storing the final data point in a data repository associated with the data pipeline, and

wherein the final data point is populated into the data pipeline to prevent the anomalous data point from being provided to a data consumer and negatively impacting operations and functionalities of a data processing system that uses the data in the data pipeline in one or more processes executed by the data processing system for providing computer-implemented services to the downstream consumer.

9 . The non-transitory machine-readable medium of claim 8 , wherein the anomalous data point is subject to multiple valid but contrasting interpretations.

10 . The non-transitory machine-readable medium of claim 9 , wherein the final data point is based on one of the multiple valid but contrasting interpretations as selected by the data source through the interaction.

11 . The non-transitory machine-readable medium of claim 8 , wherein the interaction comprises presenting information regarding the feature to the data source, the information being usable to confirm whether the feature is expected by the data source.

12 . The non-transitory machine-readable medium of claim 11 , wherein obtaining the final data point comprises:

obtaining a response from the data source that is responsive to the information;

in a first instance of the obtaining where the response confirms that the anomalous data point is expected:

using the anomalous data point as the final data point; and

in a second instance of the obtaining where the response rejects the anomalous data point as being expected:

using a data source supplied data point from the response as the final data point.

13 . The non-transitory machine-readable medium of claim 8 , wherein the operations further comprise:

removing the anomalous data point from the data pipeline.

14 . The non-transitory machine-readable medium of claim 8 , wherein the operations further comprise:

providing computer-implemented services using the final data point from the populated data pipeline.

15 . A data processing system, comprising:

a processor; and

a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for curating data by a data manager, the operations comprising:

making an identification that data obtained from a data source associated with a data pipeline comprises an anomalous data point;

identifying a feature of a set of features associated with the anomalous data point that meets importance criteria;

obtaining a final data point based on the feature and through an interaction with the data source; and

populating the data pipeline with the final data point to provide the final data point to a downstream consumer of the data pipeline using one or more application programming interfaces (APIs) associated with the data pipeline, wherein populating the data pipeline with the final data point comprises storing the final data point in a data repository associated with the data pipeline, and

wherein the final data point is populated into the data pipeline to prevent the anomalous data point from being provided to a data consumer and negatively impacting operations and functionalities of a data processing system that uses the data in the data pipeline in one or more processes executed by the data processing system for providing computer-implemented services to the downstream consumer.

16 . The data processing system of claim 15 , wherein the anomalous data point is subject to multiple valid but contrasting interpretations.

17 . The data processing system of claim 16 , wherein the final data point is based on one of the multiple valid but contrasting interpretations as selected by the data source through the interaction.

18 . The data processing system of claim 15 , wherein the interaction comprises presenting information regarding the feature to the data source, the information being usable to confirm whether the feature is expected by the data source.

19 . The method of claim 1 , wherein populating the data pipeline with the final data point further comprises:

removing the anomalous data point from the data pipeline while the data pipeline is being populated with the final data point.

20 . The method of claim 19 , wherein removing the anomalous data point from the data pipeline comprises deleting the anomalous data point from the data repository associated with the data pipeline.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2023
From: EZRIELEV, OFIR; YEHUDA, HANNA; WALSH, KRISTEN JEANNE
To: DELL PRODUCTS L.P.
Reel/Frame 064153/0639 →
Continuity (1)
Related Publication 20250004783A1 · Jan 2, 2025
References Cited (64)
US 7315805B2 · Slater · 2008 [cited by applicant]
US 9990383B2 · Brinnand · 2018 [cited by applicant]
US 10168691B2 · Zornio et al. · 2019 [cited by applicant]
US 10936479B2 · Maag et al. · 2021 [cited by applicant]
US 11101037B2 · Allen · 2021 [cited by applicant]
US 11221270B2 · Evans · 2022 [cited by applicant]
US 11341605B1 · Singh · 2022 [cited by applicant]
US 11853853B1 · Beauchesne · 2023 [cited by examiner]
US 12008046B1 · Curtis · 2024 [cited by examiner]
US 12216651B2 · Krishnan · 2025 [cited by applicant]
US 12242892B1 · Burnett · 2025 [cited by applicant]
US 20040064750A1 · Conway · 2004 [cited by applicant]
US 20060009881A1 · Ferber et al. · 2006 [cited by applicant]
US 20130205285A1 · Pizlo · 2013 [cited by applicant]
US 20130227573A1 · Morsi · 2013 [cited by applicant]
US 20140037161A1 · Rucker · 2014 [cited by applicant]
US 20140136184A1 · Hatsek · 2014 [cited by applicant]
US 20160098037A1 · Zornio · 2016 [cited by applicant]
US 20180081871A1 · Williams · 2018 [cited by applicant]
US 20190034430A1 · Das · 2019 [cited by applicant]
US 20190236204A1 · Canim · 2019 [cited by applicant]
US 20190251479A1 · Anderson et al. · 2019 [cited by applicant]
US 20200166558A1 · Weis · 2020 [cited by applicant]
US 20200167224A1 · Abali · 2020 [cited by applicant]
US 20200202478A1 · Thumpudi et al. · 2020 [cited by applicant]
US 20200293684A1 · Harris · 2020 [cited by applicant]
US 20210027771A1 · Hall · 2021 [cited by applicant]
US 20210116505A1 · Shu · 2021 [cited by applicant]
US 20210374143A1 · Neill · 2021 [cited by applicant]
US 20210377286A1 · Shukla et al. · 2021 [cited by applicant]
US 20210406110A1 · Vaid · 2021 [cited by examiner]
US 20220092234A1 · Karri · 2022 [cited by applicant]
US 20220301027A1 · Basta · 2022 [cited by applicant]
US 20220310276A1 · Wilkinson · 2022 [cited by applicant]
US 20230014438A1 · Jones · 2023 [cited by applicant]
US 20230040834A1 · Haile · 2023 [cited by applicant]
US 20230126260A1 · Elsakhawy · 2023 [cited by examiner]
US 20230153095A1 · Rahill-Marier · 2023 [cited by applicant]
US 20230161596A1 · Vadapandeshwara · 2023 [cited by applicant]
US 20230196096A1 · Milne · 2023 [cited by applicant]
US 20230213930A1 · Rakshit · 2023 [cited by applicant]
US 20230315078A1 · Sepulveda et al. · 2023 [cited by applicant]
US 20230418280A1 · Emery · 2023 [cited by applicant]
US 20240126888A1 · Kalou et al. · 2024 [cited by applicant]
US 20240235952A9 · Hicks · 2024 [cited by applicant]
US 20240281419A1 · Alfaras · 2024 [cited by applicant]
US 20240281522A1 · Kuo · 2024 [cited by applicant]
US 20240330136A1 · Furlong · 2024 [cited by applicant]
US 20240412104A1 · Zhang · 2024 [cited by applicant]
Wang, Haozhe, et al., “A graph neural network-based digital twin for network slicing management,” IEEE Transactions on Industrial Informatics 18.2 (2020): 1367-1376 (11 Pages). [cited by applicant]
Almasan, Paul, et al., “Digital Twin Network: Opportunities and challenges,” arXiv preprint arXiv:2201.01144 (2022) (7 Pages). [cited by applicant]
Hu, Weifei, et al., “Digital twin: A state-of-the-art review of its enabling technologies, applications and challenges,” Journal of Intelligent Manufacturing and Special Equipment 2.1 (2021): 1-34 (34 Pages). [cited by applicant]
Khan, Latif U., et al., “Digital-Twin-Enabled 6G: Vision, Architectural Trends, and Future Directions,” IEEE Communications Magazine 60.1 (2022): 74-80 (7 Pages). [cited by applicant]
Nguyen, Huan X., et al., “Digital Twin for 5G and Beyond,” IEEE Communications Magazine 59.2 (2021): 10-15. (12 Pages). [cited by applicant]
Wang, Danshi, et al., “The Role of Digital Twin in Optical Communication: Fault Management, Hardware Configuration, and Transmission Simulation,” IEEE Communications Magazine 59.1 (2021): 133-139 (6 Pages). [cited by applicant]
Pang, Toh Yen, et al., “Developing a digital twin and digital thread framework for an ‘Industry 4.0’Shipyard,” Applied Sciences 11.3 (2021): 1097 (22 Pages). [cited by applicant]
Isto, Pekka, et al., “5G based machine remote operation development utilizing digital twin,” Open Engineering 10.1 (2020): 265-272 (8 Pages). [cited by applicant]
Redick, William, “What is Outcome-Based Selling?” Global Performance, Web Page <https://globalperformancegroup. com/what-is-outcome-based-selling/> accessed on Feb. 14, 2023 (8 Pages). [cited by applicant]
“The Best Data Curation Tools for Computer Vision in 2022,” Web Page <https://www.lightly.ai/post/data-curation-tools-2022> accessed on Feb. 14, 2023 (9 Pages). [cited by applicant]
Bebee, Troy et al., “How to detect machine-learned anomalies in real-time foreign exchange data,” Google Cloud, Jun. 10, 2021, Web Page <https://cloud.google.com/blog/topics/financial-services/detect-anomalies-in-real-t… [cited by applicant]
Wang, Haozhe, et al., “A graph neural network-based digital twin for network slicing management,” IEEE Transactions on Industrial Informatics 18.2 (2020): 1367-1376 (10 Pages). [cited by applicant]
Bosch et al., “Towards Automated Detection of Data Pipeline Faults”, 2020 27th Asia-Pacific Software Engineering Conference (APSEC). IEEE, pp. 346-355 (Year: 2020). [cited by applicant]
Grafberger et al., Towards Interactively Improving ML Data Preparation Code via “Shadow Pipelines”, DEEM '24: Proceedings of the Eighth Workshop on Data Management for End-to-End Machine Learning, published on Jun. 9, 2… [cited by applicant]
Chowdhury et al., “An Approach for Data Pipeline with Distributed Query Engine for Industrial Applications”, published in 2020 25th IEEE International Conference on Emerging Technologies and Factory Automation (ETFA), r… [cited by applicant]