IP Library Granted Patent US 12,400,146
Granted Patent B1
US 12,400,146 · App. 17/717,527 · Granted Aug 26, 2025

Methods and apparatus for automated enterprise machine learning with data fused from multiple domains

Inventors: Harshil Shah (Centreville, VA); Manoj Kumar Goyal (Fremont, CA); Jin Yu (Union City, CA)
Assignee: CIPIO Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,146
App. No.
17/717,527
Granted
Aug 26, 2025
Kind
B1
Abstract

A method can include identifying a first set of data columns in a first data domain and a second set of data columns in a second data domain that are similar. The method can further include selecting a first subset of data columns in the first data domain and a second subset of data columns in the second data domain that have matching parameters. The method can further include arranging the first subset of data columns and the second set of data columns into a multi-domain dataset. Each column in the multi-domain dataset can be classified into structured data or unstructured data. The structured data can be used for training a machine learning model. The method can further include replacing under-performing features in the second data domain before replacing under-performing features in the first data domain. The method can further include re-training the machine learning model after replacing under-performing features.

Claims (46)

1. A method, comprising:

identifying a first set of data columns in a first data domain having a plurality of data types;

identifying a second set of data columns in a second data domain having a plurality of data types that are the same as or substantially similar to the plurality of data types of the first data domain;

selecting (1) a first subset of data columns from the first set of data columns in the first data domain and having a plurality of parameters, and (2) a second subset of data columns from the second set of data columns in the second data domain and having a plurality of parameters matching the plurality of parameters of the first subset of data columns;

arranging the first subset of data columns and the second set of data columns into a multi-domain dataset, each column in the multi-domain dataset including an attribute from the first set of data columns or an attribute the second set of data columns, each row in the multi-domain dataset including multi-domain data from the first data domain and the second data domain;

classifying each column in the multi-domain dataset into structured data or unstructured data, columns in the multi-domain dataset with structured data being mapped to a set of features for training at least one machine learning model, each feature from the set of features associated with the first data domain or the second data domain;

replacing a first set of under-performing features from a subset of features from the set of features in the second data domain, the under-performing features from the first set of under-performing features identified by evaluating the at least one machine learning model;

replacing a second set of under-performing features from a subset of features from the set of features in the first data domain, after the replacing the first set of under-performing features, the under-performing features from the second set of under-performing features identified by evaluating the at least one machine learning model, a predetermined portion of features from the subset of features from the set of features in the first data domain being retained; and

re-training the at least one machine learning model after replacing the first set of under-performing features and the second set of under-performing features.

2. The method of claim 1 , wherein the first data domain includes telecom data and the second data domain includes fitness data.

3. The method of claim 1 , wherein the multi-domain dataset includes a multi-domain profile.

4. The method of claim 1 , wherein the multi-domain dataset includes a multi-domain profile generated by concatenating rows in the multi-domain dataset having a common attribute.

5. The method of claim 1 , wherein the plurality of parameters of the first subset of data columns includes at least one of an email address or a phone number.

6. The method of claim 1 , wherein the first data domain includes at least one of a diet data domain, a medical record domain, a financial domain, a network domain, a demographic domain, or a customer domain.

7. The method of claim 1 , wherein the under-performing features from the first set of under-performing features decrease an accuracy of an output generated by the at least one machine learning model.

8. An apparatus, comprising:

a memory; and

a processor operatively coupled to the memory, the processor configured to:

identify a first set of data columns in a first data domain having a plurality of data types;

identify a second set of data columns in a second data domain having a plurality of data types that are the same as or substantially similar to the plurality of data types of the first data domain;

select (1) a first subset of data columns from the first set of data columns in the first data domain and having a plurality of parameters, and (2) a second subset of data columns from the second set of data columns in the second data domain and having a plurality of parameters matching the plurality of parameters of the first subset of data columns;

arrange the first subset of data columns and the second set of data columns into a multi-domain dataset, each column in the multi-domain dataset including an attribute from the first set of data columns or an attribute the second set of data columns, each row in the multi-domain dataset including multi-domain data from the first data domain and the second data domain;

classify each column in the multi-domain dataset into structured data or unstructured data, columns in the multi-domain dataset with structured data being mapped to a set of features for training at least one machine learning model, each feature from the set of features associated with the first data domain or the second data domain;

replace a first set of under-performing features from a subset of features from the set of features in the second data domain, the under-performing features from the first set of under-performing features identified by evaluating the at least one machine learning model;

replace a second set of under-performing features from a subset of features from the set of features in the first data domain, after the replacing the first set of under-performing features, the under-performing features from the second set of under-performing features identified by evaluating the at least one machine learning model, a predetermined portion of features from the subset of features from the set of features in the first data domain being retained; and

re-train the at least one machine learning model after replacing the first set of under-performing features and the second set of under-performing features.

9. The apparatus of claim 8 , wherein evaluating the under-performing features from the first set of under-performing features includes determining a speed the at least one machine learning model generating an output, the under-performing features from the first set of under-performing features identified based on the speed.

10. The apparatus of claim 8 , wherein the first data domain includes telecom data and the second data domain includes fitness data.

11. The apparatus of claim 8 , wherein the multi-domain dataset includes multi-domain profiles.

12. The apparatus of claim 8 , wherein the multi-domain dataset includes multi-domain profiles generated by concatenating rows in the multi-domain dataset having a common attribute.

13. The apparatus of claim 8 , wherein the plurality of parameters of the first subset of data columns includes at least one of an email address or a phone number.

14. The apparatus of claim 8 , wherein the first data domain includes at least one of a diet data domain, a medical record domain, a financial domain, a network domain, a demographic domain, or a customer domain.

15. The apparatus of claim 8 , wherein evaluating the under-performing features from the first set of under-performing features includes determining an accuracy of an output generated by the at least one machine learning model, the under-performing features from the first set of under-performing features identified based on the accuracy.

16. A non-transitory, processor readable medium storing instructions that, when executed by a processor, cause the processor to:

identify a first set of data columns in a first data domain having a plurality of data types;

identify a second set of data columns in a second data domain having a plurality of data types that are the same as or substantially similar to the plurality of data types of the first data domain;

select (1) a first subset of data columns from the first set of data columns in the first data domain and having a plurality of parameters, and (2) a second subset of data columns from the second set of data columns in the second data domain and having a plurality of parameters matching the plurality of parameters of the first subset of data columns;

arrange the first subset of data columns and the second set of data columns into a multi-domain dataset, each column in the multi-domain dataset including an attribute from the first set of data columns or an attribute the second set of data columns, each row in the multi-domain dataset including multi-domain data from the first data domain and the second data domain;

classify each column in the multi-domain dataset into structured data or unstructured data, columns in the multi-domain dataset with structured data being mapped to a set of features for training at least one machine learning model, each feature from the set of features associated with the first data domain or the second data domain;

replace a first set of under-performing features from a subset of features from the set of features in the second data domain, the under-performing features from the first set of under-performing features identified by evaluating the at least one machine learning model;

replace a second set of under-performing features from a subset of features from the set of features in the first data domain, after the replacing the first set of under-performing features, the under-performing features from the second set of under-performing features identified by evaluating the at least one machine learning model, a predetermined portion of features from the subset of features from the set of features in the first data domain being retained; and

re-train the at least one machine learning model after replacing the first set of under-performing features and the second set of under-performing features.

17. The non-transitory, processor readable medium of claim 16 , wherein the first data domain includes telecom data and the second data domain includes fitness data.

18. The non-transitory, processor readable medium of claim 16 , wherein the under-performing features from the first set of under-performing features decrease an accuracy of an output generated by the at least one machine learning model.

19. The non-transitory, processor readable medium of claim 16 , wherein the plurality of parameters of the first subset of data columns includes an email address and a phone number.

20. The non-transitory, processor readable medium of claim 16 , wherein the multi-domain dataset includes multi-domain profiles.

Assignments (2)
CHANGE OF NAME Recorded May 20, 2026
From: CIPIO, INC.
To: VIDEOFORCEAI, INC.
Reel/Frame 075583/0687 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2022
From: SHAH, HARSHIL; GOYAL, MANOJ KUMAR; YU, JIN
To: CIPIO INC.
Reel/Frame 059582/0837 →
Continuity (1)
Provisional Application 63173737 · Apr 12, 2021
References Cited (20)
US 6925454B2 · Lam · 2005 [cited by examiner]
US 10733230B2 · Jo · 2020 [cited by applicant]
US 11197036B2 · Chao · 2021 [cited by applicant]
US 11704893B2 · Ren et al. · 2023 [cited by applicant]
US 11886311B2 · Shen · 2024 [cited by examiner]
US 11898932B2 · Zhu · 2024 [cited by examiner]
US 20100042563A1 · Livingston · 2010 [cited by examiner]
US 20140214895A1 · Higgins · 2014 [cited by examiner]
US 20140272884A1 · Allen · 2014 [cited by examiner]
US 20160092561A1 · Liu et al. · 2016 [cited by applicant]
US 20200104641A1 · Alvelda, VII · 2020 [cited by examiner]
US 20220292389A1 · Seabolt · 2022 [cited by examiner]
US 20230162502A1 · Patel et al. · 2023 [cited by applicant]
US 20230245451A1 · Zhang · 2023 [cited by applicant]
US 20240124004A1 · Donderici · 2024 [cited by applicant]
US 20250008188A1 · Maity et al. · 2025 [cited by applicant]
CN 109922373A · 2019 [cited by applicant]
CN 111726536A · 2020 [cited by applicant]
CN 112929744A · 2021 [cited by applicant]
EP 3598371A1 · 2020 [cited by applicant]