IP Library Granted Patent US 12,547,652
Granted Patent B2
US 12,547,652 · App. 18/175,783 · Granted Feb 10, 2026

Systems and methods for automated data set matching services

Inventors: Wai Ho Lau (Manhasset, NY); Kalabe Gizaw Haile (Harrison, NJ); Scott Colin Laliberte (Sewell, NJ); Arun Kumar Tripathi (Los Angeles, CA)
Assignee: Protiviti Inc.
G06F16/3347G06F16/3329G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,547,652
App. No.
18/175,783
Granted
Feb 10, 2026
Kind
B2
Abstract

The present solution provides systems and methods to receive a data set comprising a representation of one or more questions from a survey and provide the data set as input to each of a plurality of machine learning models trained to predict a domain associated with the one or more questions. The systems and methods can receiving as output a first domain prediction for the domain from each of the plurality of machine learning models and determine a second domain prediction for the domain for each question of the one or more questions based on applying a function to each of the first domain predictions. The systems and methods can select, based on the data set and the second domain prediction, an enumerated list of one or more answers from an answer set and cause a display of the enumerated list via a user interface for a selection.

Claims (68)

1 . A method comprising,

receiving, by a data processing system, a data set comprising a representation of one or more questions from a survey;

providing, by the data processing system, the data set as input to each instantiation of a plurality of different types of machine learning models trained to predict a domain associated with the one or more questions;

receiving as output, by the data processing system, a first domain prediction for the domain from each instantiation of a plurality of different types of machine learning models;

determining, by the data processing system, a second domain prediction for the domain for each question of the one or more questions based on applying a function to each of the first domain predictions;

selecting, by the data processing system, based on the data set and the second domain prediction, an enumerated list of one or more answers from an answer set;

and

causing, by the data processing system, a display of the enumerated list via a user interface for a selection;

receiving, by the data processing system, an indication of the selection of one of the answers from the enumerated list; and

retraining, by the data processing system, one or more machine learning models of the plurality of different types of machine learning models based on the indication of the selection.

2 . The method of claim 1 , further comprising:

generating, by each of the plurality of machine learning models, a first confidence level for each of the first domain predictions;

determining, by the data processing system, a second confidence level based on each of the first confidence levels; and

causing, by the data processing system, a display of the second confidence level for the selection of each answer of the enumerated list via the user interface.

3 . The method of claim 1 , wherein applying the function to each of the first domain predictions comprises summing a number of the machine learning models sharing a predicted domain for each of the one or more questions.

4 . The method of claim 1 , wherein applying the function to each of the first domain predictions comprises:

receiving a confidence level for the first domain predictions from each of the plurality of machine learning models;

summing the confidence level of the first domain predictions for each of the predicted domains; and

selecting a domain of the predicted domains associated with a highest confidence level summation.

5 . The method of claim 1 , wherein an answer of the answer set includes a time interval for which the answer is valid.

6 . The method of claim 1 , wherein causing the display of the enumerated list comprises:

causing the display of a plurality of domains; and

causing the display of an association between each of the answers of the enumerated list and one of the plurality of domains.

7 . A system comprising:

at least one processor associated with a data processing system;

at least one memory storing computer-readable instructions, wherein the at least one processor is operable to access the at least one memory and execute the computer-readable instructions to:

receive a data set comprising a representation of one or more questions from a survey;

provide the data set as input to each instantiation of a plurality of different types of machine learning models trained to predict a domain associated with the one or more questions;

receive as output a first domain prediction for the domain from each instantiation of a plurality of different types of machine learning models;

determine a second domain prediction for the domain for each question of the one or more questions based on applying a function to each of the first domain predictions;

select, based on the data set and the second domain prediction, an enumerated list of one or more answers from an answer set;

cause a display of the enumerated list via a user interface for a selection;

receive an indication of the selection of one of the answers from the enumerated list; and

retrain a machine learning model of the plurality of different types of machine learning models based on the indication of the selection.

8 . The system of claim 7 , wherein the processors execute computer-readable instructions to:

generate, by each of the plurality of machine learning models, a first confidence level for each of the first domain predictions;

determine a second confidence level based on each of the first confidence levels; and

cause a display of the second confidence level for the selection of each answer of the enumerated list via the user interface.

9 . The system of claim 7 , wherein, to apply the function to each of the first domain predictions, the processors execute computer-readable instructions to sum a number of the machine learning models sharing a predicted domain for each of the one or more questions.

10 . The system of claim 7 , wherein, to apply the function to each of the first domain predictions comprises, the processors execute computer-readable instructions to:

receive a confidence level for the first domain predictions from each of the plurality of machine learning models;

sum the confidence level of the first domain predictions for each of the predicted domains; and

select a domain of the predicted domains associated with a highest confidence level summation.

11 . The system of claim 7 , wherein each answer of the answer set includes a time interval for which the answer is valid.

12 . The system of claim 7 , wherein to cause the display of the enumerated list, the processors execute computer-readable instructions to:

cause the display of a plurality of domains; and

cause the display of an association between each of the answers of the enumerated list and one of the plurality of domains.

13 . A non-transitory computer-readable media comprising computer-readable instructions stored thereon that when executed by one or more processors of a data processing system cause the one or more processors to:

receive a data set comprising a representation of one or more questions from a survey;

provide the data set as input to each instantiation of a plurality of different types of machine learning models trained to predict a domain associated with the one or more questions;

receive as output a first domain prediction for the domain from each instantiation of a plurality of different types of machine learning models;

determine a second domain prediction for the domain for each question of the one or more questions based on applying a function to each of the first domain predictions;

select, based on the data set and the second domain prediction, an enumerated list of one or more answers from an answer set; and

cause a display of the enumerated list via a user interface for a selection;

receive an indication of the selection of one of the answers from the enumerated list; and

retrain one or more machine learning models of the plurality of different types of machine learning models based on the indication of the selection.

14 . The non-transitory computer-readable media of claim 13 , wherein the computer-readable instructions comprise instructions to:

generate, by each of the plurality of machine learning models, a first confidence level for each of the first domain predictions;

determine a second confidence level based on each of the first confidence levels; and

cause a display of the second confidence level for the selection of each answer of the enumerated list via the user interface.

15 . The non-transitory computer-readable media of claim 13 , wherein the computer-readable instructions comprise instructions to sum a number of the machine learning models sharing a predicted domain for each of the one or more questions.

16 . The non-transitory computer-readable media of claim 13 , wherein, to apply the function to each of the first domain predictions comprises, the processors execute computer-readable instructions to:

receive a confidence level for the first domain predictions from each of the plurality of machine learning models;

sum the confidence level of the first domain predictions for each of the predicted domains; and

select a domain of the predicted domains associated with a highest confidence level summation.

17 . The non-transitory computer-readable media of claim 13 , wherein to cause the display of the enumerated list, the processors execute computer-readable instructions to:

cause the display of a plurality of domains; and

cause the display of an association between each of the answers of the enumerated list and one of the plurality of domains.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2023
From: LAU, WAI HO; HAILE, KALABE GIZAW; LALIBERTE, SCOTT COLIN; TRIPATHI, ARUN KUMAR
To: PROTIVITI INC.
Reel/Frame 062899/0571 →
Continuity (1)
Related Publication 20240289364A1 · Aug 29, 2024
References Cited (22)
US 11631021B1 · Benjamin · 2023 [cited by examiner]
US 20100257167A1 · Liu · 2010 [cited by examiner]
US 20160203208A1 · Anderson · 2016 [cited by examiner]
US 20170124484A1 · Thompson · 2017 [cited by examiner]
US 20170372190A1 · Bishop · 2017 [cited by examiner]
US 20180032589A1 · Allen · 2018 [cited by examiner]
US 20180349377A1 · Verma · 2018 [cited by examiner]
US 20190370629A1 · Liu · 2019 [cited by examiner]
US 20200073895A1 · Vira · 2020 [cited by examiner]
US 20210391039A1 · Laszlo · 2021 [cited by examiner]
US 20220084513A1 · Sgobba · 2022 [cited by examiner]
US 20230316150A1 · Phan · 2023 [cited by examiner]
US 20230342167A1 · Radkoff · 2023 [cited by examiner]
US 20230385861A1 · Ghose · 2023 [cited by examiner]
US 20240256988A1 · Morato · 2024 [cited by examiner]
US 20240273105A1 · Martigny · 2024 [cited by examiner]
Puerto, Haritz, Gözde Gül ahin, and Iryna Gurevych. “Metaqa: Combining expert agents for multi-skill question answering.” arXiv preprint arXiv:2112.01922 (2021). (Year: 2021). [cited by examiner]
Aroussi Said Alami et al: “Improving question answering systems by using the explicit semantic analysis method”, 2016 11th International Conference on Intelligent Systems: Theories and Applications (SITA), IEEE, Oct. 19… [cited by applicant]
Foreign Search Report on PCT Dtd Apr. 22, 2024. [cited by applicant]
Haritz Puerto et al: “MetaQA:Combining Expert Agents for Multi-Skill Question Answering”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY, 14853, Jan. 22, 2023 (Jan. 22, 2023), XP09… [cited by applicant]
Aroussi Said Alami et al: “Improving question answering systems by using the explicit semantic analysis method”. 2016 11th International Conference on Intelligent Systems: Theories and Applications (SITA), IEEE, Oct. 19… [cited by applicant]
International Preliminary Report on PCT AppIn PCT/US2024/017224 dated Sep. 11, 2025. [cited by applicant]