IP Library › Granted Patent US 12,625,872
Granted Patent B2
US 12,625,872 · App. 18/659,616 · Granted May 12, 2026

Generating overlap queries on a database system

Inventors: Matthew J. Glickman (Larchmont, NY); Orestis Kostakis (Redmond, WA); Justin Langseth (Kailua, HI)
Assignee: Snowflake Inc.
G06F16/24568G06F16/244G06F16/2456G06F16/24564
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,625,872
App. No.
18/659,616
Granted
May 12, 2026
Kind
B2
Abstract

An advanced system for refining overlap queries in a database system based on user feedback. The system monitors interactions of a first user with a first dataset on the database system, where the first dataset is associated with the first user. Feedback regarding the quality of a results dataset, generated from an executed overlap query, is received from the first user. This feedback informs the generation of a similarity score dataset that enhances the creation of new overlap queries. These new overlap queries are designed to output refined overlap datasets between the first dataset and a second dataset associated with a second user. A new joined dataset is generated by executing these overlap queries, comprising data from both the first and second datasets. A new results dataset is generated, providing the first user with refined recommendations based on additional feedback.

Claims (54)

1 . A method comprising:

monitoring, by at least one hardware processor, a first user interaction with a first dataset on a database system, the first dataset being associated with a first user;

receiving, from the first user, feedback regarding a quality of a results dataset generated from an executed overlap query, the results dataset comprising a joined dataset of data from the first dataset of the first user and data from a second dataset of a second user;

based on the feedback, generating a similarity score dataset indicating a similarity between the first dataset and the second dataset;

generating, using the similarity score dataset, a new set of overlap queries specifying join operations between the first dataset and the second dataset;

executing the new set of overlap queries to materialize a new joined dataset comprising the data from the first dataset and the data from the second dataset, the materializing comprises persisting the new joined dataset as a physical table indexed and stored in the database system for subsequent reuse;

generating a new results dataset by applying the new set of overlap queries to the new joined dataset; and

providing, to the first user, the new results dataset comprising a refinement recommendation for the overlap queries based on additional user feedback, the refinement recommendation comprising a ranked set of proposed modifications to at least one join operation associated with the overlap queries.

2 . The method of claim 1 , wherein each overlap query is generated based on non-matching data between the first dataset and a corresponding dataset of the new joined dataset, the non-matching data comprises non-matching columns between the first dataset and the new joined dataset.

3 . The method of claim 2 , wherein generating the new set of overlap queries comprises identifying, from an overlap query repository, an overlap query that is associated with the non-matching data.

4 . The method of claim 1 , further comprising:

prior to generating the new joined dataset, executing a data cleaning process on the first dataset and the second dataset to ensure data quality and consistency.

5 . The method of claim 1 , wherein the feedback from the first user is received via an interactive user interface and further comprises:

enabling the first user to rate the quality of the results dataset on a predefined scale.

6 . The method of claim 1 , wherein the new set of overlap queries is generated based on a machine learning model that predicts potential utility of overlap datasets based on historical user feedback and query outcomes.

7 . The method of claim 6 , wherein the machine learning model is retrained periodically with new user feedback and results datasets to improve accuracy of the new set of overlap queries.

8 . The method of claim 1 , further comprising:

categorizing each dataset in the database system by a semantic type of a data field within each dataset, wherein the semantic type includes at least one of a ZIP code, a date, or a numerical identifier.

9 . The method of claim 8 , wherein generating the similarity score dataset includes comparing the semantic type of the first dataset with the semantic type of the second dataset to identify matching and non-matching data fields.

10 . A system comprising:

one or more hardware processors; and

one or more computer-readable mediums storing instructions that, when executed by the one or more hardware processors, cause the system to perform operations comprising:

monitoring a first user interaction with a first dataset on a database system, the first dataset being associated with a first user;

receiving, from the first user, feedback regarding a quality of a results dataset generated from an executed overlap query, the results dataset comprising a joined dataset of data from the first dataset of the first user and data from a second dataset of a second user;

based on the feedback, generating a similarity score dataset indicating a similarity between the first dataset and the second dataset;

generating, using the similarity score dataset, a new set of overlap queries specifying join operations between the first dataset and the second dataset;

executing the new set of overlap queries to materialize a new joined dataset comprising the data from the first dataset and the data from the second dataset, the materializing comprises persisting the new joined dataset as a physical table indexed and stored in the database system for subsequent reuse;

generating a new results dataset by applying the new set of overlap queries to the new joined dataset; and

providing, to the first user, the new results dataset comprising a refinement recommendation for the overlap queries based on additional user feedback, the refinement recommendation comprising a ranked set of proposed modifications to at least one join operation associated with the overlap queries.

11 . The system of claim 10 , wherein each overlap query is generated based on non- matching data between the first dataset and a corresponding dataset of the new joined dataset, the non-matching data comprises non-matching columns between the first dataset and the new joined dataset.

12 . The system of claim 10 , the operations further comprising:

prior to generating the new joined dataset, executing a data cleaning process on the first dataset and the second dataset to ensure data quality and consistency.

13 . The system of claim 10 , wherein the feedback from the first user is received via an interactive user interface and the operations further comprising:

enabling the first user to rate the quality of the results dataset on a predefined scale.

14 . The system of claim 10 , wherein the new set of overlap queries is generated based on a machine learning model that predicts potential utility of overlap datasets based on historical user feedback and query outcomes, the machine learning model is retrained periodically with new user feedback and results datasets to improve accuracy of the new set of overlap queries.

15 . The system of claim 10 , the operations further comprising:

categorizing each dataset in the database system by a semantic type of a data field within each dataset, wherein the semantic type includes at least one of a ZIP code, a date, or a numerical identifier; and

comparing the semantic type of the first dataset with the semantic type of the second dataset to identify matching and non-matching data fields.

16 . A machine-readable storage device embodying instructions that, when executed by a machine, cause the machine to perform operations comprising:

monitoring a first user interaction with a first dataset on a database system, the first dataset being associated with a first user;

receiving, from the first user, feedback regarding a quality of a results dataset generated from an executed overlap query, the results dataset comprising a joined dataset of data from the first dataset of the first user and data from a second dataset of a second user;

based on the feedback, generating a similarity score dataset indicating a similarity between the first dataset and the second dataset;

generating, using the similarity score dataset, a new set of overlap queries specifying join operations between the first dataset and the second dataset;

executing the new set of overlap queries to materialize a new joined dataset comprising the data from the first dataset and the data from the second dataset, the materializing comprises persisting the new joined dataset as a physical table indexed and stored in the database system for subsequent reuse;

generating a new results dataset by applying the new set of overlap queries to the new joined dataset; and

providing, to the first user, the new results dataset comprising a refinement recommendation for the overlap queries based on additional user feedback, the refinement recommendation comprising a ranked set of proposed modifications to at least one join operation associated with the overlap queries.

17 . The machine-readable storage device of claim 16 , wherein each overlap query is generated based on non-matching data between the first dataset and a corresponding dataset of the new joined dataset, the non-matching data comprises non-matching columns between the first dataset and the new joined dataset.

18 . The machine-readable storage device of claim 16 , the operations further comprising:

prior to generating the new joined dataset, executing a data cleaning process on the first dataset and the second dataset to ensure data quality and consistency.

19 . The machine-readable storage device of claim 16 , wherein the feedback from the first user is received via an interactive user interface and the operations further comprising:

enabling the first user to rate the quality of the results dataset on a predefined scale.

20 . The machine-readable storage device of claim 16 , the operations further comprising:

categorizing each dataset in the database system by a semantic type of a data field within each dataset, wherein the semantic type includes at least one of a ZIP code, a date, or a numerical identifier; and

comparing the semantic type of the first dataset with the semantic type of the second dataset to identify matching and non-matching data fields.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: GLICKMAN, MATTHEW J.; KOSTAKIS, ORESTIS; LANGSETH, JUSTIN
To: SNOWFLAKE INC.
Reel/Frame 067363/0923 →
Continuity (2)
Continuation 17804434 · May 27, 2022
Related Publication 20240296162A1 · Sep 5, 2024
References Cited (45)
US 8190628B1 · Yang et al. · 2012 [cited by applicant]
US 11216580B1 · Holboke et al. · 2022 [cited by applicant]
US 11836138B1 · Glickman et al. · 2023 [cited by applicant]
US 12008001B2 · Glickman et al. · 2024 [cited by applicant]
US 20040002956A1 · Chaudhuri et al. · 2004 [cited by applicant]
US 20040064441A1 · Tow · 2004 [cited by applicant]
US 20100037161A1 · Stading · 2010 [cited by examiner]
US 20140229498A1 · Dillon · 2014 [cited by examiner]
US 20160307113A1 · Calapodescu et al. · 2016 [cited by applicant]
US 20170060950A1 · Budhiraja et al. · 2017 [cited by applicant]
US 20170262491A1 · Brewster et al. · 2017 [cited by applicant]
US 20170364561A1 · Wu · 2017 [cited by examiner]
US 20190236192A1 · Zou et al. · 2019 [cited by applicant]
US 20190251184A1 · Shan · 2019 [cited by examiner]
US 20190303801A1 · Horton et al. · 2019 [cited by applicant]
US 20190384762A1 · Hill · 2019 [cited by examiner]
US 20200272651A1 · Luo et al. · 2020 [cited by applicant]
US 20200356873A1 · Nawrocke et al. · 2020 [cited by applicant]
US 20210019318A1 · Leung et al. · 2021 [cited by applicant]
US 20210232592A1 · Liao · 2021 [cited by examiner]
US 20210326369A1 · Roitman · 2021 [cited by examiner]
US 20220067056A1 · Sexton · 2022 [cited by applicant]
US 20220067591A1 · Patel et al. · 2022 [cited by applicant]
US 20220138559A1 · Gangi Reddy et al. · 2022 [cited by applicant]
US 20230385284A1 · Glickman et al. · 2023 [cited by applicant]
US 20230385286A1 · Glickman et al. · 2023 [cited by applicant]
“U.S. Appl. No. 17/804,434, Non Final Office Action mailed Aug. 2, 2022”, 16 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Examiner Interview Summary mailed Nov. 1, 2022”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Response filed Nov. 2, 2022 to Non Final Office Action mailed Aug. 2, 2022”, 18 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Final Office Action mailed Dec. 15, 2022”, 18 pgs. [cited by applicant]
“U.S. Appl. No. 18/162,688, Preliminary Amendment filed Feb. 2, 2023”, 10 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Response filed Mar. 15, 2023 to Final Office Action mailed Dec. 15, 2022”, 15 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Examiner Interview Summary mailed Mar. 17, 2023”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Amendment and Response Filed Mar. 10, 2023 to Final Office Action Mailed Dec. 15, 2022”, 15 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Non Final Office Action mailed Apr. 6, 2023”, 18 pgs. [cited by applicant]
“U.S. Appl. No. 18/162,688, Non Final Office Action mailed Apr. 6, 2023”, 29 pgs. [cited by applicant]
“U.S. Appl. No. 18/162,688, Response filed Jun. 30, 2023 to Non Final Office Action mailed Apr. 6, 2023”, 14 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Response filed Jul. 6, 2023 to Non Final Office Action mailed Apr. 6, 2023”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Examiner Interview Summary mailed Jul. 10, 2023”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 18/162,688, Notice of Allowance mailed Jul. 19, 2023”, 14 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Final Office Action mailed Sep. 20, 2023”, 18 pgs. [cited by applicant]
“U.S. Appl. No. 18/162,688, Notice of Allowance mailed Oct. 12, 2023”, 14 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Response filed Dec. 20, 2023 to Final Office Action mailed Sep. 20, 2023”, 15 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Examiner Interview Summary mailed Dec. 22, 2023”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/804,434, Notice of Allowance mailed Feb. 2, 2024”, 18 pgs. [cited by applicant]