IP Library Granted Patent US 12,373,450
Granted Patent B2
US 12,373,450 · App. 18/353,831 · Granted Jul 29, 2025

Query-time data sessionization and analysis

Inventors: Eric Tschetter (Tokyo, JP); Rohan Garg (New Delhi, IN)
Assignee: Imply Data, Inc.
G06F16/2462G06F16/2264G06F16/2282G06F16/2433G06F16/245G06F16/24539G06F16/24542G06F16/24578G06F16/26G06F16/283G06F16/287
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,450
App. No.
18/353,831
Granted
Jul 29, 2025
Kind
B2
Abstract

A system and method for implementing an iterative query mechanism to facilitate query-time sessionization and analysis of data is disclosed. At least, the method includes determining a table of independent events, determining a query, the query including a parameter specifying a size limit on a sample of sessions, executing the query against the table of independent events, at the time of query execution, processing the table of independent events to reconstruct the sample of sessions approaching the size limit, and at the time of query execution, analyzing the sample of reconstructed sessions to generate a result.

Claims (36)

1. A computer-implemented method comprising:

determining a table of independent events based on a stream of incoming, up-to-date, and raw event-based data, a plurality of data segments comprising the table of independent events being distributed across a cluster of computing devices;

determining a query, the query including a criteria for grouping a plurality of independent events from the table of independent events into a session and a parameter for specifying a size limit on reconstructing a sample of sessions from the table of independent events;

distributing the query in parallel to the cluster of computing devices for executing the query against the table of independent events;

at the time of query execution on each computing device in the cluster of computing devices, processing, in real time, the plurality of data segments comprising the table of independent events to dynamically and randomly reconstruct the sample of sessions using a random sampling method, the sample of sessions matching the criteria and the size limit in the query, a sampling rate for dynamically and randomly reconstructing the sample of sessions being a function of the size limit and the plurality of data segments comprising the table of independent events, the sample of sessions being representative of a whole set of possible sessions from the table of independent events, each session in the sample of sessions including a time-ordered series of independent events occurring within a period of time and mapped to a unique identifier; and

at the time of query execution on each computing device in the cluster of computing devices, analyzing, in real time, the sample of sessions to generate a result.

2. The computer-implemented method of claim 1 , further comprising:

at the time of query execution, processing the table of independent events to aggregate a plurality of independent events into the time-ordered series of independent events and reconstruct the sample of sessions based on the time-ordered series of independent events;

transforming the sample of sessions into a data structure; and

rendering a visualization based on the data structure.

3. The computer-implemented method of claim 2 , wherein determining the query further comprises:

receiving a user interaction in association with the visualization; and

determining the query based on the user interaction.

4. The computer-implemented method of claim 1 , further comprising performing funnel analysis using the result.

5. The computer-implemented method of claim 1 , wherein the criteria includes at least one from a timeframe, a set of users, and a geographical location.

6. The computer-implemented method of claim 1 , wherein the unique identifier includes at least one from a session identifier, client device identifier, and a user identifier.

7. The computer-implemented method of claim 1 , wherein the table of independent events is loaded with data retrieved from one of a streaming data source and a batch data source.

8. A system comprising:

one or more processors; and

a memory, the memory storing instructions, which when executed cause the one or more processors to:

determine a table of independent events based on a stream of incoming, up-to-date, and raw event-based data, a plurality of data segments comprising the table of independent events being distributed across a cluster of computing devices;

determine a query, the query including a criteria for grouping a plurality of independent events from the table of independent events into a session and a parameter for specifying a size limit on reconstructing a sample of sessions from the table of independent claims;

distributing the query in parallel to the cluster of computing devices to execute the query against the table of independent events;

at the time of query execution on each computing device in the cluster of computing devices, process, in real time, the plurality of data segments comprising the table of independent events to dynamically and randomly reconstruct the sample of sessions using a random sampling method, the sample of sessions matching the criteria and the size limit in the query, a sampling rate for dynamically and randomly reconstructing the sample of sessions being a function of the size limit and the plurality of data segments comprising the table of independent events, the sample of sessions being representative of a whole set of possible sessions from the table of independent events, each session in the sample of sessions including a time-ordered series of independent events occurring within a period of time and mapped to a unique identifier; and

at the time of query execution on each computing device in the cluster of computing devices, analyze, in real time, the sample of sessions to generate a result.

9. The system of claim 8 , wherein the instructions further cause the one or more processors to:

at the time of query execution, process the table of independent events to aggregate a plurality of independent events into the time-ordered series of independent events and reconstruct the sample of sessions based on the time-ordered series of independent events;

transform the sample of sessions into a data structure; and

render a visualization based on the data structure.

10. The system of claim 9 , wherein to determine the query, the instructions further cause the one or more processors to:

receive a user interaction in association with the visualization; and

determine the query based on the user interaction.

11. The system of claim 8 , wherein the instructions further cause the one or more processors to perform funnel analysis using the result.

12. The system of claim 8 , wherein the criteria includes at least one from a timeframe, a set of users, and a geographical location.

13. The system of claim 8 , wherein the unique identifier includes at least one from a session identifier, client device identifier, and a user identifier.

14. The system of claim 8 , wherein the table of independent events is loaded with data retrieved from one of a streaming data source and a batch data source.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2023
From: TSCHETTER, ERIC; GARG, ROHAN
To: IMPLY DATA, INC.
Reel/Frame 064719/0738 →
Continuity (2)
Provisional Application 63389443 · Jul 15, 2022
Related Publication 20240020311A1 · Jan 18, 2024
References Cited (38)
US 10108648B2 · Rajan · 2018 [cited by examiner]
US 10324917B2 · Tutuk · 2019 [cited by examiner]
US 11204898B1 · Waas · 2021 [cited by examiner]
US 11620541B1 · Ghosh · 2023 [cited by examiner]
US 20040225639A1 · Jakobsson · 2004 [cited by examiner]
US 20060173878A1 · Bley · 2006 [cited by examiner]
US 20070174256A1 · Morris · 2007 [cited by examiner]
US 20070271145A1 · Vest · 2007 [cited by applicant]
US 20110055198A1 · Mitchell · 2011 [cited by examiner]
US 20110319080A1 · Bienas et al. · 2011 [cited by applicant]
US 20120084287A1 · Lakshminarayan · 2012 [cited by examiner]
US 20140156683A1 · de Castro Alves · 2014 [cited by examiner]
US 20150066966A1 · O'Donnell · 2015 [cited by examiner]
US 20150088823A1 · Chen · 2015 [cited by examiner]
US 20150248464A1 · Desai · 2015 [cited by examiner]
US 20170286496A1 · Pang · 2017 [cited by examiner]
US 20170286532A1 · Horowitz · 2017 [cited by examiner]
US 20170364558A1 · Thorne · 2017 [cited by examiner]
US 20180018388A1 · Mulla · 2018 [cited by examiner]
US 20180121856A1 · Song · 2018 [cited by examiner]
US 20190138639A1 · Pal et al. · 2019 [cited by applicant]
US 20200201927A1 · Narula · 2020 [cited by examiner]
US 20210042302A1 · Miao · 2021 [cited by examiner]
US 20210287250A1 · Wang et al. · 2021 [cited by applicant]
US 20210397616A1 · Iranmanesh · 2021 [cited by examiner]
US 20230118230A1 · Tschetter · 2023 [cited by examiner]
US 20230229659A1 · Beresniewicz · 2023 [cited by examiner]
EP 2843567B1 · 2017 [cited by examiner]
WO WO2006009822A2 · 2006 [cited by examiner]
WO WO2007022560A1 · 2007 [cited by examiner]
WO WO2008043082A2 · 2008 [cited by examiner]
WO WO2014169265A1 · 2014 [cited by examiner]
Stefan Esser et al., “Multi-Dimensional Event Data in Graph Databases”, Journal on Data Semantics (2021) 10:109-141 Published online: May 27, 2021. [cited by examiner]
Yuanzhe Hao et al., “TS-Benchmark: A Benchmark for Time Series Databases”, 2021 IEEE 37th International Conference on Data Engineering (ICDE), Jun. 2021. [cited by examiner]
T. Johnson et al., “Query-Aware Sampling for Data Streams”, 2007 IEEE 23rd International Conference on Data Engineering Workshop (2007, pp. 664-673). [cited by examiner]
Ushijima, K et al., “SUPRA: a sampling-query optimization method for large-scale OLAP”, Proceedings Ninth International Workshop on Database and Expert Systems Applications (Cat. No. 98EX130) (1998, pp. 232-237). [cited by examiner]
Massimo Melucci, “Impact of Query Sample Selection Bias on Information Retrieval System Ranking”, 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA) (Dec. 2016, pp. 341-350). [cited by examiner]
PCT International Search Report and Written Opinion, Imply Data, Inc. PCT/US23/27949, filing date Jul. 17, 2023, mail date Oct. 24, 2023, 11 pgs. [cited by applicant]