IP Library Granted Patent US 12,591,579
Granted Patent B2
US 12,591,579 · App. 18/754,315 · Granted Mar 31, 2026

Aggregation operations in a distributed database

Inventors: Ashok Anand (Bengaluru, IN); Ambareesh Sreekumaran Nair Jayakumari (Cupertino, CA); Prateek Gaur (San Jose, CA); Donko Donjerkovic (San Mateo, CA)
Assignee: ThoughtSpot, Inc.
G06F16/24556G06F16/2282G06F16/248G06F16/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,579
App. No.
18/754,315
Granted
Mar 31, 2026
Kind
B2
Abstract

A distributed database that includes multiple database instances receives a data-query that includes an aggregation clause on a first column of a table. The table is partitioned into shards according to a sharding criterion based on the first column such that all rows having the same value for the first column are included in the same shard. The shards are distributed to the multiple database instances. Respective intermediate results are received from at least some of the database instances. Each intermediate result received from a respective database instance that includes a respective shard aggregates values of the first column in the respective shard. The respective intermediate results are combined to obtain a final result of the data-query. The final result is then output.

Claims (40)

1 . A method, comprising:

receiving, at a distributed database that includes multiple database instances, a data-query including an aggregation clause on a first column of a table that is partitioned into shards according to a sharding criterion based on the first column such that all rows having a same value for the first column are included in a same shard, wherein the shards are distributed to the multiple database instances;

receiving, from at least some of the database instances, respective intermediate results, wherein each respective intermediate result received from a respective database instance that includes a respective shard is an aggregation of values of the first column in the respective shard;

combining the respective intermediate results to obtain a combined result of the data-query; and

outputting the combined result.

2 . The method of claim 1 , wherein combining the respective intermediate results to obtain the combined result of the data-query comprises:

obtaining the combined result by using a union operation to remove duplicate values from the respective intermediate results.

3 . The method of claim 1 , wherein the aggregation clause includes a distinct count of the first column.

4 . The method of claim 3 , wherein the combined result includes a count of distinct values of the first column.

5 . The method of claim 1 , wherein the sharding criterion includes at least one column of the table.

6 . The method of claim 1 , wherein the data-query includes a sampling clause.

7 . The method of claim 1 , wherein the combined result includes a sum of aggregated values from the respective intermediate results.

8 . The method of claim 1 , wherein the data-query includes a grouping clause on a second column of the table.

9 . The method of claim 8 , wherein the respective intermediate results are aggregated based on the values of the second column.

10 . The method of claim 9 , further comprising:

calculating a distinct count of the first column for each group of the second column.

11 . The method of claim 1 , wherein the respective intermediate results include minimum values of the first column for each shard.

12 . The method of claim 1 , wherein the respective intermediate results include maximum values of the first column for each shard.

13 . A system, comprising:

one or more memories; and

one or more processors, the one or more processors configured to execute instructions stored in the one or more memories to:

receive, at a distributed database that includes multiple database instances, a data-query including an aggregation clause on a first column of a table that is partitioned into shards according to a sharding criterion based on the first column such that all rows having a same value for the first column are included in a same shard, wherein the shards are distributed to the multiple database instances;

receive, from at least some of the database instances, respective intermediate results, wherein each respective intermediate result received from a respective database instance that includes a respective shard is an aggregation of values of the first column in the respective shard;

combine the respective intermediate results to obtain a combined result of the data-query; and

output the combined result.

14 . The system of claim 13 , wherein the instructions to combine the respective intermediate results to obtain the combined result of the data-query comprise instructions to:

obtain the combined result by using a union operation to remove duplicate values from the respective intermediate results.

15 . The system of claim 13 , wherein the one or more processors configured to execute instructions stored in the one or more memories to:

calculate a distinct count of the first column for each group of a second column.

16 . The system of claim 13 , wherein the respective intermediate results include minimum values of the first column for each shard or maximum values of the first column for the each shard.

17 . One or more non-transitory computer readable media storing instructions operable to cause one or more processors to perform operations comprising:

receiving, at a distributed database that includes multiple database instances, a data-query including an aggregation clause on a first column of a table that is partitioned into shards according to a sharding criterion based on the first column such that all rows having a same value for the first column are included in a same shard, wherein the shards are distributed to the multiple database instances;

receiving, from at least some of the database instances, respective intermediate results, wherein each respective intermediate result received from a respective database instance that includes a respective shard is an aggregation of values of the first column in the respective shard;

combining the respective intermediate results to obtain a combined result of the data-query; and

outputting the combined result.

18 . The one or more non-transitory computer readable media of claim 17 , wherein combining the respective intermediate results to obtain the combined result of the data-query comprises:

obtaining the combined result by using a union operation to remove duplicate values from the respective intermediate results.

19 . The one or more non-transitory computer readable media of claim 17 , wherein the operations further comprise:

calculating a distinct count of the first column for each group of a second column.

20 . The one or more non-transitory computer readable media of claim 17 , wherein the respective intermediate results include minimum values of the first column for each shard or maximum values of the first column for the each shard.

Assignments (2)
SECURITY INTEREST Recorded Mar 7, 2025
From: THOUGHTSPOT, INC.; THOUGHTSPOT, LLC
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 070442/0499 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2024
From: ANAND, ASHOK; JAYAKUMARI, AMBAREESH SREEKUMARAN NAIR; GAUR, PRATEEK; DONJERKOVIC, DONKO
To: THOUGHT SPOT, INC.
Reel/Frame 067841/0480 →
Continuity (3)
Continuation 18333688 · Jun 13, 2023
Continuation 17214247 · Mar 26, 2021
Related Publication 20240354303A1 · Oct 24, 2024
References Cited (26)
US 11487668B2 · Anand · 2022 [cited by examiner]
US 11720570B2 · Anand · 2023 [cited by examiner]
US 11748264B1 · Anand · 2023 [cited by examiner]
US 12210572B2 · Saupe · 2025 [cited by examiner]
US 12216648B1 · Chintala · 2025 [cited by examiner]
US 12321361B2 · Cai · 2025 [cited by examiner]
US 12400247B2 · Saito · 2025 [cited by examiner]
US 12430311B1 · Wang · 2025 [cited by examiner]
US 20190102436A1 · Bishnoi · 2019 [cited by examiner]
US 20220092069A1 · Hartsing · 2022 [cited by examiner]
US 20220309067A1 · Anand · 2022 [cited by examiner]
US 20220318147A1 · Anand · 2022 [cited by examiner]
US 20230401210A1 · Anand · 2023 [cited by examiner]
US 20240126760A1 · Lui · 2024 [cited by examiner]
US 20240411815A1 · Saupe · 2024 [cited by examiner]
US 20240427786A1 · Nhan · 2024 [cited by examiner]
US 20240427790A1 · Cai · 2024 [cited by examiner]
US 20250094386A1 · Higgins · 2025 [cited by examiner]
US 20250225133A1 · Han · 2025 [cited by examiner]
US 20250245232A1 · Lui · 2025 [cited by examiner]
US 20250272306A1 · Cai · 2025 [cited by examiner]
US 20250307534A1 · Mi · 2025 [cited by examiner]
US 20250330395A1 · Ahmed · 2025 [cited by examiner]
US 20250335455A1 · Pineda · 2025 [cited by examiner]
Unsupervised Clustering for Sharding Key Formulation and the Effects of Aggregation Computations (Year: 2019). [cited by examiner]
Using Aggregation and Dynamic Queries for Exploring Large Data Sets (Year: 1994). [cited by examiner]