IP Library › Granted Patent US 12,216,648
Granted Patent B1
US 12,216,648 · App. 18/393,324 · Granted Feb 4, 2025

Row-order dependent dataframe workloads

Inventors: Srilakshmi Chintala (Seattle, WA); Jianzhun Du (Kirkland, WA); Naresh Kumar (Santa Clara, CA); Srinath Shankar (Belmont, CA); Leonhard Franz Spiegelberg (San Francisco, CA); Eric Shawn Vandenberg (Saratoga, CA); Andong Zhan (San Mateo, CA); Yun Zou (Sunnyvale, CA)
Assignee: Snowflake Inc.
G06F16/244G06F16/2453G06F16/256
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,216,648
App. No.
18/393,324
Filed
Dec 21, 2023
Granted
Feb 4, 2025
Kind
B1
Examiner
VY, HUNG T
Art Unit
2163
USPC
707/715
Abstract

A method includes receiving instructions to perform order-dependent DataFrame operation on data. In response to receiving the instructions, the framework analyzes the instructions to identify the order-dependent DataFrame operation, and generates an executable query corresponding to the identified order-dependent DataFrame operation. The framework executes the generated executable query on the data stored in the first database in the cloud data platform, creates a row position column that generates for each row a unique integer identifier, reproducible across multiple sessions and/or queries for the same underlying data to make the data accessible via positional indexing, and assigns each row of the first database a unique row index value based on the row position column and the row position column order. The framework orders a result from performing the order-dependent DataFrame operation based on the unique row index value assigned to each row of the first database, and returns the ordered result to the user.

Claims (92)

1. A method comprising:

receiving, by at least one hardware processor, instructions to perform an order-dependent DataFrame operation on data stored in a first database in a cloud data platform, the instructions specified using code authored in a programming language for executing the order-dependent DataFrame operation within the cloud data platform;

analyzing the instructions to identify the order-dependent DataFrame operation;

generating an executable query corresponding to the identified order-dependent DataFrame operation;

executing the generated executable query on the data stored in the first database in the cloud data platform;

creating a row position column that generates a row position column order to make the data accessible via positional indexing;

assigning each row of the first database a unique row position value based on the row position column and the row position column order;

ordering a result from performing the order-dependent DataFrame operation based on the unique row position value assigned to each row of the first database; and

returning, to a user, the ordered result.

2. The method of claim 1 , wherein the order-dependent DataFrame operation comprises at least one of: displaying a specified number of rows, displaying a specified set of rows based on their indices, appending rows, or aggregating rows.

3. The method of claim 2 , wherein displaying the specified number of rows further comprises at least one of displaying the rows from a beginning of the first database or displaying the rows from an end of the first database.

4. The method of claim 1 , wherein assigning each row of the first database the unique row position value comprises:

determining an offset associated with a data file; and

assigning each row of the first database the unique row position value based on the offset of the data file and the unique row position value and a position of each row within the data file.

5. The method of claim 4 , wherein ordering the result comprises:

ordering a specified number of rows based on the unique row position value assigned to each row; and

modifying the generated executable query to include an ORDER BY clause corresponding to the unique row position value assigned to each row of the first database.

6. The method of claim 1 , wherein the order-dependent DataFrame operation comprises displaying rows from the first database based on specified row indices, and wherein ordering the result comprises ordering rows corresponding to the specified row indices based on the unique row position value assigned to each row.

7. The method of claim 1 , further comprising:

appending rows to the first database;

ordering the first database with the appended rows based on the unique row position value assigned to each row of the first database; and

assigning each of the appended rows with the unique row position value.

8. The method of claim 1 , further comprising:

performing an aggregation operation on the first database; and

ordering rows of an aggregation result based on values of grouping columns from the first database.

9. The method of claim 1 , further comprising:

joining the first database with a second database; and

ordering rows of a join result based on the unique row position value assigned to the rows of the first database and the rows of the second database.

10. The method of claim 9 , wherein ordering the rows of the join result further comprises:

prioritizing ordering the rows based on the unique row position value assigned to the first database over the unique row position value assigned to the second database.

11. A system comprising:

one or more hardware processors; and

at least one memory storing instructions that, when executed by the one or more hardware processors, cause the system to perform operations comprising:

receiving, by at least one hardware processor, instructions to perform an order-dependent DataFrame operation on data stored in a first database in a cloud data platform, the instructions specified using code authored in a programming language for executing the order-dependent DataFrame operation within the cloud data platform;

analyzing the instructions to identify the order-dependent DataFrame operation;

generating an executable query corresponding to the identified order-dependent DataFrame operation;

executing the generated executable query on the data stored in the first database in the cloud data platform;

creating a row position column that generates a row position column order to make the data accessible via positional indexing;

assigning each row of the first database a unique row position value based on the row position column and the row position column order;

ordering a result from performing the order-dependent DataFrame operation based on the unique row position value assigned to each row of the first database; and

returning, to a user, the ordered result.

12. The system of claim 11 , wherein the order-dependent DataFrame operation comprises at least one of: displaying a specified number of rows, displaying a specified set of rows based on their indices, appending rows, or aggregating rows.

13. The system of claim 12 , wherein displaying the specified number of rows further comprises at least one of displaying the rows from a beginning of the first database or displaying the rows from an end of the first database.

14. The system of claim 11 , wherein assigning each row of the first database the unique row position value comprises:

determining an offset associated with a data file; and

assigning each row of the first database the unique row position value based on the offset of the data file and the unique row position value and a position of each row within the data file.

15. The system of claim 14 , wherein ordering the result comprises:

ordering a specified number of rows based on the unique row position value assigned to each row; and

modifying the generated executable query to include an ORDER BY clause corresponding to the unique row position value assigned to each row of the first database.

16. The system of claim 11 , wherein the order-dependent DataFrame operation comprises displaying rows from the first database based on specified row indices, and wherein ordering the result comprises ordering rows corresponding to the specified row indices based on the unique row position value assigned to each row.

17. The system of claim 11 , the operations comprising:

appending rows to the first database;

ordering the first database with the appended rows based on the unique row position value assigned to each row of the first database; and

assigning each of the appended rows with the unique row position value.

18. The system of claim 11 , the operations comprising:

performing an aggregation operation on the first database; and

ordering rows of an aggregation result based on values of grouping columns from the first database.

19. The system of claim 11 , the operations comprising:

joining the first database with a second database; and

ordering rows of a join result based on the unique row position value assigned to the rows of the first database and the rows of the second database.

20. The system of claim 19 , wherein ordering the rows of the join result further comprises:

prioritizing ordering the rows based on the unique row position value assigned to the first database over the unique row position value assigned to the second database.

21. A machine-storage medium embodying instructions that, when executed by a machine, cause the machine to perform operations comprising:

receiving, by at least one hardware processor, instructions to perform an order-dependent DataFrame operation on data stored in a first database in a cloud data platform, the instructions specified using code authored in a programming language for executing the order-dependent DataFrame operation within the cloud data platform;

analyzing the instructions to identify the order-dependent DataFrame operation;

generating an executable query corresponding to the identified order-dependent DataFrame operation;

executing the generated executable query on the data stored in the first database in the cloud data platform;

creating a row position column that generates a row position column order to make the data accessible via positional indexing;

assigning each row of the first database a unique row position value based on the row position column and the row position column order;

ordering a result from performing the order-dependent DataFrame operation based on the unique row position value assigned to each row of the first database; and

returning, to a user, the ordered result.

22. The machine-storage medium of claim 21 , wherein the order-dependent DataFrame operation comprises at least one of: displaying a specified number of rows, displaying a specified set of rows based on their indices, appending rows, or aggregating rows.

23. The machine-storage medium of claim 22 , wherein displaying the specified number of rows further comprises at least one of displaying the rows from a beginning of the first database or displaying the rows from an end of the first database.

24. The machine-storage medium of claim 21 , wherein assigning each row of the first database the unique row position value comprises:

determining an offset associated with a data file; and

assigning each row of the first database the unique row position value based on the offset of the data file and the unique row position value and a position of each row within the data file.

25. The machine-storage medium of claim 24 , wherein ordering the result comprises:

ordering a specified number of rows based on the unique row position value assigned to each row; and

modifying the generated executable query to include an ORDER BY clause corresponding to the unique row position value assigned to each row of the first database.

26. The machine-storage medium of claim 21 , wherein the order-dependent DataFrame operation comprises displaying rows from the first database based on specified row indices, and wherein ordering the result comprises ordering rows corresponding to the specified row indices based on the unique row position value assigned to each row.

27. The machine-storage medium of claim 21 , wherein the operations comprise:

appending rows to the first database;

ordering the first database with the appended rows based on the unique row position value assigned to each row of the first database; and

assigning each of the appended rows with the unique row position value.

28. The machine-storage medium of claim 21 , wherein the operations comprise:

performing an aggregation operation on the first database; and

ordering rows of an aggregation result based on values of grouping columns from the first database.

29. The machine-storage medium of claim 21 , wherein the operations comprise:

joining the first database with a second database; and

ordering rows of a join result based on the unique row position value assigned to the rows of the first database and the rows of the second database.

30. The machine-storage medium of claim 29 , wherein ordering the rows of the join result further comprises:

prioritizing ordering the rows based on the unique row position value assigned to the first database over the unique row position value assigned to the second database.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 12, 2024
From: CHINTALA, SRILAKSHMI; DU, JIANZHUN; KUMAR, NARESH; SHANKAR, SRINATH; SPIEGELBERG, LEONHARD FRANZ; VANDENBERG, ERIC SHAWN; ZHAN, ANDONG; ZOU, YUN
To: SNOWFLAKE INC.
Reel/Frame 066439/0194 →
Continuity (1)
Provisional Application 63583525 · Sep 18, 2023
References Cited (20)
US 20200097325A1 · Vadapandeshwara · 2020 [cited by examiner]
US 20200327252A1 · McFall · 2020 [cited by examiner]
U.S. Appl. No. 18/393,279, filed Dec. 21, 2023, Dataframe Workloads Using Read-Only Data Snapshots. [cited by applicant]
“Best Practices—Pyspark Master Documentation”, [Online]. Retrieved from the Internet: <https://spark.apache.org/docs/latest/api/python/user_guide/pandas_on_spark/best_practice s.html#best-practices>, (Accessed on Jul. 9… [cited by applicant]
“Categorical data—Pandas 2.2”, [Online]. Retrieved from the Internet: <https://pandas.pydata.org/docs/user_guide/categorical.html>, (Accessed on Jul. 9, 2024), 41 pages. [cited by applicant]
“Duplicate Labels—pandas 2.2”, [Online]. Retrieved from the Internet: <https://pandas.pydata.org/docs/user_guide/duplicates.html>, (Accessed on Jul. 9, 2024), 10 pages. [cited by applicant]
“Iris—UCI Machine Learning Repository”, [Online]. Retrieved from the Internet: <https://archive.ics.uci.edu/dataset/53/iris>, (Accessed on Jul. 9, 2024), 5 pages. [cited by applicant]
“MultiIndex/advanced indexing—pandas 2.2.pdf”, [Online]. Retrieved from the Internet: <https://pandas.pydata.org/docs/user_guide/advanced.html>, (Accessed on Jul. 9, 2024), 47 pages. [cited by applicant]
“New Features in 19c Release Updates”, [Online]. Retrieved from the Internet: <https://docs.oracle.com/en/database/oracle/oracle-database/19/newft/new-features-19c-release-updates.html#GUID-ACODAFEE-2DFF-4EA8-A9B9-DB0E0… [cited by applicant]
“Options and settings—PySpark master documentation”, [Online]. Retrieved from the Internet: <https://spark.apache.org/docs/latest/api/python/user_guide/pandas_on_spark/options.html#default-index-type>, (Accessed on Jul.… [cited by applicant]
“Pandas.ArrowDtype—pandas 2.2”, [Online]. Retrieved from the Internet: <https://pandas.pydata.org/docs/reference/api/pandas.ArrowDtype.html>, (Accessed on Jul. 9, 2024), 2 pages. [cited by applicant]
“Pandas.CategoricalDtype—pandas 2.2”, [Online]. Retrieved from the Internet: <https:/pandas.pydata.org/docs/reference/api/pandas.CategoricalDtype.html>, (Accessed on Jul. 9, 2024), 2 pages. [cited by applicant]
“Pandas.DataFrame.pivot—pandas 2.2”, [Online]. Retrieved from the Internet: </https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.pivot.html>, (Accessed on Jul. 9, 2024), 4 pages. [cited by applicant]
“Pandas.DataFrame.pivot_table—pandas 2.2”, [Online]. Retrieved from the Internet: <https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.pivot_table.html>, (Accessed on Jul. 9, 2024), 4 pages. [cited by applicant]
“Pandas.DataFrame.transpose—pandas 2.2”, [Online]. Retrieved from the Internet: <https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.transpose.html>, (Accessed on Jul. 9, 2024), 3 pages. [cited by applicant]
“Pandas.SparseDtype—pandas 2.2”, [Online]. Retrieved from the Internet: <https://pandas.pydata.org/docs/reference/api/pandas.SparseDtype.html>, (Accessed on Jul. 9, 2024), 2 pages. [cited by applicant]
“Pandas.Timestamp”, [Online]. Retrieved from the Internet: <https://pandas.pydata.org/docs/reference/api/pandas.Timestamp.html>, (Accessed on Jul. 9, 2024), 7 pages. [cited by applicant]
“Snapshot Concepts & Architecture”, [Online]. Retrieved from the Internet: <https://docs.oracle.com/cd/A87860_01/doc/server.817/a76959/mview.htm>, (Accessed on Jul. 9, 2024), 28 pages. [cited by applicant]
McKinney, Wes, “Apache Arrow and the “10 Things I Hate About pandas””, [Online]. Retrieved from the Internet: <Iris—UCI Machine Learning Repository>, (Accessed on Jul. 9, 2024), 8 pages. [cited by applicant]
Modin, “System Architecture”, [Online]. Retrieved from the Internet: <https://modin.readthedocs.io/en/latest/development/architecture.html>, (Accessed online Jul. 10, 2024), 11 pages. [cited by applicant]
Cited By (2)
US 12,591,579 US 12,688,009