IP Library Granted Patent US 12,468,686
Granted Patent B2
US 12,468,686 · App. 18/304,795 · Granted Nov 11, 2025

Annotating datasets without redundant copying

Inventors: George Aleksandrovich (Hoffman Estates, IL); Allie K. Watfa (Urbana, IL); Robin Sahner (Urbana, IL); Mike Pippin (Sunnyvale, CA)
Assignee: YAHOO ASSETS LLC
G06F16/2379G06F7/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,686
App. No.
18/304,795
Granted
Nov 11, 2025
Kind
B2
Abstract

Disclosed embodiments are methods, apparatuses, and computer-readable media for annotating distributed data without redundant data copying. In one embodiment, a method is disclosed comprising reading a raw dataset, the raw dataset comprising a first set of columns and a first set of rows; generating an annotation dataset, the annotation dataset comprising a second set of columns and a second set of rows; assigning row identifiers to each row in the second set of rows, the row identifiers aligning the second set of rows with the first set of rows based on the underlying storage of the raw dataset and annotation dataset; and writing the annotation dataset to a distributed storage medium.

Claims (45)

1 . A method comprising:

reading, by a device, a dataset comprising data stored in an unordered manner in a first set of columns and a first set of rows in a database;

determining, by the device, an attribute for the data of the dataset;

partitioning, by the device, the data based on the determined attribute;

reordering, by the device, the partitioned data based on boundaries of each partition of the data, the reordering performed via a stripe-based alignment of the partitioned data that results in a reduction of a quantity of data files associated with the dataset;

modifying, by the device, the dataset by flattening the reordered data; and

storing, by the device, the modified dataset in the database.

2 . The method of claim 1 , wherein the reordering comprises modifying a configuration of the data within the dataset respective to the first set of columns and first set of rows.

3 . The method of claim 1 , wherein the reordering comprises grouping the data within the dataset based at least on the determined attribute.

4 . The method of claim 1 , further comprising:

sorting, based on attributes of the data, the data within the dataset, the sorting being contingent upon the reordering of the data within the database based on the determined attribute.

5 . The method of claim 4 , wherein the sorting is performed in relation to the first set of rows.

6 . The method of claim 4 , wherein the flattening of the reordered data is based on the sorted data.

7 . The method of claim 1 , further comprising:

formatting the flattened reordered data into a first format, the formatting enabling the storage of the modified dataset.

8 . A non-transitory computer-readable storage medium tangibly encoded with computer-executable instructions, that when executed by a device, perform a method comprising:

reading, by a device, a dataset comprising data stored in an unordered manner in a first set of columns and a first set of rows in a database;

determining, by the device, an attribute for the data of the dataset;

partitioning, by the device, the data based on the determined attribute;

reordering, by the device, the partitioned data based on boundaries of each partition of the data, the reordering performed via a stripe-based alignment of the partitioned data that results in a reduction of a quantity of data files associated with the dataset;

modifying, by the device, the dataset by flattening the reordered data; and

storing, by the device, the modified dataset in the database.

9 . The non-transitory computer-readable storage medium of claim 8 , wherein the reordering comprises modifying a configuration of the data within the dataset respective to the first set of columns and first set of rows.

10 . The non-transitory computer-readable storage medium of claim 8 , wherein the reordering comprises grouping the data within the dataset based at least on the determined attribute.

11 . The non-transitory computer-readable storage medium of claim 8 , further comprising:

sorting, based on attributes of the data, the data within the dataset, the sorting being contingent upon the reordering of the data within the database based on the determined attribute.

12 . The non-transitory computer-readable storage medium of claim 11 , wherein the sorting is performed in relation to the first set of rows.

13 . The non-transitory computer-readable storage medium of claim 11 , wherein the flattening of the reordered data is based on the sorted data.

14 . The non-transitory computer-readable storage medium of claim 8 , further comprising:

formatting the flattened reordered data into a first format, the formatting enabling the storage of the modified dataset.

15 . A device comprising:

a processor configured to:

read a dataset comprising data stored in an unordered manner in a first set of columns and a first set of rows in a database;

determine an attribute for the data of the dataset;

partition the data based on the determined attribute;

reorder the partitioned data based on boundaries of each partition of the data, the reordering performed via a stripe-based alignment of the partitioned data that results in a reduction of a quantity of data files associated with the dataset;

modify the dataset by flattening the reordered data; and

store the modified dataset in the database.

16 . The device of claim 15 , wherein the reordering comprises modifying a configuration of the data within the dataset respective to the first set of columns and first set of rows.

17 . The device of claim 15 , wherein the reordering comprises grouping the data within the dataset based at least on the determined attribute.

18 . The device of claim 15 , wherein the processor is further configured to:

sort, based on attributes of the data, the data within the dataset, the sorting being contingent upon the reordering of the data within the database based on the determined attribute, wherein the flattening of the reordered data is based on the sorted data.

19 . The device of claim 18 , wherein the sorting is performed in relation to the first set of rows.

20 . The device of claim 15 , wherein the processor is further configured to:

format the flattened reordered data into a first format, the formatting enabling the storage of the modified dataset.

Assignments (4)
SUPPLEMENTAL PATENT SECURITY AGREEMENT Recorded Sep 17, 2025
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 072915/0540 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2023
From: ALEKSANDROVICH, GEORGE; WATFA, ALLIE K.; SAHNER, ROBIN; PIPPIN, MIKE
To: OATH INC.
Reel/Frame 063695/0983 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2023
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 063698/0340 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2023
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 063698/0644 →