IP Library › Granted Patent US 12,265,509
Granted Patent B2
US 12,265,509 · App. 18/186,642 · Granted Apr 1, 2025

Loading data files into hierarchical storage system

Inventors: William H. Bishop (Raleigh, NC); Christopher A. Connor (Youngsville, NC); Sloan H. Holliday (Apex, NC)
Assignee: Optum, Inc.
G06F16/185G06F16/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,509
App. No.
18/186,642
Granted
Apr 1, 2025
Kind
B2
Abstract

A data file comprising a plurality of rows, each of the rows includes at least a first column and a second column, the first column contains a first-level resource ID that identifies a first-level resource, the second column provides information regarding the first-level resource. For each respective unique first-level resource ID, the processing circuitry identifies a row set for the respective first-level resource ID, performs a sequential deduplication process on the row set, and enqueues remaining rows in a queue for a thread assigned to the respective first-level resource ID. For each row enqueued in the queue for the thread, the thread dequeues the row from the queue for the respective thread, requests creation of a second-level resource that stores a version of the data element contained in the second column of the dequeued row, and requests creation of relationship data for the second-level resource.

Claims (97)

1. A computing system comprising:

a memory configured to store a data file comprising a plurality of rows each including a first column and a second column, wherein:

for each row of the plurality of rows, the first column of the row contains a first-level resource identifier (ID), of a plurality of first-level resource IDs, that identifies a respective first-level resource, of a plurality of first-level resources, and the second column contains a data element, of a plurality of data elements, that provides information regarding the respective first-level resource; and

one or more processors communicatively coupled to the memory, the one or more processors configured to:

initiate one or more threads assigned to one or more unique first-level resource IDs of the plurality of first-level resource IDs; and

for each respective unique first-level resource ID of the one or more unique first-level resource IDs:

identify a first row set that comprises one or more of the plurality of rows that contain the respective unique first-level resource ID; and

for at least a first row of the first row set:

determine whether a first data element in the second column of the first row is different from a second data element in the second column of a predecessor row from the first row set that precedes the first row;

in response to determining that the first data element is different from the second data element, enqueue the first row in a first queue for a first thread, from the one or more threads, assigned to the respective unique first-level resource ID;

dequeue the first row from the first queue for the first thread; and

cause a server to create a second-level resource that stores a version of the first data element.

2. The computing system of claim 1 , wherein the one or more processors cause the server to create the second-level resource by sending a request to the server, causing the server to create, in a hierarchical data store, the second-level resource.

3. The computing system of claim 1 , wherein the one or more processors are further configured to send a request causing the server to create, in a hierarchical data store, relationship data for the second-level resource that specify that the second-level resource is a parent of a relevant first-level resource for the second-level resource, wherein the relevant first-level resource for the second-level resource is identified by a first-level resource ID, of the plurality of first-level resource IDs, contained in the first column of the first row.

4. The computing system of claim 3 , wherein:

the first thread is configured to:

obtain a first logical ID of the relevant first-level resource for the second-level resource and a first version ID of the relevant first-level resource for the second-level resource; and

obtain, from the server, a second logical ID of the second-level resource and a second version ID of the second-level resource, and

wherein the relationship data for the second-level resource specifies that the second-level resource is the parent of the relevant first-level resource for the second-level resource by specifying: (1) the second logical ID and the second version ID and (2) the first logical ID and the first version ID.

5. The computing system of claim 4 , wherein:

the second-level resource is a first second-level resource,

the first thread is further configured to store a row number of the first row, the second logical ID, and the second version ID in a cache,

each of the plurality of rows further includes a third column that contains a third-level resource ID that identifies a third-level resource,

a set of unique third-level resource IDs is defined as including only unique third-level resource IDs contained in the third column of the plurality of rows,

for each respective third-level resource ID in the set of unique third-level resource IDs, the one or more threads include a second thread assigned to the respective third-level resource ID, and

the one or more processors are further configured to:

identify a second row set for the respective third-level resource ID that comprises ones of the plurality of rows that contain the respective third-level resource ID;

for at least a second row of the second row set:

determine whether a third data contained in any subordinate column of the second row is different from a fourth data element contained in a corresponding subordinate column of a predecessor row for the second row, wherein the predecessor row for the second row precedes the second row in the second row set;

in response to determining that the third data element is different from the fourth data element, enqueue the second row in a second queue for the second thread, and

wherein the second thread is configured to, for each third row enqueued in the second queue:

dequeue the third row from the second queue;

use a row number of the third row to search the cache to obtain a third logical ID of a second-level resource and a third version ID of the second second-level resource;

send a request causing the server to create, in the hierarchical data store, a third-level resource that specifies the third-level resource ID contained in the third column of the third row;

obtain, from the server, a fourth logical ID and a fourth version ID of the third-level resource; and

send a request causing the server to create, in the hierarchical data store, relationship data for the third-level resource that specify that the third-level resource is a parent of the second second-level resource by specifying the fourth logical ID and the fourth version ID and the third logical ID and the third version ID.

6. The computing system of claim 5 , wherein, for a fourth row of the second row set, the one or more processors are further configured to, in response to determining that a fifth data element in the second column of the fourth row is not different from a sixth data element in the second column of the fourth row:

obtain a fifth logical ID of a predecessor second-level resource and a fifth version ID of the predecessor second-level resource, the predecessor second-level resource being a second-level resource created for the predecessor row for the fourth row of the second row set; and

store a row number of the fourth row of the second row set, the fifth logical ID, and the fifth version ID in the cache.

7. The computing system of claim 1 , wherein two or more of the threads that are assigned to the unique first-level resource IDs operate in parallel.

8. The computing system of claim 1 , wherein the server is a Fast Healthcare Interoperability Resources (FHIR) server.

9. The computing system of claim 1 , wherein:

the plurality of first-level resource IDs includes a patient identifier, and the plurality of data elements provides information regarding one or more of patient first name, patient last name, patient address, or patient birth date, or

the plurality of first-level resource IDs includes a provider ID, and the plurality of data elements provides information regarding one or more of a provider name, or provider address.

10. A computer-implemented method comprising:

obtaining, by one or more processors, a data file comprising a plurality of rows each including a first column and a second column, wherein:

for each row of the plurality of rows, the first column of the row contains a first-level resource identifier (ID), of a plurality of first-level resource IDs, that identifies a respective first-level resource, of a plurality of first-level resources, and the second column contains a data element, of a plurality of data elements, that provides information regarding the respective first-level resource;

initiating, by the one or more processors, one or more threads assigned to one or more unique first-level resource IDs of the plurality of first-level resource IDs; and

for each respective unique first-level resource ID of the one or more unique first-level resource IDs:

identifying, by the one or more processors, a first row set that comprises one or more of the plurality of rows that contain the respective unique first-level resource ID; and

for at least a first row of the first row set:

determining, by the one or more processors, whether a first data element in the second column of the first row is different from a second data element in the second column of a predecessor row from the first row set that precedes the first row;

in response to determining that the first data element is different from the data second element, enqueuing, by the one or more processors, the first row in a first queue for a first thread, from the one or more threads, assigned to the respective unique first-level resource ID;

dequeue the first row from the first queue for the first thread; and

cause a server to create a second-level resource that stores a version of the first data element.

11. The computer-implemented method of claim 10 , wherein the first thread is configured to, as part of causing the server to create the second-level resource, send a request to the server, causing the server to create, in a hierarchical data store, the second-level resource.

12. The computer-implemented method of claim 10 , wherein the first thread is configured to create, in a hierarchical data store, relationship data for the second-level resource that specify that the second-level resource is a parent of a relevant first-level resource for the second-level resource, wherein the relevant first-level resource for the second-level resource is identified by a first-level resource ID, of the plurality of first-level resource IDs, contained in the first column of the first row.

13. The computer-implemented method of claim 12 , wherein:

the first thread is further configured to:

obtain a first logical ID of the relevant first-level resource for the second-level resource and a first version ID of the relevant first-level resource for the second-level resource; and

obtain, from the server, a second logical ID of the second-level resource and a second version ID of the second-level resource, and

wherein the relationship data for the second-level resource specifies that the second-level resource is the parent of the relevant first-level resource for the second-level resource by specifying: (1) the second logical ID and the second version ID and (2) the first logical ID and the first version ID.

14. The computer-implemented method of claim 13 , wherein:

the second-level resource is a first second-level resource,

the first thread is further configured to store a row number of the first row, the second logical ID, and the second version ID in a cache,

each of the plurality of rows further includes a third column that contains a third-level resource ID that identifies a third-level resource,

a set of unique third-level resource IDs is defined as including only unique third-level resource IDs contained in the third column of the plurality of rows,

for each respective third-level resource ID in the set of unique third-level resource IDs, the one or more threads include a second thread assigned to the respective third-level resource ID, and

the computer-implemented method further comprises:

identifying, by the one or more processors, a second row set for the respective third-level resource ID that comprises ones of the plurality of rows that contain the respective third-level resource ID;

for at least a second row of the second row set:

determining, by the one or more processors, whether a third data element contained in any subordinate column of the second row is different from a fourth data element contained in a corresponding subordinate column of a predecessor row for the second row, wherein the predecessor row for the second row precedes the second row in the second row set;

in response to determining that the third data element is different from the fourth data element, enqueuing the second row in a second queue for the second thread, and

wherein the second thread is configured to, for each third row enqueued in the second queue dequeue the third row from the second queue;

use a row number of the third row to search the cache to obtain a third logical ID of a second second-level resource and a third version ID of the second second-level resource;

send a request causing the server to create, in the hierarchical data store, a third-level resource that specifies the third-level resource ID contained in the third column of the third row;

obtain, from the server, a fourth logical ID and a fourth version ID of the third-level resource; and

send a request causing the server to create, in the hierarchical data store, relationship data for the third-level resource that specify that the third-level resource is a parent of the second second-level resource by specifying the fourth logical ID and the fourth version ID and the third logical ID and the third version ID.

15. The computer-implemented method of claim 14 , further comprising, for a fourth row of the second row set, in response to determining that a fifth data element in the second column of the fourth row is not different from a sixth data element in the second column of the fourth row:

obtaining, by the one or more processors, a fifth logical ID of a predecessor second-level resource and a fifth version ID of the predecessor second-level resource, the predecessor second-level resource being a second-level resource created for the predecessor row for the fourth row of the second row set; and

storing, by the one or more processors, a row number of the fourth row of the second row set, the fifth logical ID, and the fifth version ID in the cache.

16. The computer-implemented method of claim 10 , wherein two or more of the threads that are assigned to the unique first-level resource IDs operate in parallel.

17. The computer-implemented method of claim 10 , wherein the server is a Fast Healthcare Interoperability Resources (FHIR) server.

18. The computer-implemented method of claim 10 , wherein:

the plurality of first-level resource IDs include a patient identifier, and the plurality of data elements provides information regarding one or more of patient first name, patient last name, patient address, or patient birth date, or

the plurality of first-level resource IDs includes a provider ID, and the plurality of data elements provides information regarding one or more of a provider name, or provider address.

19. One or more non-transitory computer readable storage media comprising instructions stored thereon that, when executed by one or more processors, cause the one or more processors to:

obtain a data file comprising a plurality of rows each including a first column and a second column, wherein:

for each row of the plurality of rows, the first column of the row contains a first-level resource identifier (ID), of a plurality of first-level resource IDs, that identifies a respective first-level resource, of a plurality of first-level resources, and the second column contains a data element, of a plurality of data elements, that provides information regarding the respective first-level resource;

initiate one or more threads assigned to one or more unique first-level resource IDs of the plurality of first-level resource IDs; and

for each respective unique first-level resource ID of the one or more unique first-level resource IDs:

identify a first row set that comprises one or more of the plurality of rows that contain the respective unique first-level resource ID; and

for at least a first row of the first row set:

determine whether a first data element in the second column of the first row is different from a second data element in the second column of a predecessor row from the first row set that precedes the first row;

in response to determining that the first data element is different from the second data element, enqueue the first row in a first queue for a first thread, from the one or more threads, assigned to the respective unique first-level resource ID;

dequeue the first row from the first queue for the first thread; and

cause a server to create a second-level resource that stores a version of the first data element.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2023
From: BISHOP, WILLIAM H.; CONNOR, CHRISTOPHER A.; HOLLIDAY, SLOAN H.
To: OPTUM, INC.
Reel/Frame 063036/0633 →
Continuity (1)
Related Publication 20240411730A1 · Dec 12, 2024
References Cited (26)
US 8010391B2 · Wait · 2011 [cited by examiner]
US 11011258B1 · Temkin et al. · 2021 [cited by applicant]
US 11120899B1 · Rai et al. · 2021 [cited by applicant]
US 20050071194A1 · Bormann · 2005 [cited by examiner]
US 20050240447A1 · Kil · 2005 [cited by examiner]
US 20080103816A1 · Kaplan · 2008 [cited by examiner]
US 20100211413A1 · Tholl · 2010 [cited by examiner]
US 20100211414A1 · Tholl · 2010 [cited by examiner]
US 20150149362A1 · Baum et al. · 2015 [cited by applicant]
US 20180150609A1 · Kim · 2018 [cited by examiner]
US 20200202988A1 · Gold et al. · 2020 [cited by applicant]
US 20200251194A1 · Vogt et al. · 2020 [cited by applicant]
US 20210058229A1 · Jiang et al. · 2021 [cited by applicant]
US 20210210160A1 · Shelton · 2021 [cited by applicant]
US 20210343418A1 · Thompson et al. · 2021 [cited by applicant]
US 20220101970A1 · Graham · 2022 [cited by applicant]
US 20220310217A1 · Hennie et al. · 2022 [cited by applicant]
US 20220319645A1 · Scipioni et al. · 2022 [cited by applicant]
US 20220343419A1 · Rolls et al. · 2022 [cited by applicant]
WO 2021174169A1 · 2021 [cited by applicant]
Grimes et al., “Pathling: analytics on FHIR”, Journal of Biomedical Semantics, vol. 13, No. 1, Springer Nature Switzerland AG, Sep. 8, 2022, 19 pp., URL: https://link.springer.com/article/10.1186/s13326-022-00277-1. [cited by applicant]
HL7 FHIR, “Base Resource Definitions”, HL7 International, vol. 5, 18 pp., Retrieved from the Internet on May 11, 2022 from URL: https://www.hl7.org/fhir/resource.html#id. [cited by applicant]
HL7 FHIR, “Extended Operations on the RESTful API”, HL7 International, vol. 5, 6 pp., Retrieved from the Internet on May 11, 2022 from URL: https://www.hl7.org/fhir/operations.html. [cited by applicant]
HL7 FHIR, “Resource Provenance—Content”, HL7 International, vol. 5, 7 pp., Retrieved from the Internet on May 11, 2022 from URL: https://www.hl7.org/fhir/provenance.html. [cited by applicant]
HL7 FHIR, “Resource References”, HL7 International, vol. 5, 17 pp., Retrieved from the Internet on May 11, 2022 from URL: https://www.hl7.org/fhir/references.html. [cited by applicant]
Robinson et al., “Fast and simple comparison of semi-structured data, with emphasis on electronic health records”, bioRxiv, Apr. 2, 2018, 13 pp., URL: https://www.biorxiv.org/content/10.1101/293183v2.full.pdf. [cited by applicant]