IP Library Granted Patent US 12,373,462
Granted Patent B2
US 12,373,462 · App. 18/535,818 · Granted Jul 29, 2025

Methods and apparatus to organize an object store namespace

Inventors: Prashant Pogde (Sunnyvale, CA); Siddharth Jivan Wagle (Saratoga, CA); Uma Maheswara Rao Gangumalla (Milpitas, CA); Arpit Ashok Agarwal (Sunnyvale, CA)
Assignee: Cloudera, Inc.
G06F16/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,462
App. No.
18/535,818
Granted
Jul 29, 2025
Kind
B2
Abstract

Disclosed examples create at least first and second database shards in a leader node, the leader node located in a consensus ring; and cause replication of first namespace metadata in the at least the first and second database shards of the leader node and in at least first and second database shards in a follower node, the follower node located in the consensus ring.

Claims (42)

1. An apparatus comprising:

interface circuitry;

machine-readable instructions; and

programmable circuitry to at least one of instantiate or execute the machine-readable instructions to:

create at least first and second database shards in a leader node, the leader node located in a consensus ring; and

cause replication of first namespace metadata to the at least the first and second database shards of the leader node and to at least first and second database shards in corresponding ones of first and second namespace databases in a follower node, the follower node located in the consensus ring.

2. The apparatus of claim 1 , comprising:

determine that the at least the first and second database shards are filled to a threshold capacity;

create a third database shard in the leader node and a third database shard in the follower node; and

cause replication of second namespace metadata in the third database shard of the leader node and in the third database shard of the follower node.

3. The apparatus of claim 2 , wherein the programmable circuitry is to create the third database shards in the leader node and in the follower node to scale up an object store namespace based on at least one additional object identified by a client device.

4. The apparatus of claim 1 , wherein the programmable circuitry is to cause replication of cross-shard metadata in a master shard of the leader node and in a master shard of the follower node.

5. The apparatus of claim 4 , wherein the cross-shard metadata includes volume information corresponding to volumes of the first namespace metadata stored in the at least the first and second database shards in the leader node.

6. The apparatus of claim 1 , wherein the programmable circuitry is to cause the replication of the first namespace metadata to the at least the first and second database shards of the leader node and to the at least the first and second database shards in the follower node by writing the first namespace metadata to the leader node and the follower node in parallel.

7. The apparatus of claim 1 , wherein the programmable circuitry is to cause the replication of the first namespace metadata to the at least the first and second database shards of the leader node and to the at least the first and second database shards in the follower node by causing transmission of transaction commands to command logs of the leader and follower nodes.

8. The apparatus of claim 1 , wherein the programmable circuitry is to create a third database shard in the leader node to increase input/output operations (IOPS), the IOPS to execute transaction commands in a command log of the leader node.

9. A non-transitory computer-readable medium comprising instructions to cause programmable circuitry to at least:

establish database shards for a leader node and a follower node of a consensus ring;

generate cross-shard metadata, the cross-shard metadata to describe the database shards; and

initiate transactions to command logs in the leader node and the follower node, the transactions to write namespace metadata and the cross-shard metadata in first ones of the database shards in the leader node and replicate the namespace metadata and the cross-shard metadata to second ones of the database shards in corresponding namespace databases in the follower node.

10. The non-transitory computer-readable medium of claim 9 , wherein the instructions are to cause the programmable circuitry to:

determine that the first ones of the database shards are filled to a threshold capacity;

create an additional database shard in the leader node and an additional database shard in the follower node; and

cause replication of second namespace metadata in the additional database shards of the leader node and the follower node.

11. The non-transitory computer-readable medium of claim 10 , wherein the instructions are to cause the programmable circuitry to create the additional database shards in the leader node and in the follower node to scale up a namespace based on a client device requesting to represent at least one additional object in the namespace.

12. The non-transitory computer-readable medium of claim 9 , wherein the instructions are to cause the programmable circuitry to cause replication of the cross-shard metadata in a master shard of the leader node and in a master shard of the follower node.

13. The non-transitory computer-readable medium of claim 12 , wherein the cross-shard metadata includes volume information corresponding to volumes of the namespace metadata stored in the database shards of the leader node.

14. The non-transitory computer-readable medium of claim 9 , wherein the instructions are to cause the programmable circuitry to initiate the transactions to the command logs by causing transmission of transaction commands to the command logs of the leader node and the follower node.

15. The non-transitory computer-readable medium of claim 9 , wherein the instructions are to cause the programmable circuitry to create an additional database shard in the leader node to increase input/output operations (IOPS), the IOPS to execute transaction commands in the leader node.

16. A method comprising:

creating at least first and second database shards in a leader node, the leader node located in a consensus ring; and

causing replication of first namespace metadata across the at least the first and second database shards of the leader node and at least first and second database shards in corresponding ones of first and second namespace databases in a follower node, the follower node located in the consensus ring.

17. The method of claim 16 , comprising:

determining that the at least the first and second database shards are filled to a threshold capacity;

creating a third database shard in the leader node and a third database shard in the follower node; and

causing replication of second namespace metadata in the third database shard of the leader node and in the third database shard of the follower node.

18. The method of claim 17 , wherein the creating of the third database shards in the leader node and in the follower node is to scale up a namespace based on a client device requesting to represent at least one additional object in the namespace.

19. The method of claim 16 , including causing replication of cross-shard metadata in a master shard of the leader node and in a master shard of the follower node.

20. The method of claim 19 , wherein the cross-shard metadata includes volume information corresponding to volumes of the first namespace metadata stored in the at least the first and second database shards in the leader node.

21. The method of claim 16 , wherein the replication of the first namespace metadata across the at least the first and second database shards of the leader node and the at least the first and second database shards in the follower node is in parallel.

22. The method of claim 16 , wherein the replication of the first namespace metadata across the at least the first and second database shards of the leader node and the at least the first and second database shards in the follower node is performed by causing transmission of transaction commands to command logs of the leader and follower nodes.

23. The method of claim 16 , including creating a third database shard in the leader node to increase input/output operations (IOPS), the IOPS to execute transaction commands in a command log of the leader node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2024
From: POGDE, PRASHANT; WAGLE, SIDDHARTH JIVAN; GANGUMALLA, UMA MAHESWARA RAO; AGARWAL, ARPIT ASHOK
To: CLOUDERA, INC.
Reel/Frame 066122/0146 →
Continuity (2)
Provisional Application 63588249 · Oct 5, 2023
Related Publication 20250117397A1 · Apr 10, 2025
References Cited (15)
US 11144510B2 · Sharma · 2021 [cited by examiner]
US 11194763B1 · Grider · 2021 [cited by examiner]
US 11853263B2 · Shvachko · 2023 [cited by examiner]
US 20130218934A1 · Lin · 2013 [cited by examiner]
US 20180246916A1 · Fathalla · 2018 [cited by examiner]
US 20180260409A1 · Sundar · 2018 [cited by examiner]
Apache, “[RATIS 1557] Support Linearizable Read from Followers,” Mar. 21, 2022, retrieved from <https://issues.apache.org/jira/si/jira.issueviews:issue-html/RATIS-1557/RATIS-1557.html> on Dec. 8, 2023, 4 pages. [cited by applicant]
Apache Software Foundation,“Apache Ozone,” published on Aug. 23, 2023, retrieved from <https://web.archive.org/web/20230823034522/http://apache.org/> on Dec. 11, 2023, 3 pages. [cited by applicant]
Apache Software Foundation, “Apache Ratis,” published on Aug. 23, 2023, retrieved from <https://web.archive.org/web/20230823034538/http://apache.org/> on Dec. 11, 2023, 5 pages. [cited by applicant]
Apache, “Documentation for Apache Hadoop Ozone,” published on Mar. 24, 2023, retrieved from <https://web.archive.org/web/20230324054301/https://ozone.apache.org/docs/1.0.0/concept/ozonemanager.html> on Dec. 11, 2023, 4 … [cited by applicant]
Apache, “Documentation for Apache Hadoop Ozone,” retrieved from <https://ozone.apache.org/docs/1.0.0/concept/ozonemanager.html>on Dec. 8, 2023, 5 pages. (Included to show figure labeled ‘Key Reads’ that is missing from … [cited by applicant]
Github, “The Raft Consensus Algorithm,” published on Sep. 27, 2023, retrieved from <https://web.archive.org/web/20230927155628/https://raft.github.io/> on Dec. 11, 2023, 14 pages. [cited by applicant]
Ongaro et al., “In Search of an Understandable Consensus Algorithm (Extended Version),” May 20, 2014, 18 pages. [cited by applicant]
Rocksdb, “A Persistent Key-Value Store,” Meta Open Source, 2022, retrieved from <https://rocksdb.org> on Dec. 8, 2023, 6 pages. [cited by applicant]
Scylladb, “Paxos Consensus Algorithm,” published on May 31, 2023, retrieved from <https://www.scylladb.com/glossary/paxos-consensus-algorithm/> on Dec. 11, 2023, 7 pages. [cited by applicant]