IP Library Granted Patent US 9,798,735
Granted Patent B2
US 9,798,735 · App. 15/381,733 · Granted Oct 24, 2017

Map-reduce ready distributed file system

Inventors: Mandayam C. Srivas (Union City, CA); Pindikura Ravindra (Hyderabad, IN); Uppaluri Vijaya Saradhi (Hyderabad, IN); Arvind Arun Pande (Mumbai, IN); Chandra Guru Kiran Babu Sanapala (Hyderabad, IN); Lohit Vijaya Renu (Sunnyvale, CA); Vivekanand Vellanki (Hyderabad, IN); Sathya Kavacheri (Fremont, CA); Amit Ashoke Hadke (San Jose, CA)
Assignee: MapR Technologies, Inc.
G06F17/30174G06F17/30194G06F17/30215G06F17/30345G06F17/30227G06F17/30575
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,798,735
App. No.
15/381,733
Granted
Oct 24, 2017
Kind
B2
Abstract

A map-reduce compatible distributed file system that consists of successive component layers that each provide the basis on which the next layer is built provides transactional read-write-update semantics with file chunk replication and huge file-create rates. Containers provide the fundamental basis for data replication, relocation, and transactional updates. A container location database allows containers to be found among all file servers, as well as defining precedence among replicas of containers to organize transactional updates of container contents. Volumes facilitate control of data placement, creation of snapshots and mirrors, and retention of a variety of control and policy information. Also addressed is the use of distributed transactions in a map-reduce system; the use of local and distributed snapshots; replication, including techniques for reconciling the divergence of replicated data after a crash; and mirroring.

Claims (33)

1. A map-reduce compatible distributed file system comprising replicated containers preventing data loss comprising:

a container location database (CLDB) configured to maintain information about where each of a plurality of containers is located;

a plurality of cluster nodes, each cluster node containing one or more storage pools, each storage pool containing zero or more containers; and

a plurality of inodes for structuring said file system objects within said containers; wherein said containers comprise file system objects, and said replicated containers preventing data loss comprise said containers replicated to other cluster nodes with one container designated as master container for each replication chain controlling transactions for said replication chain, said replication chain arranged in a linear pattern, a star pattern, or any combination of said linear and said star pattern, wherein said replication chain for said container is changed if a node holding any replica fails or is taken out of service, or if a node that previously contained a replica returns to service;

wherein said maintained information about where each of said plurality of containers is located that is maintained in said CLDB is stored as inodes in containers;

wherein said CLDB inodes are configured to maintain a database that contains at least following information about all of said containers:

nodes that have replicas of a container; and

an ordering of said replication chain for said container;

wherein updates to said container are sent to said master container for said updated container;

wherein changes to content of said container are propagated to said replicas of said container by said master container;

wherein some file system objects are larger than a single container; and

wherein some file system objects are spread over a larger number of nodes than a set represented by said replication chain of a single container.

2. The system of claim 1 , wherein said replication chain is arranged in a linear replication pattern, in which a node holding said master container of said container propagates updates to a first slave node that contains another replica; and

wherein said first slave node, in turn, propagates updates to a second slave node that contains a third replica of said container.

3. The system of claim 1 , wherein said replication chain is arranged in a star replication pattern, in which a node containing said master container of a container propagates updates directly and simultaneously to all other nodes containing replicas of said container.

4. The system of claim 1 , wherein when an update is received by a node containing said master container of said container, a lock is taken on a particular portion of a file being updated to guarantee that updates that change a same portion of said file are transactionally serialized; and

wherein updates are made consistently on all replicas of said container.

5. The system of claim 1 , wherein a number of update transactions that are allowed to be pending for a single file is limited to increase amount of parallelism available for updates.

6. The system of claim 1 , wherein updates to said containers are acknowledged once enough of said nodes containing replicas have acknowledged that they have received said update.

7. The system of claim 1 , wherein multiple updates are pending at a same time to increase potential parallelism, but all updates to said container are assigned a serial update identifier by said master container.

8. The system of claim 1 , wherein updates are applied on different replicas in different orders, but only updates that have been acknowledged by all replicas are reported back to a client as committed.

9. The system of claim 1 , wherein, before reporting an update as completed, all updates with smaller update identifiers are also completed.

10. The system of claim 1 , wherein potentially out-of-date replicas are resynchronized to a current state of said master container of said container.

11. The system of claim 1 , wherein when a node containing an out-of-date replica joins an existing replica chain, said node containing said out-of-date replica contacts said CLDB;

wherein said CLDB may decide that enough replicas are already known, in which case said node is instructed to discard said replica; and

wherein said CLDB may also assign said out-of-date replica to said replication chain, at which point a resynchronization process is executed to bring said out-of-date container up to a current state.

12. The system of claim 1 , wherein when a set of failing nodes does not include a node containing said master container, failing nodes are removed from said replication chain, but said master container survives; and

wherein all other replicas are considered to be out-of-date.

13. The system of claim 1 , wherein a file chunk and an original inode are updated in a coordinated fashion.

14. The system of claim 1 , wherein an original inode and multiple file chunks are updated together.

15. The system of claim 1 , further comprising: a distributed transaction in a form of a snapshot of a file system volume, consisting of directories and files spread over a number of containers;

wherein all data and meta-data for a volume is organized into a single name container and zero or more data containers; and

wherein all cross-container references to data are segregated into said name container while keeping all of said data in data containers.

Assignments (7)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2019
From: MAPR (ABC), LLC
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 050835/0730 →
NUNC PRO TUNC ASSIGNMENT Recorded Oct 27, 2019
From: MAPR TECHNOLOGIES, INC.
To: MAPR (ABC), LLC
Reel/Frame 050835/0717 →
RELEASE OF SECURITY INTEREST Recorded Aug 5, 2019
From: SILICON VALLEY BANK
To: MAPR TECHNOLOGIES, INC.
Reel/Frame 049962/0587 →
RELEASE OF SECURITY INTEREST Recorded Aug 5, 2019
From: LIGHTSPEED VENTURE PARTNERS SELECT, L.P.; LIGHTSPEED VENTURE PARTNERS VIII, L.P.; NEW ENTERPRISES ASSOCIATES 13, LIMITED PARTNERSHIP; CAPITALG II LP; MAYFIELD XIII, A CAYMAN ISLANDS EXEMPTED LIMITED PARTNERSHIP; MAYFIELD SELECT, A CAYMAN ISLANDS EXEMPTED LIMITED PARTNERSHIP
To: MAPR TECHNOLOGIES, INC.
Reel/Frame 049962/0462 →
SECURITY INTEREST Recorded Jun 28, 2019
From: MAPR TECHNOLOGIES, INC.
To: LIGHTSPEED VENTURE PARTNERS VIII, L.P.; LIGHTSPEED VENTURE PARTNERS SELECT, L.P.; NEW ENTERPRISE ASSOCIATES 13, LIMITED PARTNERSHIP; CAPITALG II LP; MAYFIELD XIII, A CAYMAN ISLANDS EXEMPTED LIMITED PARTNERSHIP; MAYFIELD SELECT, A CAYMAN ISLANDS EXEMPTED LIMITED PARTNERSHIP
Reel/Frame 049626/0030 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 21, 2019
From: MAPR TECHNOLOGIES, INC.
To: SILICON VALLEY BANK
Reel/Frame 049555/0484 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2017
From: SRIVAS, MANDAYAM C.; RAVINDRA, PINDIKURA; SARADHI, UPPALURI VIJAYA; PANDE, ARVIND ARUN; SANAPALA, CHANDRA GURU KIRAN BABU; RENU, LOHIT VIJAYA; VELLANKI, VIVEKANAND; KAVACHERI, SATHYA; HADKE, AMIT ASHOKE
To: MAPR TECHNOLOGIES, INC.
Reel/Frame 042064/0417 →
Continuity (5)
Continuation 14951437 · Nov 24, 2015
Continuation 13340532 · Dec 29, 2011
Continuation In Part 13162439 · Jun 16, 2011
Provisional Application 61356582 · Jun 19, 2010
Related Publication 20170116221A1 · Apr 27, 2017