PROCESSING STRUCTURED AND UNSTRUCTURED DATA USING OFFLOAD PROCESSORS
A data processing system for unstructured data is disclosed. A plurality of modules can be connected to a memory bus, each including at least one processor. A central processing unit (CPU) can be connected to the modules by the memory bus, with the CPU configured to process computationally intensive data processing tasks while directing parallel computation tasks to a plurality of the modules.
1 . A structured data processing system, comprising:
a plurality of modules connected to a memory bus in a first server, each including at least one processor, and at least one module configured to provide an in-memory database; and
a central processing unit (CPU) in the first server connected to the in-line modules by the memory bus, with the CPU configured to process and direct structured queries to the plurality of in-line modules.
2 . The structured data processing system of claim 1 , wherein the modules are configured to communicate with each other without requiring access to the CPU of the first server.
3 . The structured data processing system of claim 1 , wherein the modules are mounted on different servers in a same rack, and further comprising a top of the rack switch to mediate communication therebetween.
4 . The structured data processing system of claim 1 , further including a module driver configured to execute a memory mapping function to transfer a query from the CPU to at least one module in the form of a memory read or writes.
5 . The structured data processing system of claim 1 , wherein the modules are configured for insertion into a dual-in-line-memory module (DIMM) socket, and the module further comprises the processors being connected to memory and a computational field programmable gate array (FPGA).
6 . A data processing system for unstructured data, comprising:
a plurality of modules connected to a memory bus, each including at least one processor; and
a central processing unit (CPU) connected to the modules by the memory bus, with the CPU configured to process computationally intensive data processing tasks while directing parallel computation tasks to a plurality of the modules.
7 . The data processing system for unstructured data of claim 6 , wherein a Map/Reduce algorithm is processed, with results of a Map step being stored and available in main memory, and collected through DMA operations and parsed by the modules.
8 . The data processing system for unstructured data of claim 6 , wherein the modules together define a massively parallel input/output (I/O) mid-plane defined by the modules.
9 . The data processing system for unstructured data of claim 6 , further including:
the modules are disposed on multiple servers;
a Map/Reduce algorithm is processed by the multiple servers, and the servers are configured exchange intermediate (key, value) pairs across the multiple servers through switching of the modules.
10 . The data processing system for unstructured data of claim 6 , wherein each module is configured for insertion into a dual-in-line-memory-module (DIMM) socket, and each module further comprises offload processors connected to memory and a computational field programmable gate array (FPGA).