IP Library Granted Patent US 8,359,281
Granted Patent B2
US 8,359,281 · App. 12/478,073 · Granted Jan 22, 2013

System and method for parallelizing and accelerating learning machine training and classification using a massively parallel accelerator

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,359,281
App. No.
12/478,073
Granted
Jan 22, 2013
Kind
B2
Abstract

A method system for training an apparatus to recognize a pattern includes providing the apparatus with a host processor executing steps of a machine learning process; providing the apparatus with an accelerator including at least two processors; inputting training pattern data into the host processor; determining coefficient changes in the machine learning process with the host processor using the training pattern data; transferring the training data to the accelerator; determining kernel dot-products with the at least two processors of the accelerator using the training data; and transferring the dot-products back to the host processor.

Claims (40)

1. A method for training an apparatus to recognize a pattern, the method comprising the steps of:

providing the apparatus with a host processor executing steps of a machine learning process;

providing the apparatus with an accelerator including at least two processors;

inputting training data into the host processor;

reducing the precision of the training data with the host processor

transferring the training data with reduced precision to the accelerator;

determining coefficient changes in the machine learning process with the host processor using the training data with reduced precision;

transferring indices pertaining to one or more training vectors to the accelerator

determining kernel dot-products with the at least two processors of the accelerator using the training data with reduced precision; and

transferring the dot-products back to the host processor.

2. The method of claim 1 , further comprising the step of determining kernels of the machine learning process with the host processor using the kernel dot-products.

3. The method of claim 2 , further comprising the step of determining gradients of the machine learning process with the host processor using the kernels.

4. The method of claim 1 , wherein the step of reducing the precision of the training data with the host processor is performed prior to the step of transferring the training data to the accelerator.

5. The method of claim 4 , further comprising the step of reducing the precision of the dot-products with the accelerator prior to the step of transferring the dot-products back to the host processor.

6. The method of claim 1 , further comprising the step of reducing the precision of the kernel dot-products with the accelerator prior to the step of transferring the dot-products back to the host processor.

7. The method of claim 6 , wherein the accelerator further includes a memory bank associated with each one of the at least two processors, and further comprising the step of partitioning the reduced precision kernel dot-products into groups and storing each of the groups of the kernel dot-products in one of the memory banks prior to the step of transferring the kernel dot-products back to the host processor.

8. The method of claim 1 , wherein the kernel dot-products are determined in a parallel manner with the at least two processors of the accelerator.

9. The method of claim 1 , wherein the kernel dot-products are determined in separate and discrete chunks.

10. The method of claim 1 , wherein the accelerator further includes a memory banks associated with each one of the at least two processors, and further comprising the step of partitioning the kernel dot-products into groups and storing each of the groups of the kernel dot-product in one of the memory banks prior to the step of transferring the kernel dot-products back to the host processor.

11. A system for training an apparatus to recognize a pattern, the system comprising:

a host processor of the apparatus for determining coefficient changes of a machine learning process from input training data and for reducing the precision of the training data;

an accelerator including at least two processors for determining kernel dot-products using the training data with reduced precision received from the host processor; and

at least one conduit for transferring the training data with reduced precision from the host processor to the accelerator and for transferring the kernel dot-products from the accelerator to the host processor.

12. The system of claim 11 , wherein the host processor uses the kernel dot-products to determine kernels of the machine learning process.

13. The system of claim 12 , wherein the host processor uses the kernels to determine gradients of the machine learning process.

14. The system of claim 11 , wherein the host processor reduces the precision of the training data prior to its transfer to the accelerator.

15. The system of claim 14 , wherein the accelerator reduces the precision of the kernel dot-products prior to their transfer to the host processor.

16. The system of claim 11 , wherein the accelerator reduces the precision of the kernel dot-products prior to their transfer to the host processor.

17. The system of claim 16 , wherein the accelerator further includes a memory bank associated with each one of the at least two processors, and wherein the kernel dot-products are partitioned into groups and each of the groups of the kernel dot-products are stored in one of the memory banks prior to being transferred to the host processor.

18. The system of claim 11 , wherein the kernel dot-products are determined in a parallel manner by the at least two processors of the accelerator.

19. The system of claim 11 , wherein the kernel dot-products are determined in separate and discrete chunks.

20. The system of claim 11 , wherein the accelerator further includes a memory bank associated with each one of the at least two processors, and wherein the kernel dot-products are partitioned into groups and each of the groups of the kernel dot-products are stored in one of the memory banks prior to being transferred to the host processor.

21. A method for recognizing patterns, the method comprising the steps of:

providing host processor executing steps of a support vector machine learning process;

providing an accelerator including at least two processors and a memory bank associated with each of the at least two processors;

storing support vectors in the memory banks of the accelerator;

reducing the precision of unlabeled pattern data with the host processor;

transferring unlabeled pattern data from the host processor to the accelerator;

calculating labels for the unlabeled pattern data with the at least two processors of the accelerator using the support vectors stored in the memory banks of the accelerator;

and transferring the labeled pattern data back to the host processor.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE 8538896 AND ADD 8583896 PREVIOUSLY RECORDED ON REEL 031998 FRAME 0667. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 30, 2017
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 042754/0703 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2014
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 031998/0667 →