IP Library › Granted Patent US 10,417,525
Granted Patent B2
US 10,417,525 · App. 14/663,233 · Granted Sep 17, 2019

Object recognition with reduced neural network weight precision

Inventors: Zhengping Ji (Pasadena, CA); Ilia Ovsiannikov (Studio City, CA); Yibing Michelle Wang (Temple City, CA); Lilong Shi (Pasadena, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06K9/6232G06K9/4628G06K9/6267G06K9/6272G06N3/0454G06N3/084G06K2209/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,417,525
App. No.
14/663,233
Granted
Sep 17, 2019
Kind
B2
Abstract

A client device configured with a neural network includes a processor, a memory, a user interface, a communications interface, a power supply and an input device, wherein the memory includes a trained neural network received from a server system that has trained and configured the neural network for the client device. A server system and a method of training a neural network are disclosed.

Claims (45)

1. A client device configured with a trained neural network, the client device comprising:

a processor, a memory, a user interface, a communications interface, a power supply and an input device;

the memory comprising the trained neural network received from a server system, wherein the server system has trained and configured a server-based neural network to be used as the trained neural network for the client device;

wherein:

the trained neural network is configured to generate a feature map, the feature map comprising a plurality of weight values derived from an input image; and

the trained neural network is configured to perform a unitary quantizing operation or a supervised iterative quantization operation on the feature map to reduce a number of bits of each weight of the plurality of weight values from a first predetermined number to a second predetermined number that is less than the first predetermined number without changing a dimension of the feature map.

2. The client device as in claim 1 , wherein the input device is configured to capture an image and to store image input data in the memory.

3. The client device as in claim 1 , further comprising a multilayer perceptron (MLP) classifier configured to map image input data.

4. The client device as in claim 1 , wherein the trained neural network comprises a convolutional neural network.

5. The client device as in claim 1 , wherein the quantization operation performs back-propagation (BP) of image input data.

6. The client device as in claim 1 , wherein the trained neural network is configured to perform object recognition.

7. The client device as in claim 1 , comprising one of a smartphone, a tablet computer and a portable electronic device.

8. The client device as in claim 1 , wherein the trained neural network is a quantized low-bit version of the server-based neural network.

9. The client device as in claim 1 , wherein network weights for the trained neural network are quantized for lower bit resolution by the server system.

10. A method that comprises performing the following using a client device:

receiving a trained neural network from a server system, wherein the server system has trained and configured a server-based neural network to be used as the trained neural network for the client device;

capturing an input image;

processing the input image using the trained neural network;

generating a feature map using the trained neural network, the feature map comprising a plurality of weight values derived from an input image;

performing a unitary quantizing operation or a supervised iterative quantization operation on the feature map using the trained neural network to reduce a number of bits of each weight of the plurality of weight values from a first predetermined number to a second predetermined number that is less than the first predetermined number without changing a dimension of the feature map; and

recognizing an object in the input image based on a result of the processing.

11. The method of claim 10 , wherein receiving the trained neural network includes:

receiving a quantized low-bit version of the server-based neural network as the trained neural network.

12. The method of claim 11 , wherein receiving the quantized low-bit version includes:

receiving network weights associated with the trained neural network, wherein the network weights are quantized for lower bit resolution by the server system.

13. The method of claim 10 , wherein recognizing the object includes:

analyzing the result of the processing using a multilayer perceptron (MLP) classifier to recognize the object in the input image.

14. A non-transitory computer-readable medium storing program code, which, when executed by a processor, cause the processor to perform the following:

receive a trained neural network from a server system, wherein the server system has trained and configured a server-based neural network to be used as the trained neural network;

capture an input image;

generating a feature map using the trained neural network, the feature map comprising a plurality of weight values derived from an input image;

performing a unitary quantizing operation or supervised iterative quantization operation on the feature map using the trained neural network to reduce a number of bits of each weight of the plurality of weight values from a first predetermined number to a second predetermined number that is less than the first predetermined number without changing a dimension of the feature map; and

process the input image using the trained neural network to recognize an object in the input image.

15. The non-transitory computer-readable medium of claim 14 , wherein the program code, when executed by the processor, cause the processor to further perform the following:

receive network weights associated with the trained neural network, wherein the network weights are quantized for lower bit resolution by the server system.

16. The non-transitory computer-readable medium of claim 14 , wherein the program code, when executed by the processor, cause the processor to further perform the following:

importing a low-resolution configuration of the server-based neural network; and

storing the imported configuration as the trained neural network.

17. A client device configured with a trained neural network, the client device comprising:

a processor, a memory, a user interface, a communications interface, a power supply and an input device;

the memory comprising the trained neural network received from a server system, wherein the server system has trained and configured a server-based neural network to be used as the trained neural network for the client device;

wherein:

the trained neural network is configured to generate a feature map, the feature map comprising a plurality of first weight values derived from an input image;

the trained neural network is configured to convert the first weights of the feature map into second weights by a unitary or a supervised iterative quantizing operation; and

the second weights are encoded using a number of bits lower than that used to encode the first weights without changing a dimension of the feature map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2015
From: JI, ZHENGPING; OVSIANNIKOV, ILIA; WANG, YIBING M.; SHI, LILONG
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 035257/0605 →
Continuity (2)
Provisional Application 62053692 · Sep 22, 2014
Related Publication 20160086078A1 · Mar 24, 2016
Cited By (3)
US 12,591,776 US 12,657,898 US 12,699,903