IP Library Granted Patent US 11,410,014
Granted Patent B2
US 11,410,014 · App. 16/272,997 · Granted Aug 9, 2022

Customizable chip for AI applications

Inventors: Saman Naderiparizi (Seattle, WA); Mohammad Rastegari (Bothell, WA); Sayyed Karen Khatamifard (Seattle, WA)
Assignee: Apple Inc.
G06N3/02G06F3/0604G06F3/0676G06F3/0677G06N3/0454G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,410,014
App. No.
16/272,997
Granted
Aug 9, 2022
Kind
B2
Abstract

In one embodiment, a computing device includes an input sensor providing an input data; a programmable logic device (PLD) implementing a convolutional neural network (CNN), wherein: each compute block of the PLD corresponds to one of a multiple of convolutional layers of the CNN, each compute block of the PLD is placed in proximity to at least two memory blocks, a first one of the memory blocks serves as a buffer for the corresponding layer of the CNN, and a second one of the memory blocks stores model-specific parameters for the corresponding layer of the CNN.

Claims (43)

1. A computing device comprising:

a programmable logic device (PLD) implementing a convolutional neural network (CNN), wherein:

each of a plurality of logical units of the PLD corresponds to one of a plurality of layers of the CNN; and

at least one of the logical units corresponds to a convolutional layer of the plurality of layers and comprises a compute block and at least two memory blocks, wherein:

the compute block is positioned in the PLD in proximity to the at least two memory blocks relative to at least one other memory block of the PLD;

a first one of the at least two memory blocks serves as a buffer for the convolutional layer; and

a second one of the at least two memory blocks stores model-specific parameters for the convolutional layer.

2. The computing device of claim 1 , wherein data in the second memory block is written into locations in the second memory block at consecutive addresses.

3. The computing device of claim 1 , wherein the model-specific parameters comprise weights or filters for the convolutional layer.

4. The computing device of claim 1 , further comprising a communication module for receiving over-the-air (OTA) updates for parameter configurations or transmitting an output data.

5. The computing device of claim 4 , wherein the output data comprises classification data corresponding to an input data.

6. The computing device of claim 4 , wherein the communication module communicates with other devices within a wireless network.

7. The computing device of claim 4 , wherein the communication module comprises at least two wireless transmitters, and wherein one of the at least two wireless transmitters is selected to be used for receiving the updates or transmitting the output data based on a supply power available from an energy source.

8. The computing device of claim 4 , wherein the output data is batched for transmission.

9. The computing device of claim 5 , further comprising an external memory to store the output data comprising classification data corresponding to the input data.

10. The computing device of claim 1 , wherein the computing device is made from a bio-degradable material.

11. The computing device of claim 1 , further comprising a camera used for capturing images or video frames, a microphone to capture audio signals, or any other sensor device.

12. The computing device of claim 1 , wherein input data is reduced, based on a supply power available from an energy source:

by reducing a sampling rate of the input data; or

by reducing a resolution at which the input data is captured.

13. The computing device of claim 1 , wherein the compute block in at least one of the logical units accesses at least one shared memory block in at least one other logical unit to read or write data.

14. The computing device of claim 13 , further comprising a memory controller implemented on the PLD, wherein the memory controller manages shared access to the at least one shared memory block.

15. The computing device of claim 1 , wherein the at least two memory blocks comprise dedicated on-chip memory blocks.

16. A system, comprising:

an input sensor providing input data;

an energy source for supplying power to the system;

a communication module; and

a programmable logic device (PLD) implementing a convolutional neural network (CNN), wherein:

each of a plurality of logical units of the PLD corresponds to one of a plurality of layers of the CNN; and

at least one of the logical units corresponds to a convolutional layer of the plurality of layers and comprises a compute block and at least two memory blocks, wherein:

the compute block is positioned in the PLD in proximity to the at least two memory blocks relative to at least one other memory block of the PLD;

a first one of the at least two memory blocks serves as a buffer for the convolutional layer; and

a second one of the at least two memory blocks stores model-specific parameters for the convolutional layer.

17. The system of claim 16 , wherein the communication module comprises at least two wireless transmitters, and wherein one of the at least two wireless transmitters is selected to be used for receiving the updates or transmitting an output data based on a supply power available from the energy source.

18. The system of claim 17 , further comprising an external memory to store the output data.

19. The system of claim 17 , wherein the output data is batched for transmission.

20. The system of claim 16 , wherein the power supplied by the energy source corresponds to a duty cycle of the energy source, and wherein the duty cycle is a rate at which the energy source charges and discharges.

21. A method for processing a computing device, comprising:

initializing a programmable logic device (PLD) with an initializing configuration for a convolutional neural network (CNN);

receiving input data;

processing, by a plurality of logical units of the PLD, the input data, wherein each of a plurality of logical units of the PLD corresponds to one of a plurality of layers of the CNN; and

wherein at least one of the logical units corresponds to a convolutional layer of the plurality of layers and comprises a compute block and at least two memory blocks, wherein the compute block is positioned in the PLD in proximity to the at least two memory blocks relative to at least one other memory block of the PLD, a first one of the at least two memory blocks serves as a buffer for the convolutional layer, and a second one of the at least two memory blocks stores model-specific parameters for the convolutional layer; and

transmitting an output data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2021
From: XNOR.AI, INC.
To: APPLE INC.
Reel/Frame 058390/0589 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2019
From: NADERIPARIZI, SAMAN; RASTEGARI, MOHAMMAD; KHATAMIFARD, SAYYED KAREN
To: XNOR.AI, INC.
Reel/Frame 048299/0271 →
Continuity (1)
Related Publication 20200257955A1 · Aug 13, 2020