IP Library Granted Patent US 11,822,600
Granted Patent B2
US 11,822,600 · App. 17/248,386 · Granted Nov 21, 2023

Content tagging

Inventors: Xiaoyu Wang (Playa Vista, CA); Ning Xu (Irvine, CA); Ning Zhang (Los Angeles, CA); Vitor R. Carvalho (San Diego, CA); Jia Li (Marina Del Rey, CA)
Assignee: Snap Inc.
G06F16/5866G06F16/9038G06F18/24G06N3/04G06N3/045G06N3/08G06T1/0007G06V10/751G06V10/764G06V10/82H04N23/63G06N5/022G06V2201/09
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,822,600
App. No.
17/248,386
Granted
Nov 21, 2023
Kind
B2
Abstract

Systems, methods, devices, media, and computer readable instructions are described for local image tagging in a resource constrained environment. One embodiment involves processing image data using a deep convolutional neural network (DCNN) comprising at least a first subgraph and a second subgraph, the first subgraph comprising at least a first layer and a second layer, processing, the image data using at least the first layer of the first subgraph to generate first intermediate output data; processing, by the mobile device, the first intermediate output data using at least the second layer of the first subgraph to generate first subgraph output data, and in response to a determination that each layer reliant on the first intermediate data have completed processing, deleting the first intermediate data from the mobile device. Additional embodiments involve convolving entire pixel resolutions of the image data against kernels in different layers if the DCNN.

Claims (36)

1. A method comprising:

accessing, by one or more processors of a mobile device, image data;

convolving using a first convolution layer, a first plurality of kernels of a deep convolutional neural network (DCNN) with the image data to generate first intermediate data, wherein the first plurality of kernels comprise pixel heights less than a pixel height of the image data and pixel widths less than a pixel width of the image data, and wherein the first plurality of kernels are associated with a first plurality of tags;

convolving using a second convolution layer, a second plurality of kernels of the DCNN with the first intermediate data to generate second intermediate data, wherein the second plurality of kernels are associated with a second plurality of tags;

determining, with a fully connected layer of the DCNN, a third intermediate data from the second intermediate data, the third intermediate data associated with a third plurality of tags; and

assigning one or more tags of the third plurality of tags to portions of the image data based on a comparison of a plurality of output values with threshold values, wherein each tag of the first plurality of tags, the second plurality of tags, and the third plurality of tags indicates a prediction score map for an image category.

2. The method of claim 1 wherein the method further comprises:

capturing, using an image sensor of the mobile device, the image data.

3. The method of claim 1 further comprising:

converting the fully connected layer of the DCNN to a third convolution layer.

4. The method of claim 3 wherein the third convolution layer comprises: a third plurality of kernels of the DCNN and wherein the third plurality of kernels are associated with the third plurality of tags.

5. The method of claim 1 wherein the first convolution layer and the second convolution layer convolve across the pixel width of the image data and the pixel height of the image data.

6. The method of claim 1 wherein the image data is not split into multiple sub-windows.

7. The method of claim 1 wherein the DCNN comprises additional convolution layers.

8. The method of claim 1 wherein 16-bit half precision values are used to store first intermediate data and second intermediate data.

9. The method of claim 1 further comprising: receiving, by one or more processors of the mobile device, a plurality of weights for the DCNN, wherein the plurality of weights are floating point weights compressed to weight indices.

10. The method of claim 9 further comprising:

in response to convolving processing, using the first convolution layer, decompressing a first set of weight indices.

11. The method of claim 1 wherein convolving, using the first convolution layer, further comprises: using a corresponding first set of floating point weights as decompressed from weight indices.

12. The method of claim 1 wherein the DCNN comprises only convolutional layers and a fully connected layer.

13. The method of claim 1 wherein the DCNN uses a plurality of weights comprising 16 bit weights, 32 bit weights, or 64 bit weights.

14. The method of claim 1 wherein the DCNN refrains from using a max-pooling layer.

15. The method of claim 1 further comprising:

storing the image data and indications of the one or more assigned tags to portions of the image data.

16. A mobile device for image tagging comprising: a memory; an image sensor coupled to the memory; and one or more processors coupled to the memory and configured to:

access, by one or more processors of a mobile device, image data;

convolve using a first convolution layer, a first plurality of kernels of a deep convolutional neural network (DCNN) with the image data to generate first intermediate data, wherein the first plurality of kernels comprise pixel heights less than a pixel height of the image data and pixel widths less than a pixel width of the image data, and wherein the first plurality of kernels are associated with a first plurality of tags;

convolve using a second convolution layer, a second plurality of kernels of the DCNN with the first intermediate data to generate second intermediate data, wherein the second plurality of kernels are associated with a second plurality of tags;

determine, with a fully connected layer of the DCNN, a third intermediate data from the second intermediate data, the third intermediate data associated with a third plurality of tags; and

assign one or more tags of the third plurality of tags to portions of the image data based on a comparison of a plurality of output values with threshold values, wherein each tag of the first plurality of tags, the second plurality of tags, and the third plurality of tags indicates a prediction score map for an image category.

17. A non-transitory storage medium comprising instructions that, when executed by one or more processors of a mobile device, cause the mobile device to perform operations for local image tagging, the operations comprising:

accessing, by one or more processors of a mobile device, image data;

convolving using a first convolution layer, a first plurality of kernels of a deep convolutional neural network (DCNN) with the image data to generate first intermediate data, wherein the first plurality of kernels comprise pixel heights less than a pixel height of the image data and pixel widths less than a pixel width of the image data, and wherein the first plurality of kernels are associated with a first plurality of tags;

convolving using a second convolution layer, a second plurality of kernels of the DCNN with the first intermediate data to generate second intermediate data, wherein the second plurality of kernels are associated with a second plurality of tags;

determining, with a fully connected layer of the DCNN, a third intermediate data from the second intermediate data, the third intermediate data associated with a third plurality of tags; and

assigning one or more tags of the third plurality of tags to portions of the image data based on a comparison of a plurality of output values with threshold values, wherein each tag of the first plurality of tags, the second plurality of tags, and the third plurality of tags indicates a prediction score map for an image category.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2023
From: WANG, XIAOYU; XU, NING; ZHANG, NING; CARVALHO, VITOR R.; LI, JIA
To: SNAPCHAT, INC.
Reel/Frame 064933/0302 →
CHANGE OF NAME Recorded Sep 18, 2023
From: SNAPCHAT, INC.
To: SNAP INC.
Reel/Frame 064934/0284 →
Continuity (5)
Continuation 16192419 · Nov 15, 2018
Continuation 15247697 · Aug 25, 2016
Provisional Application 62358461 · Jul 5, 2016
Provisional Application 62218965 · Sep 15, 2015
Related Publication 20210216830A1 · Jul 15, 2021
Cited By (4)
US 12,197,543 US 12,380,159 US 12,411,890 US 12,645,736