IP Library › Granted Patent US 12,265,915
Granted Patent B2
US 12,265,915 · App. 17/138,709 · Granted Apr 1, 2025

Composable neural network kernels

Inventors: Chao Liu (Austin, TX); Daniel Isamu Lowell (Austin, TX); Wen Heng Chung (Austin, TX); Jing Zhang (Austin, TX)
Assignee: Advanced Micro Devices, Inc.
G06N3/10G06F8/41G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,915
App. No.
17/138,709
Granted
Apr 1, 2025
Kind
B2
Abstract

A technique for manipulating a generic tensor is provided. The technique includes receiving a first request to perform a first operation on a generic tensor descriptor associated with the generic tensor, responsive to the first request, performing the first operation on the generic tensor descriptor, receiving a second request to perform a second operation on generic tensor raw data associated with the generic tensor, and responsive to the second request, performing the second operation on the generic tensor raw data, the performing the second operation including mapping a tensor coordinate specified by the second request to a memory address, the mapping including evaluating a delta function to determine an address delta value to add to a previously determined address for a previously processed tensor coordinate.

Claims (53)

1. A method for manipulating a generic tensor, the method comprising:

receiving a request to perform an operation on generic tensor raw data associated with a generic tensor; and

responsive to the request, performing the operation on generic tensor raw data, the performing the operation including mapping a tensor coordinate specified by the request to a memory address, the mapping including evaluating a delta function associated with a type of the request,

wherein the operation generates a modified generic tensor descriptor based on a generic tensor descriptor without modifying the generic tensor raw data,

wherein the delta function indicates a difference between a previously calculated memory address and the memory address, the difference determined based on a difference between the tensor coordinate and a previously calculated tensor coordinate for the previously calculated memory address.

2. The method of claim 1 , wherein the operation comprises one of:

a slice, a strided slice, a pad, a reorder, a fold, a merge, an embed, and a move slicing window operation.

3. The method of claim 1 , wherein the operation comprises one of:

a sliced copy, a general matrix multiply or batched general matrix multiply, a reduction, and an algorithm-specific transformation.

4. The method of claim 1 , wherein the generic tensor raw data comprises:

data elements of the generic tensor.

5. The method of claim 1 , wherein:

the request and operation are specified by instructions of a program.

6. The method of claim 1 , wherein:

the request is specified by a compiled program; and

the operation is performed by the compiled program.

7. The method of claim 1 , wherein:

the request is specified by a program; and

the operation is performed by a specialized hardware circuit configured to perform at least one operation on generic tensor descriptors or on generic tensor raw data.

8. The method of claim 1 , wherein the generic tensor descriptor indicates any one or a combination of the following for the generic tensor: a type of generic tensor, a number of dimensions of the generic tensor, lengths for each dimension of the generic tensor, or a base address for the generic tensor raw data.

9. The method of claim 1 , wherein the generic tensor descriptor indicates how to translate one or more index values into one or more memory addresses for generic tensor raw data associated with the generic tensor.

10. The method of claim 9 , wherein the one or more index values are translated into one or more memory addresses for the generic tensor raw data where the translating includes a reduction of dimensionality for a merged tensor or an increase of dimensionality for an embedded tensor.

11. A system for manipulating a generic tensor, the system comprising:

a memory storing a generic tensor descriptor; and

a processor, configured to:

receive a request to perform an operation on generic tensor raw data associated with a generic tensor; and

responsive to the request, performing the operation on generic tensor raw data, the perform the operation including mapping a tensor coordinate specified by the request to a memory address, the mapping including evaluating a delta function associated with a type of the request,

wherein the operation generates a modified generic tensor descriptor based on the generic raw tensor descriptor without modifying the generic tensor raw data,

wherein the delta function indicates a difference between a previously calculated memory address and the memory address, the difference determined based on a difference between the tensor coordinate and a previously calculated tensor coordinate for the previously calculated memory address.

12. The system of claim 11 , wherein the operation comprises one of:

a pad, a slice, a strided slice, a reorder, a fold, a merge, an embed, and a move slicing window operation.

13. The system of claim 11 , wherein the operation comprises one of:

a sliced copy, a general matrix multiply or batched general matrix multiply, a reduction, and an algorithm-specific transformation.

14. The system of claim 11 , wherein the generic tensor raw data comprises:

data elements of the generic tensor.

15. The system of claim 11 , wherein:

the request and operation are specified by instructions of a program.

16. The system of claim 11 , wherein:

the request is specified by a compiled program; and

the operation is performed by the compiled program.

17. The system of claim 11 , wherein:

the request is specified by a program; and

the operation is performed by a specialized hardware circuit configured to perform at least one operation on generic tensor descriptors or on generic tensor raw data.

18. The system of claim 11 , wherein the generic tensor descriptor indicates any one or a combination of the following for the generic tensor: a type of generic tensor, a number of dimensions of the generic tensor, lengths for each dimension of the generic tensor, or a base address for the generic tensor raw data.

19. The system of claim 11 , wherein the generic tensor descriptor indicates how to translate one or more index values into one or more memory addresses for generic tensor raw data associated with the generic tensor.

20. The system of claim 19 , wherein the one or more index values are translated into one or more memory addresses for the generic tensor raw data where the translating includes a reduction of dimensionality for a merged tensor or an increase of dimensionality for an embedded tensor.

21. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to manipulate a generic tensor, by:

receiving a request to perform an operation on generic tensor raw data associated with a generic tensor; and

responsive to the request, performing the operation on generic tensor raw data, the performing the operation including mapping a tensor coordinate specified by the request to a memory address, the mapping including evaluating a delta function associated with a type of the request,

wherein the operation generates a modified generic tensor descriptor based on a generic tensor descriptor without modifying the generic tensor raw data,

wherein the delta function indicates a difference between a previously calculated memory address and the memory address, the difference determined based on a difference between the tensor coordinate and a previously calculated tensor coordinate for the previously calculated memory address.

22. The non-transitory computer-readable medium of claim 21 , wherein the operation comprises one of:

a slice, a strided slice, a pad, a reorder, a fold, a merge, an embed, and a move slicing window operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2021
From: LIU, CHAO; LOWELL, DANIEL ISAMU; CHUNG, WEN HENG; ZHANG, JING
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 055449/0945 →
Continuity (6)
Continuation In Part 16779557 · Jan 31, 2020
Provisional Application 63024376 · May 13, 2020
Provisional Application 62927603 · Oct 29, 2019
Provisional Application 62925168 · Oct 23, 2019
Provisional Application 62867766 · Jun 27, 2019
Related Publication 20210117806A1 · Apr 22, 2021
References Cited (15)
US 6367069B1 · Corbett · 2002 [cited by applicant]
US 9753762B1 · Emelyanov · 2017 [cited by applicant]
US 20150370544A1 · Santry et al. · 2015 [cited by examiner]
US 20160232175A1 · Zhou et al. · 2016 [cited by applicant]
US 20160323199A1 · Bryant et al. · 2016 [cited by applicant]
US 20180322386A1 · Sridharan et al. · 2018 [cited by examiner]
US 20190042940A1 · Sakairi et al. · 2019 [cited by applicant]
US 20190130270A1 · Nicol et al. · 2019 [cited by examiner]
CN 109886399A · 2019 [cited by applicant]
WO 2020200244A1 · 2020 [cited by applicant]
De Almeida Maia, H., et al., “A Video Tensor Self-descriptor Based on Block Matching”, International Conference on Computational Science and Its Applications, Jun. 2014; retrieved on Sep. 10, 2020 from <URL: https://doi… [cited by applicant]
Gardner, J., et al., “Driving with Data: Modeling and Forecasting Vehicle Fleet Maintenance in Detroit”, arXiv:1710.06839v1; Oct. 18, 2017; retrieved on Sep. 10, 2020 from <URL: https://arxiv.org.1710.06839V1.pdf> secti… [cited by applicant]
Mota, V. F., et al., “A tensor montion descriptor based on histograms of gradients and optical flow”, Pattern Recognition Letters, vol. 39, Issue 1, Apr. 2014; retrieved on Sep. 10, 2020 from <URL: https://doi.org/10/10… [cited by applicant]
Asakawa Shinichi, “Experience Deep Learning with Python-Caffe, Theano, Chainer, TensorFlow-,” First Edition, Corona Publishing Co., Ltd. Publication date 2016010/25, pp. 33, 37, Independent Book 2018-00275-001. [cited by applicant]
Ariyama Keiji, “Getting Started with TensorFlow3 Object Detection”, First Edition, Impress R&D, Publication date Feb. 16, 2018, p. 13, Independet Book 2020-01302-001. [cited by applicant]