IP Library Granted Patent US 11,237,971
Granted Patent B1
US 11,237,971 · App. 17/023,015 · Granted Feb 1, 2022

Compile time logic for detecting streaming compatible and broadcast compatible data access patterns

Inventors: Kevin James Brown (Belmont, CA); David Alan Koeplinger (Menlo Park, CA); Weiwei Chen (Mountain View, CA); Xiaoming Gu (Campbell, CA)
Assignee: SambaNova Systems, Inc.
G06F12/0842G06F5/10G06F11/3072G06F2205/123G06F2212/1016G06F2212/45
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,237,971
App. No.
17/023,015
Granted
Feb 1, 2022
Kind
B1
Abstract

A dataflow graph for an application has operation units that are configured to be producers and consumers of tensors. A write access pattern of a particular producer specifies an order in which the particular producer generates elements of a tensor, and a read access pattern of a corresponding consumer specifies an order in which the corresponding consumer processes the elements of the tensor. The technology disclosed detects conflicts between the producers and the corresponding consumers that have mismatches between the write access patterns and the read access patterns. A conflict occurs when the order in which the particular producer generates the elements of the tensor is different from the order in which the corresponding consumer processes the elements of the tensor. The technology disclosed resolves the conflicts by inserting buffers between the producers and the corresponding consumers.

Claims (33)

1. A data processing system, comprising:

memory storing a dataflow graph for an application, the dataflow graph having operation units that are configured to be producers to produce tensors for execution of the application, and to be consumers to consume the tensors for execution of the application;

the memory storing write access patterns of the producers, and read access patterns of the consumers, wherein a write access pattern of a particular producer specifies an order in which the particular producer generates elements of a tensor, and a read access pattern of a corresponding consumer specifies an order in which the corresponding consumer processes the elements of the tensor; and

a compile time logic having access to the memory and configured to process the dataflow graph to

detect conflicts between certain ones of the producers and corresponding ones of the consumers that have mismatches between the write access patterns and the read access patterns, wherein a conflict occurs when the order in which the particular producer generates the elements of the tensor is different from the order in which the corresponding consumer processes the elements of the tensor; and

resolve the conflicts by inserting buffers between the certain ones of the producers and the corresponding ones of the consumers.

2. The data processing system of claim 1 , wherein the certain ones of the producers are configured to write elements of the tensors in the buffers in accordance with the write access patterns.

3. The data processing system of claim 2 , wherein the corresponding ones of the consumers are configured to read the elements of the tensors from the buffers in accordance with the read access patterns.

4. The data processing system of claim 1 , wherein the compile time logic is further configured to detect a group of corresponding consumers that have a same read access pattern for processing the elements of the tensor.

5. The data processing system of claim 4 , wherein the compile time logic is further configured to insert a single buffer between the group of corresponding consumers and the particular producer.

6. The data processing system of claim 5 , wherein the particular producer is configured to write the elements of the tensor in the single buffer in accordance with the write access pattern.

7. The data processing system of claim 6 , wherein consumers in the group of corresponding consumers are configured to read the elements of the tensor from the single buffer in accordance with the read access pattern.

8. The data processing system of claim 1 , wherein the write access patterns and the read access patterns are defined based on operation types implemented by the operation units.

9. The data processing system of claim 1 , wherein the compile time logic is further configured to generate a modified version of the dataflow graph with buffers inserted between the certain ones of the producers and the corresponding ones of the consumers.

10. The data processing system of claim 9 , wherein runtime logic is configured to allocate physical compute units and physical memory units of a reconfigurable processor to the modified version of the dataflow graph.

11. The data processing system of claim 10 , wherein the runtime logic is configured to execute the modified version of the dataflow graph on the reconfigurable processor based on the allocation.

12. A computer-implemented method, including:

storing a dataflow graph for an application, the dataflow graph having operation units that are configured to be producers to produce tensors for execution of the application, and to be consumers to consume the tensors for execution of the application;

storing write access patterns of the producers, and read access patterns of the consumers, wherein a write access pattern of a particular producer specifies an order in which the particular producer generates elements of a tensor, and a read access pattern of a corresponding consumer specifies an order in which the corresponding consumer processes the elements of the tensor; and

detecting conflicts between certain ones of the producers and corresponding ones of the consumers that have mismatches between the write access patterns and the read access patterns, wherein a conflict occurs when the order in which the particular producer generates the elements of the tensor is different from the order in which the corresponding consumer processes the elements of the tensor; and

resolving the conflicts by inserting buffers between the certain ones of the producers and the corresponding ones of the consumers.

13. The computer-implemented method of claim 12 , wherein the certain ones of the producers are configured to write elements of the tensors in the buffers in accordance with the write access patterns.

14. The computer-implemented method of claim 13 , wherein the corresponding ones of the consumers are configured to read the elements of the tensors from the buffers in accordance with the read access patterns.

15. The computer-implemented method of claim 12 , wherein the compile time logic is further configured to detect a group of corresponding consumers that have a same read access pattern for processing the elements of the tensor.

16. The computer-implemented method of claim 15 , wherein the compile time logic is further configured to insert a single buffer between the group of corresponding consumers and the particular producer.

17. The computer-implemented method of claim 16 , wherein the particular producer is configured to write the elements of the tensor in the single buffer in accordance with the write access pattern.

18. A non-transitory computer readable storage medium impressed with computer program instructions to process data, the instructions, when executed on a processor, implement a method comprising:

storing a dataflow graph for an application, the dataflow graph having operation units that are configured to be producers to produce tensors for execution of the application, and to be consumers to consume the tensors for execution of the application;

storing write access patterns of the producers, and read access patterns of the consumers, wherein a write access pattern of a particular producer specifies an order in which the particular producer generates elements of a tensor, and a read access pattern of a corresponding consumer specifies an order in which the corresponding consumer processes the elements of the tensor; and

detecting conflicts between certain ones of the producers and corresponding ones of the consumers that have mismatches between the write access patterns and the read access patterns, wherein a conflict occurs when the order in which the particular producer generates the elements of the tensor is different from the order in which the corresponding consumer processes the elements of the tensor; and

resolving the conflicts by inserting buffers between the certain ones of the producers and the corresponding ones of the consumers.

19. The non-transitory computer readable storage medium of claim 18 , wherein the certain ones of the producers are configured to write elements of the tensors in the buffers in accordance with the write access patterns.

20. The non-transitory computer readable storage medium of claim 19 , wherein the corresponding ones of the consumers are configured to read the elements of the tensors from the buffers in accordance with the read access patterns.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2020
From: BROWN, KEVIN JAMES; KOEPLINGER, DAVID ALAN; CHEN, WEIWEI; GU, XIAOMING
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 053805/0186 →
Cited By (10)
US 12,189,570 US 12,197,379 US 12,380,060 US 12,413,530 US 12,430,109 US 12,547,581 US 12,572,341 US 12,602,349 US 12,681,806 US 12,705,205