IP Library Granted Patent US 12,632,290
Granted Patent B2
US 12,632,290 · App. 17/559,612 · Granted May 19, 2026

Method and apparatus for dynamically adjusting pipeline depth to improve execution latency

Inventors: Saurabh Gayen (Portland, OR); Dhananjay Joshi (Portland, OR); Philip Lantz (Cornelius, OR); Rajesh Sankaran (Portland, OR); Narayan Ranganathan (Bangalore, IN)
Assignee: Intel Corporation
G06F9/4881G06F9/30079G06F9/45558G06F9/485G06F2009/45591
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,290
App. No.
17/559,612
Granted
May 19, 2026
Kind
B2
Abstract

Apparatus and method for managing pipeline depth of a data processing device. For example, one embodiment of an accelerator comprises: an interface to receive a plurality of work requests from a plurality of clients; and a plurality of engines to perform the plurality of work requests; wherein the work requests are to be dispatched to the plurality of engines from work queues, the work queues to store a work descriptor per work request, each work descriptor to include information needed to perform a corresponding work request, wherein the work queues include a first and a second work queue to store work descriptors associated with first latency characteristics and second latency characteristics, respectively; engine configuration circuitry to configure a first engine to have a first pipeline depth based on the first latency characteristics and to configure a second engine to have a second pipeline depth based on the second latency characteristics.

Claims (37)

1 . An accelerator comprising:

an interface configured to receive work requests from a plurality of clients;

a plurality of engines configured to perform the work requests, wherein the work requests are to be dispatched to the plurality of engines from a plurality of work queues, the plurality of work queues to store a work descriptor per work request, each work descriptor corresponding to a work request and including an identifier of a client and one or more privileges of the client, and wherein the plurality of work queues include a first work queue to store work descriptors associated with work requests that share first latency characteristics and a second work queue to store work descriptors associated with work requests that share second latency characteristics; and

engine configuration circuitry configured to cause a first engine to have a first pipeline with a first pipeline depth based on the first latency characteristics and to cause a second engine to have a second pipeline with a second pipeline depth based on the second latency characteristics, the first pipeline depth being longer than the second pipeline depth with more slots in the first pipeline for work descriptors based on the first latency characteristics including being less sensitive to latency than the second latency characteristics.

2 . The accelerator of claim 1 wherein the first and second latency characteristics are obtained based on data received by the engine configuration circuitry from a client associated with the work requests.

3 . The accelerator of claim 2 wherein the client associated with the work requests comprises an application, virtual machine, or supervisory application.

4 . The accelerator of claim 1 wherein the first and second latency characteristics comprise latency requirements associated with the corresponding work requests.

5 . The accelerator of claim 1 wherein the first and second latency characteristics comprise a maximum allowable latency value and/or a desired latency value.

6 . The accelerator of claim 1 wherein the first latency characteristics comprise a first latency value and the second latency characteristics comprise a second latency value larger than the first latency value, then the engine configuration circuitry is to configure the second pipeline depth to be deeper than the first pipeline depth.

7 . The accelerator of claim 1 wherein the engine configuration circuitry is to configure the first and second pipeline depths of the first engine and the second engine, respectively, based further on first and second throughput values associated with the corresponding work requests.

8 . The accelerator of claim 7 wherein the first and second throughput values each comprise a minimum allowable throughput value and/or a desired throughput value.

9 . A method comprising:

receiving, by an accelerator, work requests from a plurality of clients;

performing the work requests on a plurality of engines of the accelerator;

dispatching the work requests to the plurality of engines from a plurality of work queues, the plurality of work queues to store a work descriptor per work request, each work descriptor corresponding to a work request and including an identifier of a client and one or more privileges of the client, the plurality of work queues including a first work queue to store work descriptors associated with work requests that share first latency characteristics and a second work queue to store work descriptors associated with work requests that share second latency characteristics;

configuring, by engine configuration circuitry, a first engine to have a first pipeline with a first pipeline depth based on the first latency characteristics and a second engine to have a second pipeline with a second pipeline depth based on the second latency characteristics, the first pipeline depth being longer than the second pipeline depth with more slots in the first pipeline for work descriptors based on the first latency characteristics including being less sensitive to latency than the second latency characteristics; and

processing work descriptors in the first work queue by the first engine in the first pipeline and processing work descriptors in the second work queue by the second engine in the second pipeline.

10 . The method of claim 9 wherein the first and second latency characteristics are obtained based on data received by the engine configuration circuitry from a client associated with the work requests.

11 . The method of claim 10 wherein the client associated with the work requests comprises an application, virtual machine, or supervisory application.

12 . The method of claim 9 wherein the first and second latency characteristics comprise latency requirements associated with the corresponding work requests.

13 . The method of claim 9 wherein the first and second latency characteristics comprise a maximum allowable latency value and/or a desired latency value.

14 . The method of claim 9 wherein the first latency characteristics comprise a first latency value and the second latency characteristics comprise a second latency value larger than the first latency value, then the engine configuration circuitry is to configure the second pipeline depth to be deeper than the first pipeline depth.

15 . The method of claim 9 wherein the engine configuration circuitry is to configure the first and second pipeline depths of the first engine and the second engine, respectively, based further on first and second throughput values associated with the corresponding work requests.

16 . The method of claim 15 wherein the first and second throughput values each comprise a minimum allowable throughput value and/or a desired throughput value.

17 . A non-transitory computer-readable medium having program code stored thereon which, when executed by a machine, causes the machine to perform:

receiving, by an accelerator, work requests from a plurality of clients;

performing the work requests on a plurality of engines of the accelerator;

dispatching the work requests to the plurality of engines from a plurality of work queues, the plurality of work queues to store a work descriptor per work request, each work descriptor corresponding to a work request and including an identifier of a client and one or more privileges of the client, the plurality of work queues including a first work queue to store work descriptors associated with work requests that share first latency characteristics and a second work queue to store work descriptors associated with work requests that share second latency characteristics;

configuring, by engine configuration circuitry, a first engine to have a first pipeline with a first pipeline depth based on the first latency characteristics and a second engine to have a second pipeline with a second pipeline depth based on the second latency characteristics, the first pipeline depth being longer than the second pipeline depth with more slots in the first pipeline for work descriptors based on the first latency characteristics including being less sensitive to latency than the second latency characteristics; and

processing work descriptors in the first work queue by the first engine in the first pipeline and processing work descriptors in the second work queue by the second engine in the second pipeline.

18 . The non-transitory computer-readable medium of claim 17 wherein the first and second latency characteristics are obtained based on data received by the engine configuration circuitry from a client associated with the work requests.

19 . The non-transitory computer-readable medium of claim 18 wherein the client associated with the work requests comprises an application, virtual machine, or supervisory application.

20 . The non-transitory computer-readable medium of claim 17 wherein the first and second latency characteristics comprise latency requirements associated with the corresponding work requests.

21 . The non-transitory computer-readable medium of claim 17 wherein the first and second latency characteristics comprise a maximum allowable latency value and/or a desired latency value.

22 . The non-transitory computer-readable medium of claim 17 wherein the first latency characteristics comprise a first latency value and the second latency characteristics comprise a second latency value larger than the first latency value, then the engine configuration circuitry is to configure the second pipeline depth to be deeper than the first pipeline depth.

23 . The non-transitory computer-readable medium of claim 17 wherein the engine configuration circuitry is to configure the first and second pipeline depths of the first engine and the second engine, respectively, based further on first and second throughput values associated with the corresponding work requests.

24 . The non-transitory computer-readable medium of claim 23 wherein the first and second throughput values each comprise a minimum allowable throughput value and/or a desired throughput value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2022
From: GAYEN, SAURABH; JOSHI, DHANANJAY; LANTZ, PHILIP; SANKARAN, RAJESH; RANGANATHAN, NARAYAN
To: INTEL CORPORATION
Reel/Frame 060151/0432 →
Continuity (2)
Provisional Application 63226159 · Jul 27, 2021
Related Publication 20230040226A1 · Feb 9, 2023
References Cited (16)
US 11106613B2 · Lantz et al. · 2021 [cited by applicant]
US 20070271449A1 · Lichtensteiger · 2007 [cited by examiner]
US 20090060197A1 · Taylor · 2009 [cited by examiner]
US 20130152099A1 · Bass · 2013 [cited by examiner]
US 20190303324A1 · Lantz et al. · 2019 [cited by applicant]
US 20200004703A1 · Sankaran et al. · 2020 [cited by applicant]
US 20200401440A1 · Sankaran · 2020 [cited by examiner]
US 20210014177A1 · Kasichainula · 2021 [cited by examiner]
EP 3547130A1 · 2019 [cited by applicant]
Notification of Oral Proceeding, EP App. No. 22181003.9, Jul. 23, 2024, 8 pages. [cited by applicant]
Anonymous: “Intel Data Streaming Accelerator Preliminary Architecture Specification”, Order No. 341204-001US, Revision: 1.0, Nov. 2019, pp. 41 (Part II) XP055692158. [cited by applicant]
Anonymous: “Intel Data Streaming Accelerator Preliminary Architecture Specification”, Order No. 341204-001US, Revision: 1.0, Nov. 2019, pp. 84 (Part I) XP055692158. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 22181003.9, Dec. 13, 2022, 9 pages. [cited by applicant]
Brief Communication, EP App. No. 22181003.9, May 6, 2025, 2 pages. [cited by applicant]
Office Action, EP App. No. 22181003.9, Dec. 20, 2023, 4 pages. [cited by applicant]
Intention to Grant, EP App. No. 22181003.9, Jun. 13, 2025, 8 pages. [cited by applicant]