IP Library Granted Patent US 12,315,513
Granted Patent B2
US 12,315,513 · App. 18/637,771 · Granted May 27, 2025

Dynamic service level assignment system for data processing manager

Inventors: Tim Stonehocker (Sunnyvale, CA); Zizo Gowayyed (San Francisco, CA); Seyed Majid Emami (Cupertino, CA); Matthias Eichstaedt (San Jose, CA); Evelyn Jiang (Cupertino, CA); Ryan Berryhill (Toronto, CA); Mathieu Ramona (Cachan, FR); Neil Veira (Toronto, CA)
G10L15/30G10L15/16G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,513
App. No.
18/637,771
Granted
May 27, 2025
Kind
B2
Abstract

A data processing system includes a queue manager receiving data processing requests and determining a queue depth representing the number of pending requests. A load supervisor assigns a service level to each request based on the queue depth when the request is at the head of the queue. The system offers two service levels, with the second level requiring fewer computing resources than the first. This dynamic management system optimizes resource allocation by adjusting service levels based on the workload, ensuring efficient processing of data requests.

Claims (49)

1. A system for managing processing of data, the system comprising:

a queue manager configured to receive a request to process data using a service, add the request to a queue of incoming requests, and determine a queue depth representing a number of requests in the queue at a given time; and

a load supervisor configured to receive the request and the queue depth from the queue manager at a time that the request is at a head of the queue and assign a service level for the request based on the queue depth at the time that the request is at the head of the queue;

wherein the service has a first service level and a second service level that uses less computing resource than the first service level.

2. The system of claim 1 , wherein the data comprises audio data and the service comprises automatic speech recognition.

3. The system of claim 1 , the load supervisor further configured to:

select the first service level for the request in response to the queue depth being below a first value; and

select the second service level for the request in response to the queue depth being above the first value.

4. The system of claim 3 , the load supervisor further configured to compute the first value based on an existing load on a server providing the service at the time that the request is at the head of the queue or on information received with the request.

5. The system of claim 1 , the load supervisor further configured to:

select the first service level for the request in response to the queue depth being below a first value;

and select the second service level for the request in response to the queue depth being above the first value and below a second value that is greater than the first value; and

select a third service level that uses fewer computing resources than the second service level for the request in response to the queue depth being above the second value.

6. The system of claim 1 , the load supervisor further configured to reject the request in response to the queue depth being above a predetermined value.

7. The system of claim 1 , further comprising a stream processor to process the data as a stream using the assigned service level.

8. The system of claim 1 , the load supervisor further configured to:

choose a server from a set of available servers providing the service to process the request; and

send the request to the chosen server with a tag indicating the assigned service level.

9. A method of managing a load for a server processing data, the method comprising:

receiving a request to process data using a service having a first service level and a second service level that uses fewer computing resources than the first service level;

adding the request to a queue of incoming requests, the queue having a queue depth representing a number of requests in the queue at a given time;

determining the queue depth at a time that the request is at a head of the queue; and

selecting between the first service level and the second service level for the request to establish a selected service level based at least in part on the queue depth.

10. The method of claim 9 , wherein the data comprises audio data and the request comprises a request for automatic speech recognition.

11. The method of claim 9 , further comprising:

selecting the first service level for the request in response to the queue depth being below a first value; and

selecting the second service level for the request in response to the queue depth being above the first value.

12. The method of claim 11 , wherein the first value is set to a fixed value at a configuration time.

13. The method of claim 11 , wherein the first value is computed based on an existing load on a server providing the service at the time that the request is at the head of the queue.

14. The method of claim 11 , wherein the first value is computed based on information received with the request.

15. The method of claim 9 , further comprising:

selecting the first service level for the request in response to the queue depth being below a first value;

selecting the second service level for the request in response to the queue depth being above the first value and below a second value that is greater than the first value; and

selecting a third service level that uses fewer computing resources than the second service level for the request in response to the queue depth being above the second value.

16. The method of claim 9 , further comprising:

receiving a second request to process second data;

adding the second request to the queue of incoming requests;

determining the queue depth at a second time that the request is at the head of the queue; and

rejecting the request in response to the queue depth being above a predetermined value.

17. The method of claim 9 , further comprising processing the data using a selected service level, wherein a single computing system manages the queue, establishes the selected service level for the request, and processes the data.

18. The method of claim 9 , wherein a load balancing system manages the queue and performs said selecting, the method further comprising:

choosing, by the load balancing system, a server from a set of available servers providing the service to process the request; and

sending the request to the chosen server with a tag indicating a selected service level.

19. The method of claim 9 , wherein a length of the data is unknown at the time that the request is at the head of the queue.

20. A non-transitory machine readable medium comprising one or more instructions that in response to being executed on a computing device cause the computing device to carry out a method comprising:

receiving a request to process data using a service having a first service level and a second service level that uses fewer computing resources than the first service level;

adding the request to a queue of incoming requests, the queue having a queue depth representing a number of requests in the queue at a given time;

determining the queue depth at a time that the request is at a head of the queue; and

selecting between the first service level and the second service level for the request based at least in part on the queue depth.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: VEIRA, NEIL; BERRYHILL, RYAN; EMAMI, SEYED MAJID; STONEHOCKER, TIMOTHY P.; GOWAYYED, ZIZU; JIANG, EVELYN; EICHSTAEDT, MATTHIAS; RAMONA, MATHIEU
To: SOUNDHOUND, INC.
Reel/Frame 067190/0402 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 067769/0712 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 067769/0768 →
Continuity (2)
Continuation 17447823 · Sep 16, 2021
Related Publication 20240296844A1 · Sep 5, 2024
References Cited (13)
US 9575963B2 · Pasupalak · 2017 [cited by examiner]
US 10418032B1 · Mohajer · 2019 [cited by examiner]
US 11594221B2 · Thomson · 2023 [cited by examiner]
US 11978454B2 · Stonehocker · 2024 [cited by examiner]
US 20150066479A1 · Pasupalak · 2015 [cited by examiner]
US 20180182398A1 · Halstvedt · 2018 [cited by examiner]
US 20210233530A1 · Thomson · 2021 [cited by examiner]
US 20230082955A1 · Stonehocker · 2023 [cited by examiner]
US 20240296844A1 · Stonehocker · 2024 [cited by examiner]
Daghero, Francesco, Energy-Efficient Quality Adaptation for Recurrent Neural Networks, 2019. [cited by applicant]
Hinton, Geoffrey, et al., Deep Neural Networks for Acoustic Modeling in Speech Recognition, IEEE Signal Processing Magazine, Nov. 2012. [cited by applicant]
Nedevschi, et al., Hardware Speech Recognition for User Interfaces in Low Cost, Low Power Devices, DAC 2005, Jun. 13-17, 2005. [cited by applicant]
You, et al., Openmp-Based Parallel Implementation of a Continuous Speech Recognizer on a Multi-Core System, ICASSP 2009. [cited by applicant]