IP Library Granted Patent US 12,353,793
Granted Patent B2
US 12,353,793 · App. 18/676,243 · Granted Jul 8, 2025

Dynamically preventing audio artifacts

Inventors: Utkarsh Vaidya (Santa Clara, CA); Sumit Bhattacharya (Santa Clara, CA)
Assignee: NVIDIA Corporation
G06F3/165G06F3/162G06N3/045G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,793
App. No.
18/676,243
Granted
Jul 8, 2025
Kind
B2
Abstract

The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.

Claims (27)

1. A method of processing computer generated audio, comprising:

determining, using a neural network, a probability of an audio artifact during output of an audio signal generated using a computing system, wherein the determining the probability of the audio artifact is based at least on a processing frequency of a processor or processing unit of the system; and

adjusting at least one of the processing frequency or a memory frequency of the system to lower the probability of an audio artifact occurring when the probability is above a threshold.

2. The method of claim 1 , wherein the system is a part of a mobile device, and the audio signal is a streamed audio signal.

3. The method of claim 1 , wherein the audio signal includes at least one synthetically generated spoken phrase.

4. The method of claim 1 , wherein the adjusting includes increasing the processing frequency until the processing frequency reaches a maximum processing frequency value.

5. The method of claim 1 , wherein the adjusting includes increasing the memory frequency until the memory frequency reaches a maximum memory frequency value.

6. The method of claim 1 , further comprising reversing a prior adjustment of the system when the probability of an audio artifact occurring is not higher than the threshold.

7. The method of claim 1 , wherein the neural network is a deep neural network.

8. The method of claim 1 , wherein the adjusting includes increasing the processing frequency and the method further includes adjusting a size of an audio buffer of the system in response to increasing the processing frequency.

9. A method of processing computer generated audio, comprising:

determining, using a neural network, a probability of an audio artifact during output of an audio signal generated using a computing system, wherein the determining the probability of the audio artifact is based at least on a memory frequency of the system; and

adjusting at least one of the memory frequency of the system or a processing frequency of the system to lower the probability of an audio artifact occurring when the probability is above a threshold.

10. The method of claim 9 , wherein the system is a part of a mobile device, and the audio signal is a streamed audio signal.

11. The method of claim 9 , wherein the audio signal includes at least one synthetically generated spoken phrase.

12. The method of claim 9 , wherein the adjusting includes increasing the processing frequency until the processing frequency reaches a maximum processing frequency value.

13. The method of claim 9 , wherein the adjusting includes increasing the memory frequency until the memory frequency reaches a maximum memory frequency value.

14. The method of claim 9 , further comprising reversing a prior adjustment of the system when the probability of an audio artifact occurring is not higher than the threshold.

15. The method of claim 9 , wherein the neural network is a deep neural network.

16. A computing system comprising:

one or more processing units to perform operations that include:

determining, using a neural network, a probability of an audio artifact during output of a computer generated audio signal, wherein the determining the probability of the audio artifact is based at least on a memory frequency of the computing system; and

adjusting at least one of a processing frequency of the computing system or the memory frequency when the probability is higher than a threshold.

17. The computing system of claim 16 , wherein the neural network is a deep neural network.

18. The computing system of claim 16 , wherein the computing system is part of a mobile device, and the audio signal is a streamed audio signal.

19. The computing system of claim 16 , wherein the computing system is a vehicle computing system.

20. The computing system of claim 16 , wherein the audio signal includes one or more synthetically generated spoken phrases.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2024
From: VAIDYA, UTKARSH; BHATTACHARYA, SUMIT
To: NVIDIA CORPORATION
Reel/Frame 067543/0935 →
Continuity (4)
Continuation 18161326 · Jan 30, 2023
Continuation 17121373 · Dec 14, 2020
Continuation 16285941 · Feb 26, 2019
Related Publication 20240311080A1 · Sep 19, 2024
References Cited (26)
US 6594213B1 · Hayashi · 2003 [cited by applicant]
US 6901039B1 · Sugie et al. · 2005 [cited by applicant]
US 7236436B2 · Hayashi · 2007 [cited by applicant]
US 7737960B2 · Gong et al. · 2010 [cited by applicant]
US 8209731B2 · Lee et al. · 2012 [cited by applicant]
US 8331385B2 · Black et al. · 2012 [cited by applicant]
US 8527075B2 · Julian et al. · 2013 [cited by applicant]
US 9070198B2 · Carter et al. · 2015 [cited by applicant]
US 9985887B2 · Roncero Izquierdo et al. · 2018 [cited by applicant]
US 10236031B1 · Gurijala · 2019 [cited by applicant]
US 20030031336A1 · Harrison et al. · 2003 [cited by applicant]
US 20030156639A1 · Liang · 2003 [cited by applicant]
US 20070183192A1 · Barnum et al. · 2007 [cited by applicant]
US 20070203597A1 · Iyoshi · 2007 [cited by applicant]
US 20110043694A1 · Izuno et al. · 2011 [cited by applicant]
US 20120009892A1 · Yen et al. · 2012 [cited by applicant]
US 20170063692A1 · Izquierdo et al. · 2017 [cited by applicant]
US 20190373490A1 · Rahmati · 2019 [cited by examiner]
US 20200272409A1 · Vaidya et al. · 2020 [cited by applicant]
GB 2331678A · 1999 [cited by applicant]
JP 2002268662A · 2002 [cited by applicant]
JP 2007194845A · 2007 [cited by applicant]
KR 20090050768A · 2009 [cited by applicant]
Ivo Creusen, et al.; “Control of jitter buffer size using machine learning”; Technical Disclosure Commons, Defensive Publications Series; http://www.tdommons.org/dpubs_series; Dec. 6, 2017; 7 pgs. [cited by applicant]
“Audio Class Stream Data Flow”; Micrium Documentation; μC/USB Device Documentation V4.05; https://doc.micrium.com/display/USBDDOCV405/Audio+Class+Stream+Data+Flow; 2019; 8 pgs. [cited by applicant]
“Decrease Underrun”; Matlab & Simulink—MathWorks India; The MathWords, Inc.; https://in.mathworks.com/help/audio/ug;decrease-underrun.html; 2016; 6 pgs. [cited by applicant]