IP Library Granted Patent US 9,583,116
Granted Patent B1
US 9,583,116 · App. 14/804,217 · Granted Feb 28, 2017

High-efficiency digital signal processing of streaming media

Inventors: Gabor Szanto (Kapolnasnyek, HU); Alexander Patrick Vlaskovits (Irvine, CA)
Assignee: Superpowered Inc.
G10L19/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,583,116
App. No.
14/804,217
Granted
Feb 28, 2017
Kind
B1
Abstract

A phase vocoder executes a fast-Fourier transform (FFT) with respect to an input audio data stream to generate an array of frequency-domain values corresponding to respective frequencies that are nominally uniformly distributed across a frequency range of interest, each of the frequency-domain values being representative of amplitude and phase of a spectral component of the input audio data stream at the respective frequency. The phase vocoder scales the nominally uniform distribution of the respective frequencies to reduce a cumulative error across the frequency distribution resulting from finite precision of a digital representation and then implements at least one of a time-stretching operation or a pitch-shifting operation with respect to the input data stream by manipulating the frequency-domain values with respect to one another within the array.

Claims (17)

1. A method of operation within a phase vocoder, the method comprising:

executing a fast-Fourier transform (FFT) with respect to an input audio data stream to generate an array of frequency-domain values corresponding to respective frequencies that are nominally uniformly distributed across a frequency range of interest, each of the frequency domain values being representative of amplitude and phase of a spectral component of the input audio data stream at the respective frequency;

scaling the nominally uniform distribution of the respective frequencies to reduce a cumulative error across the frequency distribution resulting from finite precision of a digital representation;

implementing at least one of a time-stretching operation or a pitch-shifting operation with respect to the input data stream by manipulating the frequency-domain values with respect to one another within the array; and

executing an inverse-FFT with respect to the manipulated frequency-domain values to generate a time-domain audio output.

2. The method of claim 1 wherein scaling the nominally uniform distribution of the respective frequencies to reduce a cumulative error comprises compressing the frequency distribution by omitting at least one of a division or a multiplication by Pi in the generation of the frequency distribution.

3. The method of claim 1 wherein at least one of executing the FFT, scaling the nominally uniform distribution of the respective frequencies, implementing the time-stretching operation, implementing the pitch-shifting operation, or executing the inverse-FFT comprises executing a sequence of program code instructions within one or more processors.

4. The method of claim 3 wherein executing the sequence of program code instructions comprises executing floating point operations in which Pi is implemented as a floating point number with finite precision such that, absent scaling the nominally uniform distribution, a substantial cumulative error results across the frequency distribution due to the finite precision of the floating point implementation of Pi.

5. The method of claim 1 wherein scaling the nominally uniform distribution of the respective frequencies comprises scaling by a factor of Pi such that at least one multiplication or division by Pi is obviated.

6. A phase vocoder comprising:

a Fourier transform engine to (i) generate, through execution of a fast-Fourier transform (FFT) with respect to an input audio data stream, an array of frequency-domain values corresponding to respective frequencies that are nominally uniformly distributed across a frequency range of interest, each of the frequency domain values being representative of amplitude and phase of a spectral component of the input audio data stream at the respective frequency, and (ii) to scale the nominally uniform distribution of the respective frequencies such that a cumulative error across the frequency distribution resulting from finite precision of a digital representation is reduced;

a scaling module to execute at least one of a time-stretching operation or a pitch-shifting operation with respect to the input data stream by manipulating the frequency-domain values with respect to one another within the array; and

an inverse Fourier transform engine to generate, through execution of an inverse-FFT with respect to the manipulated frequency-domain values, a time-domain audio output.

7. The phase vocoder of claim 6 wherein the Fourier transform engine to scale the nominally uniform distribution of the respective frequencies to reduce a cumulative error comprises logic to compress the frequency distribution by omitting at least one of a division or a multiplication by Pi in the generation of the frequency distribution.

8. The phase vocoder of claim 6 wherein at least one of the Fourier transform engine, scaling module and inverse Fourier transform engine comprises one or more programmed processors.

9. The phase vocoder of claim 8 wherein the one or more programmed processors comprises a processor and a memory having a sequence of program code instructions stored therein and wherein the processor executes floating point operations in which Pi is implemented as a floating point number with finite precision such that, absent scaling the nominally uniform distribution, a substantial cumulative error results across the frequency distribution due to the finite precision of the floating point implementation of Pi.

10. The phase vocoder of claim 6 wherein the Fourier transform engine to scale the nominally uniform distribution of the respective frequencies comprises logic to scale by a factor of Pi such that at least one multiplication or division by Pi is obviated.

Assignments (3)
SECURITY INTEREST Recorded Apr 21, 2025
From: DISTRIBUTED CREATION INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 070899/0867 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2024
From: SUPERPOWERED INC.
To: DISTRIBUTED CREATION INC.
Reel/Frame 069596/0684 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2015
From: SZANTO, GABOR; VLASKOVITS, ALEXANDER PATRICK
To: SUPERPOWERED INC.
Reel/Frame 036620/0128 →
Continuity (1)
Provisional Application 62027108 · Jul 21, 2014