IP Library › Granted Patent US 12,586,600
Granted Patent B2
US 12,586,600 · App. 18/163,848 · Granted Mar 24, 2026

Streaming vocoder

Inventors: Oleg Rybakov (Mountain View, CA); Liyang Jiang (Mountain View, CA); Fadi Biadsy (Mountain View, CA)
Assignee: Google LLC
G10L21/10G10L21/003G10L21/0232G10L21/0364G10L21/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,600
App. No.
18/163,848
Granted
Mar 24, 2026
Kind
B2
Abstract

A method includes receiving a current spectrogram frame and reconstructing a phase of the current spectrogram frame by, for each corresponding committed spectrogram frame in a sequence of M number of committed spectrogram frames preceding the current spectrogram frame, obtaining a value of a committed phase of the corresponding committed spectrogram frame and estimating the phase of the current spectrogram frame based on a magnitude of the current spectrogram frame and the value of the committed phase of each corresponding committed spectrogram frame in the sequence of M number of committed spectrogram frames preceding the current spectrogram frame. The method also includes synthesizing, for the current spectrogram frame, a new time-domain audio waveform frame based on the estimated phase of the current spectrogram frame.

Claims (50)

1 . A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:

receiving a current spectrogram frame;

reconstructing a phase of the current spectrogram frame by:

for each corresponding committed spectrogram frame in a sequence of M number of committed spectrogram frames preceding the current spectrogram frame, obtaining a value of a committed phase of the corresponding committed spectrogram frame; and

estimating the phase of the current spectrogram frame by performing one or more iterations within a sliding window that contains the current spectrogram frame, wherein performing each iteration of the one or more iterations within the sliding window comprises:

estimating an uncommitted phase of the current spectrogram frame based on a sequence of N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame; and

updating a complex-valued spectrogram representation within the sliding window by combining the value of the committed phase of each corresponding committed spectrogram frame in the sequence of M number of committed spectrogram frames preceding the current spectrogram frame, the estimated uncommitted phase, and a magnitude of the current spectrogram frame; and

for the current spectrogram frame, synthesizing a new time-domain audio waveform frame based on the estimated phase of the current spectrogram frame.

2 . The method of claim 1 , wherein:

the current spectrogram frame comprises a log-magnitude spectrogram frame output from a speech conversion model; and

prior to reconstructing the phase of the current spectrogram frame, the phase of the current spectrogram frame is initialized with a value equal to zero.

3 . The method of claim 1 , wherein the M number of committed spectrogram frames preceding the current spectrogram frame is equal to one.

4 . The method of claim 1 , wherein the M number of committed spectrogram frames preceding the current spectrogram frame is at least two.

5 . The method of claim 1 , wherein estimating the uncommitted phase of the current spectrogram frame based on the N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame comprises:

for each corresponding uncommitted spectrogram frame in the sequence of N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame, obtaining a value of an uncommitted phase of the corresponding uncommitted spectrogram frame; and

estimating the uncommitted phase of the current spectrogram frame is based on the value of the uncommitted phase of each corresponding uncommitted spectrogram frame in the sequence of N number of committed spectrogram frames within the sliding window that are subsequent to the current spectrogram frame.

6 . The method of claim 1 , wherein the N number of uncommitted spectrogram frames and the M number of committed spectrogram frames are equal.

7 . The method of claim 1 , wherein the N number of uncommitted spectrogram frames and the M number of committed spectrogram frames are different.

8 . The method of claim 1 , wherein the N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame is equal to one.

9 . The method of claim 1 , wherein the N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame is at least two.

10 . The method of claim 1 , wherein the current spectrogram frame is in a Short-time Fourier transform (STFT) domain when reconstructing the phase of the current spectrogram frame.

11 . The method of claim 10 , wherein synthesizing the new time-domain audio waveform frame based on the estimated phase of the current spectrogram frame comprises running a streaming inverse STFT on an output frame corresponding to the current spectrogram frame, the output frame extracted using the estimated phase of the current spectrogram frame.

12 . The method of claim 1 , wherein the operations further comprise, after reconstructing the phase of the current spectrogram frame, designating the current spectrogram frame as a committed frame and storing the estimated phase of the current spectrogram frame as a committed phase.

13 . The method of claim 1 , wherein the data processing hardware resides on a user computing device or a server.

14 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a current spectrogram frame;

reconstructing a phase of the current spectrogram frame by:

for each corresponding committed spectrogram frame in a sequence of M number of committed spectrogram frames preceding the current spectrogram frame, obtaining a value of a committed phase of the corresponding committed spectrogram frame; and

estimating the phase of the current spectrogram frame by performing one or more iterations within a sliding window that contains the current spectrogram frame, wherein performing each iteration of the one or more iterations within the sliding window comprises:

estimating an uncommitted phase of the current spectrogram frame based on a sequence of N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame; and

updating a complex-valued spectrogram representation within the sliding window by combining the value of the committed phase of each corresponding committed spectrogram frame in the sequence of M number of committed spectrogram frames preceding the current spectrogram frame, the estimated uncommitted phase, and a magnitude of the current spectrogram frame; and

for the current spectrogram frame, synthesizing a new time-domain audio waveform frame based on the estimated phase of the current spectrogram frame.

15 . The system of claim 14 , wherein:

the current spectrogram frame comprises a log-magnitude spectrogram frame output from a speech conversion model; and

prior to reconstructing the phase of the current spectrogram frame, the phase of the current spectrogram frame is initialized with a value equal to zero.

16 . The system of claim 14 , wherein the M number of committed spectrogram frames preceding the current spectrogram frame is equal to one.

17 . The system of claim 14 , wherein the M number of committed spectrogram frames preceding the current spectrogram frame is at least two.

18 . The system of claim 14 , wherein estimating the uncommitted phase of the current spectrogram frame based on the N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame comprises:

for each corresponding uncommitted spectrogram frame in the sequence of N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame, obtaining a value of an uncommitted phase of the corresponding uncommitted spectrogram frame; and

estimating the uncommitted phase of the current spectrogram frame is further-based on the value of the uncommitted phase of each corresponding uncommitted spectrogram frame in the sequence of N number of committed spectrogram frames within the sliding window that are subsequent to the current spectrogram frame.

19 . The system of claim 14 , wherein the N number of uncommitted spectrogram frames and the M number of committed spectrogram frames are equal.

20 . The system of claim 14 , wherein the N number of uncommitted spectrogram frames and the M number of committed spectrogram frames are different.

21 . The system of claim 14 , wherein the N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame is equal to one.

22 . The system of claim 14 , wherein the N number of uncommitted spectrogram frames within the sliding window that are subsequent to the current spectrogram frame is at least two.

23 . The system of claim 14 , wherein the current spectrogram frame is in a Short-time Fourier transform (STFT) domain when reconstructing the phase of the current spectrogram frame.

24 . The system of claim 23 , wherein synthesizing the new time-domain audio waveform frame based on the estimated phase of the current spectrogram frame comprises running a streaming inverse STFT on an output frame corresponding to the current spectrogram frame, the output frame extracted using the estimated phase of the current spectrogram frame.

25 . The system of claim 14 , wherein the operations further comprise, after reconstructing the phase of the current spectrogram frame, designating the current spectrogram frame as a committed frame and storing the estimated phase of the current spectrogram frame as a committed phase.

26 . The system of claim 14 , wherein the data processing hardware resides on a user computing device or a server.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2023
From: RYBAKOV, OLEG; JIANG, LIYANG; BIADSY, FADI
To: GOOGLE LLC
Reel/Frame 062578/0019 →
Continuity (2)
Provisional Application 63312195 · Feb 21, 2022
Related Publication 20230267949A1 · Aug 24, 2023
References Cited (16)
US 10529349B2 · Le Roux · 2020 [cited by examiner]
US 10726856B2 · Le Roux · 2020 [cited by examiner]
US 11017763B1 · Aggarwal · 2021 [cited by examiner]
US 11462209B2 · Arik · 2022 [cited by examiner]
US 11776528B2 · Kang · 2023 [cited by examiner]
WO WO2020171868A1 · 2020 [cited by examiner]
Zhu, Xinglei, Gerald T. Beauregard, and Lonce L. Wyse, “Real-Time Signal Estimation From Modified Short-Time Fourier Transform Magnitude Spectra”, Jul. 2007, IEEE Transactions on Audio, Speech, and Language Processing, … [cited by examiner]
Rybakov, Oleg, Marco Tagliasacchi, Yunpeng Li, Liyang Jiang, Xia Zhang, and Fadi Biadsy, “Real time spectrogram inversion on mobile phone”, Aug. 2023, Interspeech 2023, pp. 4314-4118. (Year: 2023). [cited by examiner]
Biadsy, Fadi, Ron J. Weiss, Pedro J. Moreno, Dimitri Kanvesky, and Ye Jia, “Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech Separation”, Sep. 2019, I… [cited by examiner]
Průša, Zdeněk, Peter Balazs, and Peter Lempel Søndergaard, “A Noniterative Method for Reconstruction of Phase From STFT Magnitude”, May 2017, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, No.… [cited by examiner]
Průša, Zdeněk, and Pavel Rajmic, “Towards High Quality Real-Time Signal Reconstruction from STFT Magnitude”, Apr. 2017, IEEE Signal Processing Letters, vol. 24, No. 6, pp. 892-896. (Year: 2017). [cited by examiner]
Beauregard, Gerald T., Mithila Harish, and Lonce Wyse, “Single Pass Spectrogram Inversion”, Jul. 2015, 2015 IEEE International Conference on Digital Signal Processing (DSP), pp. 427-431. (Year: 2015). [cited by examiner]
Gerkmann, Timo, Martin Krawczyk-Becker, and Jonathan Le Roux, “Phase Processing for Single Channel Speech Enhancement: History and Recent Advances”, Mar. 2015, IEEE Signal Processing Magazine, vol. 32, No. 2, pp. 55-66.… [cited by examiner]
Gnann, Volker, and Martin Spiertz, “Inversion of short-time fourier transform magnitude spectrograms with adaptive window lengths”, Apr. 2009, 2009 IEEE International Conference on Acoustics, Speech and Signal Processin… [cited by examiner]
International Search Report and Written Opinion for the related Application No. PCT/US2023/012239, dated May 9, 2023, 17 pages. [cited by applicant]
Xinglei Zhu et al: “Real-Time Iterative Spectrum Inversion with Look- Ahead”, Proceedings / 2006 IEEE International Conference on Multimedia and Expo, ICME 2006 :Jul. 9-12, 2006, Hilton, Toronto, Toronto, Ontario, Canad… [cited by applicant]