IP Library › Granted Patent US 12,340,790
Granted Patent B2
US 12,340,790 · App. 18/194,829 · Granted Jun 24, 2025

Dynamic tempered sampling in generative models inference

Inventor: Pablo Barrera Gonzalez (Mountain View, CA)
Assignee: Google LLC
G10L13/047G06N3/04G10L13/00G10L19/005G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,790
App. No.
18/194,829
Granted
Jun 24, 2025
Kind
B2
Abstract

A method of sampling output audio samples includes, during a packet loss concealment event, obtaining a sequence of previous output audio samples. At each time step during the event, the method includes generating a probability distribution over possible output audio samples for the time step. Each sample includes a respective probability indicating a likelihood that the corresponding sample represents a portion of an utterance at the time step. The method also includes determining a temperature sampling value based on a function of a number of time steps that precedes the time step, and an initial, a minimum, and a maximum temperature sampling value. The method also includes applying the temperature sampling value to the probability distribution to adjust a probability of selecting possible samples and randomly selecting one of the possible samples based on the adjusted probability. The method also includes generating synthesized speech using the randomly selected sample.

Claims (42)

1. A computer-implemented method executed by data processing hardware that causes the data processing hardware to perform operations comprising:

during a packet loss concealment event in an active voice communication session:

obtaining a sequence of previous output audio samples during a time window having a start time and an end time, the end time occurring when the packet loss concealment event commences; and

at each time step of a plurality of time steps during the packet loss concealment event:

generating, using a speech synthesis model, a probability distribution over possible output audio samples for the corresponding time step, each possible output audio sample in the probability distribution comprising a respective probability indicating a likelihood that the corresponding possible output audio sample represents a portion of an utterance at the corresponding time step;

determining a dynamic temperature sampling value based on a function of a number of time steps in the plurality of time steps that precedes the corresponding time step, the dynamic temperature sampling value increasing as the number of time steps in the plurality of time steps preceded the corresponding time step during the packet loss concealment event increases;

applying the dynamic temperature sampling value to the probability distribution to adjust a probability of selecting possible output audio samples from the probability distribution;

selecting one of the possible output audio samples of the probability distribution based on the adjusted probability associated with each of the possible output audio samples; and

generating synthesized speech using the selected output audio sample.

2. The method of claim 1 , selecting one of the possible output audio samples of the probability distribution based on the adjusted probability associated with each of the possible output audio samples comprises a random selection.

3. The method of claim 1 , wherein determining the dynamic temperature sampling value is further based on an initial temperature sampling value, a minimum temperature sampling value, and a maximum temperature sampling value.

4. The method of claim 3 , wherein the maximum temperature sampling value is 0.85.

5. The method of claim 3 , wherein the minimum temperature sampling value is 0.25.

6. The method of claim 3 , wherein the initial temperature sampling value is the same as the minimum temperature sampling value.

7. The method of claim 1 , wherein determining the dynamic temperature sampling value comprises:

determining the number of time steps in the plurality of time steps that preceded the corresponding time step during the packet loss concealment event; and

when the number of time steps satisfies a threshold, increasing the dynamic temperature sampling value by a set amount.

8. The method of claim 7 , wherein the threshold a multiple of ten time steps.

9. The method of claim 7 , wherein the set amount is 0.1.

10. The method of claim 1 , wherein determining the dynamic temperature sampling value comprises increasing the dynamic temperature sampling value based on the number of time steps in the plurality of time steps preceded the corresponding time step during the packet loss concealment event.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

during a packet loss concealment event in an active voice communication session:

obtaining a sequence of previous output audio samples during a time window having a start time and an end time, the end time occurring when the packet loss concealment event commences; and

at each time step of a plurality of time steps during the packet loss concealment event:

generating, using a speech synthesis model, a probability distribution over possible output audio samples for the corresponding time step, each possible output audio sample in the probability distribution comprising a respective probability indicating a likelihood that the corresponding possible output audio sample represents a portion of an utterance at the corresponding time step;

determining a dynamic temperature sampling value based on a function of a number of time steps in the plurality of time steps that precedes the corresponding time step, the dynamic temperature sampling value increasing as the number of time steps in the plurality of time steps preceded the corresponding time step during the packet loss concealment event increases;

applying the dynamic temperature sampling value to the probability distribution to adjust a probability of selecting possible output audio samples from the probability distribution;

selecting one of the possible output audio samples of the probability distribution based on the adjusted probability associated with each of the possible output audio samples; and

generating synthesized speech using the selected output audio sample.

12. The system of claim 11 , selecting one of the possible output audio samples of the probability distribution based on the adjusted probability associated with each of the possible output audio samples comprises a random selection.

13. The system of claim 11 , wherein determining the dynamic temperature sampling value is further based on an initial temperature sampling value, a minimum temperature sampling value, and a maximum temperature sampling value.

14. The system of claim 13 , wherein the maximum temperature sampling value is 0.85.

15. The system of claim 13 , wherein the minimum temperature sampling value is 0.25.

16. The system of claim 13 , wherein the initial temperature sampling value is the same as the minimum temperature sampling value.

17. The system of claim 11 , wherein determining the dynamic temperature sampling value comprises:

determining the number of time steps in the plurality of time steps that preceded the corresponding time step during the packet loss concealment event; and

when the number of time steps satisfies a threshold, increasing the dynamic temperature sampling value by a set amount.

18. The system of claim 17 , wherein the threshold a multiple of ten time steps.

19. The system of claim 17 , wherein the set amount is 0.1.

20. The system of claim 11 , wherein determining the dynamic temperature sampling value comprises increasing the dynamic temperature sampling value based on the number of time steps in the plurality of time steps preceded the corresponding time step during the packet loss concealment event.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2023
From: GONZALEZ, PABLO BARRERA
To: GOOGLE LLC
Reel/Frame 063216/0174 →
Continuity (2)
Continuation 16718333 · Dec 18, 2019
Related Publication 20230237986A1 · Jul 27, 2023
References Cited (18)
US 10127918B1 · Kamath Koteshwara · 2018 [cited by examiner]
US 10546066B2 · Li · 2020 [cited by examiner]
US 20170011738A1 · Senior · 2017 [cited by examiner]
US 20190205748A1 · Fukuda · 2019 [cited by examiner]
US 20190244604A1 · Masataki · 2019 [cited by examiner]
US 20190362229A1 · Norouzi · 2019 [cited by examiner]
US 20200035223A1 · Asami · 2020 [cited by examiner]
US 20200160843A1 · Shillingford · 2020 [cited by examiner]
US 20210073438A1 · Akiyama · 2021 [cited by examiner]
US 20210082399A1 · Kurata · 2021 [cited by examiner]
US 20210117786A1 · Schwarz · 2021 [cited by examiner]
US 20220001078A1 · Siondalski · 2022 [cited by examiner]
US 20230237986A1 · Gonzalez · 2023 [cited by examiner]
WO 2019213021A1 · 2019 [cited by applicant]
Marco Lippi , Marcelo A. Montemurro , Mirko Degli Esposti, and Giampaolo Cristadoro; Natural Language Statistical Features of LSTM-Generated Texts; Nov. 2019; URL: https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=86… [cited by examiner]
International Search Report and Written Opinion for the related International Application No. PCT/US2020/065638, Dated Apr. 16, 2021, 14 pages. [cited by applicant]
Bong-Ki Lee et al: “Packet loss concealment based on deep neural networks for digital speech transmission”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, IEEE, USA, vol. 24, No. 2, Feb. 1, 2016 (Feb. … [cited by applicant]
Sercan O Arik et al: “Deep Voice: Real-time Neural Text-to-Speech”, Feb. 24, 2017, XP055489867, Retrieved from the Internet:URL:https://arxiv.org/pdf/1702. 07825.pdf., 17 pages. [cited by applicant]