IP Library › Granted Patent US 12,315,490
Granted Patent B2
US 12,315,490 · App. 17/565,826 · Granted May 27, 2025

Text-to-speech and speech recognition for noisy environments

Inventors: Daniel Bromand (Boston, MA); Björn Erik Roth (Stockholm, SE); Kåre Sjölander (Stockholm, SE)
Assignee: Spotify AB
G10L13/033G10L13/08G10L15/08G10L15/22G10L25/84G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,490
App. No.
17/565,826
Granted
May 27, 2025
Kind
B2
Abstract

The present disclosure relates generally to speech processing. Humans change their speech patterns in noisy environments. The systems and devices described herein can compensate for noisy environments to be more human-like. Thus, the configurations and implementations herein can determine a sound profile for the sound environment where the user is listening. Based on the sound profile, the devices can determine a transform to apply to output speech from the device. This transform is applied to the wake word, speech recognition, and to the output speech to compensate for the noise level of the environment by mimicking the Lombard effect.

Claims (63)

1. A method comprising:

providing, by a media delivery system, a first audio representation of a first simulated sound environment;

receiving, by the media delivery system, first speech from a user speaking subject to audio playout of in the first simulated sound environment;

providing, by the media delivery system, a second audio representation of a second simulated sound environment, wherein the second simulated sound environment has different acoustic characteristics than the first simulated sound environment;

receiving, by the media delivery system, second speech from the user speaking subject to audio playout of the second simulated sound environment;

determining a change in a speech component between the first speech and the second speech; and

based on the change in the speech component, creating a transform to adjust the speech component.

2. The method of claim 1 , wherein the change in the speech component is associated with one or more of a change to a phoneme, in a speed of the phoneme, in a duration of the phoneme, to a separation between phonemes, to a separation between morphemes, in a pitch of the phoneme, in a frequency range for the phoneme, or in a pause between words.

3. The method of claim 2 , wherein the change in the speech component mimics the Lombard Effect.

4. The method of claim 2 , wherein a first phoneme is pronounced in a first frequency range and a second phoneme is pronounced in a second frequency range, and wherein the change to the speech component involves a first change to the first phoneme that is different than a second change to the second phoneme based on a difference between the first frequency range and the second frequency range.

5. The method of claim 1 , further comprising:

assigning a desired voice;

receiving a request from the user;

determining a current sound environment for the user;

determining text to output to the user in response to the request;

synthesizing the text, by Text-To-Speech (TTS), to create speech output;

applying the transform to the speech output; and

playing the transformed speech output.

6. The method of claim 5 , wherein the request from the user is a wake word, the method further comprising:

retrieving the transform; and

adjusting a reception of the wake word based on the transform.

7. The method of claim 5 , further comprising:

receiving third speech from the user in the request;

retrieving the transform; and

adjusting a speech recognition of the third speech based on the transform.

8. The method of claim 1 , wherein the transform is associated with the second simulated sound environment.

9. The method of claim 1 , wherein the first simulated sound environment is quieter than the first simulated sound environment.

10. The method of claim 9 , wherein the first simulated sound environment has a sound level of 45 dB or less, and wherein the second simulated sound environment has a sound level of over 50 dB.

11. The method of claim 1 , wherein the transform is applied to sounds in a partial band of frequencies.

12. A system comprising:

a memory; and

a processing unit coupled to the memory, wherein the processing unit is operative to:

provide, by a media delivery system, a first audio representation of a first simulated sound environment;

receive, by the media delivery system, first speech from a user speaking subject to audio playout of the first simulated sound environment;

provide, by the media delivery system, a second audio representation of a second simulated sound environment, wherein the second simulated sound environment has different acoustic characteristics than the first simulated sound environment;

receive, by the media delivery system, second speech from the user speaking subject to audio playout of the second simulated sound environment;

determine a change in a speech component between the first speech and the second speech; and

based on the change in the speech component, create a transform to adjust the speech component.

13. The system of claim 12 , wherein the change in the speech component is associated with one or more of a change to a phoneme, in a speed of the phoneme, in a duration of the phoneme, to a separation between phonemes, to a separation between morphemes, in a pitch of the phoneme, in a frequency range for the phoneme, and wherein the change in the speech component mimics the Lombard Effect.

14. The system of claim 12 , the processing unit further operative to:

assign a desired voice;

receive a request from the user;

determine a current sound environment for the user;

determine text to output to the user in response to the request;

synthesize the text, by Text-To-Speech (TTS), to create speech output;

apply the transform to the speech output; and

play the transformed speech output.

15. The system of claim 14 , wherein the request from the user is a wake word, the processing unit further operative to:

retrieve the transform; and

adjust a reception of the wake word based on the transform.

16. The system of claim 14 , the processing unit further operative to:

receive third speech from the user in the request;

retrieve the transform; and

adjust a speech recognition of the third speech based on the transform.

17. A method comprising:

determining, by a media-playback device, a sound environment from received background noise;

selecting a sound profile with similar audio characteristics as the received background noise, wherein the sound profile is associated with a transform for speech;

determining speech output;

applying the transform to the speech output to create transformed speech; and

playing, by the media-playback device, the transformed speech.

18. The method of claim 17 , wherein a characteristic associated with the transform comprises one or more of a change to a phoneme, in a speed of the phoneme, in a duration of the phoneme, to a separation between phonemes, to a separation between morphemes, in a pause between words, in a pitch of the phoneme, in a frequency range for the phoneme, and wherein the transform mimics the Lombard effect.

19. The method of claim 17 , wherein a user can understand the transformed speech in the sound environment without changing a volume of the media-playback device.

20. The method of claim 17 , wherein determining the sound environment comprises determining a dB (A) or a dB (C).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2022
From: BROMAND, DANIEL; ROTH, BJÖRN ERIK; SJÖLANDER, KÅRE
To: SPOTIFY AB
Reel/Frame 059921/0379 →
Continuity (2)
Provisional Application 63133101 · Dec 31, 2020
Related Publication 20220208174A1 · Jun 30, 2022
References Cited (32)
US 7653761B2 · Juster · 2010 [cited by applicant]
US 8094891B2 · Andreasson · 2012 [cited by applicant]
US 8571871B1 · Stuttle · 2013 [cited by examiner]
US 8589167B2 · Baughman · 2013 [cited by examiner]
US 10237256B1 · Pena · 2019 [cited by applicant]
US 10524070B2 · Kadri · 2019 [cited by applicant]
US 10877718B2 · Gosu · 2020 [cited by applicant]
US 10885900B2 · Li · 2021 [cited by examiner]
US 20060020662A1 · Robinson · 2006 [cited by applicant]
US 20060212478A1 · Plastina · 2006 [cited by applicant]
US 20070113725A1 · Oliver · 2007 [cited by applicant]
US 20070174866A1 · Brown · 2007 [cited by applicant]
US 20070276866A1 · Bodin · 2007 [cited by applicant]
US 20080317292A1 · Baker · 2008 [cited by applicant]
US 20090044687A1 · Sorber · 2009 [cited by applicant]
US 20090055426A1 · Kalasapur · 2009 [cited by applicant]
US 20090063414A1 · White · 2009 [cited by applicant]
US 20090164516A1 · Svendsen · 2009 [cited by applicant]
US 20090172538A1 · Bates · 2009 [cited by applicant]
US 20090222392A1 · Martin · 2009 [cited by applicant]
US 20090325602A1 · Higgins · 2009 [cited by applicant]
US 20090328087A1 · Higgins · 2009 [cited by applicant]
US 20110173539A1 · Rottler · 2011 [cited by applicant]
US 20110295843A1 · Ingrassia, Jr. · 2011 [cited by applicant]
US 20150154647A1 · Suwald · 2015 [cited by applicant]
US 20150281878A1 · Roundtree · 2015 [cited by applicant]
US 20170104824A1 · Bajwa · 2017 [cited by applicant]
US 20180091913A1 · Hartung · 2018 [cited by applicant]
US 20180158447A1 · Maziewski · 2018 [cited by examiner]
US 20190228791A1 · Sun · 2019 [cited by examiner]
US 20210020162A1 · Griffin · 2021 [cited by examiner]
US 20210097980A1 · Lezzoum · 2021 [cited by examiner]