IP Library Granted Patent US 12,499,901
Granted Patent B2
US 12,499,901 · App. 18/693,411 · Granted Dec 16, 2025

Noise reduction using synthetic audio

Inventors: James Nesfield (Edinburgh, GB); Ian Ward Frank (Arlington, MA)
Assignee: Sonos, Inc.
G10L21/0216G10L25/30G10L25/78H04R3/005G10L2021/02165
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,901
App. No.
18/693,411
Granted
Dec 16, 2025
Kind
B2
Abstract

Systems and methods for noise reduction using synthetic audio are disclosed. One or more playback devices include first microphone(s) (e.g., air-conduction microphone(s)) and second microphone(s) (e.g., bone-conduction microphone(s)). In operation, based on a user voice input, first and sound data streams are captured via the first and second microphone(s), respectively. The first sound data stream is evaluated to determine whether a noise threshold is exceeded. While the noise threshold is not exceeded, a synthetic sound data model is trained based on the first and second sound data streams. An audio output stream based on the first sound data stream is communicated to at least one second playback device. When noise is detected in the first sound data stream, a synthetic audio stream is mixed into the output audio stream. The synthetic audio stream can be produced based on the synthetic sound data model.

Claims (51)

1 . One or more playback devices comprising:

a first one or more microphones;

a second one or more microphones;

one or more processors; and

data storage having instructions stored thereon that, when executed by the one or more processors, cause the one or more playback devices to perform operations comprising:

capturing a first sound data stream based on a user voice input via the first one or more microphones;

while capturing the first sound data stream, (i) concurrently capturing a second sound data stream based on the user voice input via the second one or more microphones and (ii) evaluating the first sound data stream to determine whether a noise threshold is exceeded;

while the noise threshold is not exceeded, training a synthetic sound data model based on the first sound data stream and the second sound data stream concurrently captured with first sound data;

communicating an output audio stream to at least one second playback device based on the first sound data stream;

detecting noise in the first sound data stream; and

mixing a synthetic audio stream into the output audio stream based on the noise in the first sound data stream, wherein the synthetic audio stream is produced based on the synthetic sound data model.

2 . The one or more playback devices of claim 1 , wherein the operations further comprise:

detecting that the noise is no longer present in the first sound data stream; and

ceasing the mixing of the synthetic audio stream into the output audio stream after detecting that the noise is no longer present in the first sound data stream.

3 . The one or more playback devices of claim 1 , wherein the first one or more microphones comprises an air-conduction microphone and the second one or more microphones comprises a bone-conduction microphone.

4 . The one or more playback devices of claim 1 , wherein the one or more playback devices comprises a portable playback device comprising the first and second microphones.

5 . The one or more playback devices of claim 1 , wherein the mixing is based at least in part on a determined noise level in the second sound data stream.

6 . The one or more playback devices of claim 1 , wherein the mixing comprises utilizing the synthetic sound data model over a first frequency range and utilizing the second sound data stream over a second frequency range different from the first.

7 . The one or more playback devices of claim 1 , wherein the mixing comprises combining the first sound data stream at a first relative volume level with the synthetic audio stream at a second relative volume level.

8 . The one or more playback devices of claim 1 , wherein the mixing comprises dynamically switching between the first sound data stream and the synthetic audio stream.

9 . A method comprising:

capturing a first sound data stream based on a user voice input via a first one or more microphones of at least one first playback device;

while capturing the first sound data stream, (i) concurrently capturing a second sound data stream based on the user voice input via a second one or more microphones of the at least one first playback device and (ii) evaluating the first sound data stream to determine whether a noise threshold is exceeded;

while the noise threshold is not exceeded, training a synthetic sound data model based on the first sound data stream and the second sound data stream concurrently captured with first sound data;

communicating an output audio stream to at least one second playback device based on the first sound data stream;

detecting noise in the first sound data stream; and

mixing a synthetic audio stream into the output audio stream based on the noise in the first sound data stream, wherein the synthetic audio stream is produced based on the synthetic sound data model.

10 . The method of claim 9 , further comprising:

detecting that the noise is no longer present in the first sound data stream; and

ceasing the mixing of the synthetic audio stream into the output audio stream after detecting that the noise is no longer present in the first sound data stream.

11 . The method of claim 9 , wherein the first one or more microphones comprises an air-conduction microphone and the second one or more microphones comprises a bone-conduction microphone.

12 . The method of claim 9 , wherein the at least one first playback device comprises a portable playback device comprising the first and second microphones.

13 . The method of claim 9 , wherein the mixing is based at least in part on a determined noise level in the second sound data stream.

14 . The method of claim 9 , wherein the mixing comprises utilizing the synthetic sound data model over a first frequency range and utilizing the second sound data stream over a second frequency range different from the first.

15 . The method of claim 9 , wherein the mixing comprises combining the first sound data stream at a first relative volume level with the synthetic audio stream at a second relative volume level.

16 . The method of claim 9 , wherein the mixing comprises dynamically switching between the first sound data stream and the synthetic audio stream.

17 . One or more tangible, non-transitory computer-readable media comprising instructions that, when executed by one or more processors of at least one playback device, cause the at least one playback device to perform operations comprising:

capturing a first sound data stream based on a user voice input via a first one or more microphones of the at least one playback device;

while capturing the first sound data stream, (i) concurrently capturing a second sound data stream based on the user voice input via a second one or more microphones of the at least one playback device and (ii) evaluating the first sound data stream to determine whether a noise threshold is exceeded;

while the noise threshold is not exceeded, training a synthetic sound data model based on the first sound data stream and the second sound data stream concurrently captured with first sound data;

communicating an output audio stream to at least one second playback device based on the first sound data stream;

detecting noise in the first sound data stream; and

mixing a synthetic audio stream into the output audio stream based on the noise in the first sound data stream, wherein the synthetic audio stream is produced based on the synthetic sound data model.

18 . The computer-readable media of claim 17 , wherein the operations further comprise:

detecting that the noise is no longer pres 3 nt in the first sound data stream; and

ceasing the mixing of the synthetic audio stream into the output audio stream after detecting that the noise is no longer present in the first sound data stream.

19 . The computer-readable media of claim 17 , wherein the mixing is based at least in part on a determined noise level in the second sound data stream.

20 . The computer-readable media of claim 17 , wherein the mixing comprises at least one of:

utilizing the synthetic sound data model over a first frequency range and utilizing the second sound data stream over a second frequency range different from the first;

combining the first sound data stream at a first relative volume level with the synthetic audio stream at a second relative volume level; or

dynamically switching between the synthetic audio stream and the second sound data stream.

Assignments (2)
SECURITY INTEREST Recorded Jan 30, 2026
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074533/0615 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2024
From: NESFIELD, JAMES; FRANK, IAN WARD
To: SONOS, INC.
Reel/Frame 066827/0944 →
Continuity (2)
Provisional Application 63261885 · Sep 30, 2021
Related Publication 20240249740A1 · Jul 25, 2024
References Cited (67)
US 5440644A · Farinelli et al. · 1995 [cited by applicant]
US 5761320A · Farinelli et al. · 1998 [cited by applicant]
US 5923902A · Inagaki · 1999 [cited by applicant]
US 5933506A · Aoki · 1999 [cited by examiner]
US 6032202A · Lea et al. · 2000 [cited by applicant]
US 6256554B1 · DiLorenzo · 2001 [cited by applicant]
US 6404811B1 · Cvetko et al. · 2002 [cited by applicant]
US 6469633B1 · Wachter · 2002 [cited by applicant]
US 6522886B1 · Youngs et al. · 2003 [cited by applicant]
US 6611537B1 · Edens et al. · 2003 [cited by applicant]
US 6631410B1 · Kowalski et al. · 2003 [cited by applicant]
US 6757517B2 · Chang · 2004 [cited by applicant]
US 6778869B2 · Champion · 2004 [cited by applicant]
US 7130608B2 · Hollstrom et al. · 2006 [cited by applicant]
US 7130616B2 · Janik · 2006 [cited by applicant]
US 7143939B2 · Henzerling · 2006 [cited by applicant]
US 7236773B2 · Thomas · 2007 [cited by applicant]
US 7295548B2 · Blank et al. · 2007 [cited by applicant]
US 7391791B2 · Balassanian et al. · 2008 [cited by applicant]
US 7483538B2 · McCarty et al. · 2009 [cited by applicant]
US 7571014B1 · Lambourne et al. · 2009 [cited by applicant]
US 7630501B2 · Blank et al. · 2009 [cited by applicant]
US 7643894B2 · Braithwaite et al. · 2010 [cited by applicant]
US 7657910B1 · McAulay et al. · 2010 [cited by applicant]
US 7853341B2 · McCarty et al. · 2010 [cited by applicant]
US 7987294B2 · Bryce et al. · 2011 [cited by applicant]
US 8014423B2 · Thaler et al. · 2011 [cited by applicant]
US 8045952B2 · Qureshey et al. · 2011 [cited by applicant]
US 8103009B2 · McCarty et al. · 2012 [cited by applicant]
US 8234395B2 · Millington · 2012 [cited by applicant]
US 8483853B1 · Lambourne · 2013 [cited by applicant]
US 8942252B2 · Balassanian et al. · 2015 [cited by applicant]
US 20010042107A1 · Palm · 2001 [cited by applicant]
US 20020022453A1 · Balog et al. · 2002 [cited by applicant]
US 20020026442A1 · Lipscomb et al. · 2002 [cited by applicant]
US 20020124097A1 · Isely et al. · 2002 [cited by applicant]
US 20030157951A1 · Hasty, Jr. · 2003 [cited by applicant]
US 20040024478A1 · Hans et al. · 2004 [cited by applicant]
US 20070142944A1 · Goldberg et al. · 2007 [cited by applicant]
US 20220392475A1 · Yan · 2022 [cited by examiner]
EP 1389853A1 · 2004 [cited by applicant]
EP 3737115A1 · 2020 [cited by applicant]
JP 2000261534A · 2000 [cited by examiner]
WO 200153994 · 2001 [cited by applicant]
WO 2003093950A2 · 2003 [cited by applicant]
WO 2021068120A1 · 2021 [cited by applicant]
AudioTron Quick Start Guide, Version 1.0, Mar. 2001, 24 pages. [cited by applicant]
AudioTron Reference Manual, Version 3.0, May 2002, 70 pages. [cited by applicant]
AudioTron Setup Guide, Version 3.0, May 2002, 38 pages. [cited by applicant]
Bluetooth. “Specification of the Bluetooth System: The ad hoc SCATTERNET for affordable and highly functional wireless connectivity,” Core, Version 1.0 A, Jul. 26, 1999, 1068 pages. [cited by applicant]
Bluetooth. “Specification of the Bluetooth System: Wireless connections made easy,” Core, Version 1.0 B, Dec. 1, 1999, 1076 pages. [cited by applicant]
Dell, Inc. “Dell Digital Audio Receiver: Reference Guide,” Jun. 2000, 70 pages. [cited by applicant]
Dell, Inc. “Start Here,” Jun. 2000, 2 pages. [cited by applicant]
“Denon 2003-2004 Product Catalog,” Denon, 2003-2004, 44 pages. [cited by applicant]
International Bureau, International Search Report and Written Opinion mailed on Dec. 21, 2022, issued in connection with International Application No. PCT/US2022/077154, filed on Sep. 28, 2022, 14 pages. [cited by applicant]
Jo et al., “Synchronized One-to-many Media Streaming with Adaptive Playout Control,” Proceedings of SPIE, 2002, pp. 71-82, vol. 4861. [cited by applicant]
Jones, Stephen, “Dell Digital Audio Receiver: Digital upgrade for your analog stereo,” Analog Stereo, Jun. 24, 2000 http://www.reviewsonline.com/articles/961906864.htm retrieved Jun. 18, 2014, 2 pages. [cited by applicant]
Louderback, Jim, “Affordable Audio Receiver Furnishes Homes With MP3,” TechTV Vault. Jun. 28, 2000 retrieved Jul. 10, 2014, 2 pages. [cited by applicant]
Palm, Inc., “Handbook for the Palm VII Handheld,” May 2000, 311 pages. [cited by applicant]
Presentations at WinHEC 2000, May 2000, 138 pages. [cited by applicant]
[cited by applicant]
U.S. Appl. No. 60/490,768, filed Jul. 28, 2003, entitled “Method for synchronizing audio playback between multiple networked devices,” 13 pages. [cited by applicant]
U.S. Appl. No. 60/825,407, filed Sep. 12, 2006, entitled “Controlling and manipulating groupings in a multi-zone music or media system,” 82 pages. [cited by applicant]
UPnP; “Universal Plug and Play Device Architecture,” Jun. 8, 2000; version 1.0; Microsoft Corporation; pp. 1-54. [cited by applicant]
Yamaha DME 64 Owner's Manual; copyright 2004, 80 pages. [cited by applicant]
Yamaha DME Designer 3.5 setup manual guide; copyright 2004, 16 pages. [cited by applicant]
Yamaha DME Designer 3.5 User Manual; Copyright 2004, 507 pages. [cited by applicant]