IP Library Granted Patent US 12,334,094
Granted Patent B2
US 12,334,094 · App. 18/056,611 · Granted Jun 17, 2025

Audio cancellation for voice recognition

Inventors: Richard Mitic (Stockholm, SE); Robert Swain (Stockholm, SE); Daniel Bromand (Stockholm, SE); Waqar Sheikh (Malmo, SE); James Robert Stansfield (Flyinge, SE)
Assignee: Spotify AB
G10L21/0232G10L25/51H04R3/00G10L15/20G10L15/22G10L2015/223H04R2420/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,094
App. No.
18/056,611
Granted
Jun 17, 2025
Kind
B2
Abstract

An audio cancellation system includes a voice enabled computing system that is connected to an audio output device using a wired or wireless communication network. The voice enabled computing device can provide media content to a user and receive a voice command from the user. The connection between the voice enabled computing system and the audio output device introduces a time delay between the media content being generated at the voice enabled computing device and the media content being reproduced at the audio output device. The system operates to determine a calibration value adapted for the voice enabled computing system and the audio output device. The system uses the calibration value to filter the user's voice command from a recording of ambient sound including the media content, without requiring significant use of memory and computing resources.

Claims (56)

1. A media delivery system comprising:

a processor;

a memory storing data instructions that, when executed by the processor, cause the media delivery system to:

receive, from a sound system, a time delay between a first time associated with a generation of an audio cue at a media playback device of the sound system and a second time associated with when the audio cue is recorded in a recording in an audio buffer of the media playback device, wherein the audio cue represents a first signal sent at the first time and a second signal sent at a third time, wherein the first time and the third time are different, wherein a calibration value is used by the sound system in an audio cancellation operation to reduce background noise in the recording, wherein the calibration value is based on the time delay, and wherein the media playback device determines the time delay by:

determining a first time delay associated with the first signal;

determining a second time delay associated with the second signal; and

averaging the first time delay and the second time delay associated with the first and second signals to determine the time delay;

analyze, at the media delivery system, performance of the audio cancellation operation when operated using the calibration value; and

send instructions to the sound system to adjust the calibration value.

2. The media delivery system of claim 1 , wherein the audio cue has a first root mean square (RMS) higher than a second RMS associated with the background noise.

3. The media delivery system of claim 1 , wherein the audio cue has a strong attack of less than 100 milliseconds.

4. The media delivery system of claim 1 , wherein the audio cue comprises two or more frequencies.

5. The media delivery system of claim 1 , wherein the audio cue is an emulated sound from a snare drum.

6. The media delivery system of claim 1 , wherein the background noise is a person talking.

7. The media delivery system of claim 1 , wherein the background noise is associated with an operation of a motor vehicle or is noise from a room or building where the sound system is located.

8. The media delivery system of claim 1 , wherein the background noise emanates from an engine, a home appliance, a television, an animal, wind noise, or traffic.

9. The media delivery system of claim 1 , wherein the audio cue comprises a non-verbal response, and wherein the non-verbal response comprises a beep, a signal, or a ding.

10. The media delivery system of claim 1 , wherein the audio cue comprises a verbal response, and wherein the verbal response comprises a word, a phrase, or a short sentence.

11. The media delivery system of claim 1 , wherein the data instructions when executed by the processor, further cause the media delivery system to:

receive a request from the media playback device for one or more media content items; and

transmit the one or more media content items to the media playback device for the media playback device to send the one or more media content items to an audio output device for playback, wherein the media playback device and the audio output device together comprise the sound system.

12. The media delivery system of claim 1 , wherein the second time at which the audio cue is detected in the recording occurs when a peak-to-RMS ratio crosses a predetermined threshold.

13. The media delivery system of claim 12 , wherein the predetermined threshold is 30 decibels.

14. A media playback system comprising:

a media delivery system; and

a sound system comprising:

an audio output device; and

a media playback device configured to:

generate an audio cue at a first time, wherein the audio cue represents a first signal sent at the first time and a second signal sent at a third time, and wherein the first time and the third time are different;

send the audio cue to the audio output device;

record a recording in an audio buffer using a microphone, the recording including the audio cue recorded at a second time;

detect the audio cue in the recording from the audio buffer over a background noise in a sound environment of the sound system;

determine a time delay between the first time of the generation of the audio cue and the second time that the audio cue was recorded in the recording in the audio buffer, wherein the time delay is determined by:

determining a first time delay associated with the first signal;

determining a second time delay associated with the second signal; and

averaging the first time delay and the second time delay associated with the first and second signals to determine the time delay;

send the time delay to the media delivery system for analysis;

receive a calibration value from the media delivery system, wherein the calibration value is based on the time delay; and

use the calibration value to cancel audio from the sound system from subsequent recordings.

15. The media playback system of claim 14 , wherein the audio output device is configured to:

receive the audio cue from the media playback device; and play the audio cue.

16. The media playback system of claim 14 , wherein the audio cue has a first root mean square (RMS) higher than a second RMS associated with the background noise.

17. The media playback system of claim 14 , wherein the audio cue has a strong attack of less than 100 milliseconds.

18. The media playback system of claim 14 , wherein the audio cue comprises two or more frequencies.

19. The media playback system of claim 14 , wherein the audio cue is an emulated sound from a snare drum.

20. A method of audio cancellation comprising:

generating, at a media playback device, an audio cue at a first time, wherein the audio cue represents a first signal sent at the first time and a second signal sent at a third time, and wherein the first time and the third time are different;

sending, from the media playback device, the audio cue to an audio output device;

recording, at the media playback device, a recording in an audio buffer using a microphone, the recording including the audio cue recorded at a second time;

determining a time delay between the first time of the generation of the audio cue and the second time that the audio cue was recorded in the recording in the audio buffer, wherein the time delay is determined by:

determining a first time delay associated with the first signal;

determining a second time delay associated with the second signal; and

averaging the first time delay and the second time delay associated with the first and second signals to determine the time delay;

sending the time delay to a media delivery system for analysis;

receiving, from the media delivery system, a calibration value, wherein the calibration value is based on the time delay; and

using the calibration value, cancelling audio from subsequent recordings played by a sound system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2024
From: MITIC, RICHARD; SWAIN, ROBERT; BROMAND, DANIEL; SHEIKH, WAQAR; STANSFIELD, JAMES ROBERT
To: SPOTIFY AB
Reel/Frame 068834/0439 →
Priority Claims (1)
EP 18202941 · Oct 26, 2018 · regional
Continuity (4)
Continuation 17158312 · Jan 26, 2021
Continuation 16593539 · Oct 4, 2019
Provisional Application 62820762 · Mar 19, 2019
Related Publication 20230162752A1 · May 25, 2023
References Cited (27)
US 7881460B2 · Looney et al. · 2011 [cited by applicant]
US 8503669B2 · Mao · 2013 [cited by applicant]
US 9430999B2 · Clemow · 2016 [cited by applicant]
US 9484030B1 · Meaney et al. · 2016 [cited by applicant]
US 9633671B2 · Giacobello et al. · 2017 [cited by applicant]
US 10943599B2 · Mitic · 2021 [cited by applicant]
US 20080082326A1 · Venkataraman et al. · 2008 [cited by applicant]
US 20140126745A1 · Dickins · 2014 [cited by examiner]
US 20150271616A1 · Kechichian · 2015 [cited by applicant]
US 20150371654A1 · Johnston et al. · 2015 [cited by applicant]
US 20160171988A1 · Vos · 2016 [cited by examiner]
US 20160275050A1 · Tanaka · 2016 [cited by examiner]
US 20170245079A1 · Sheen et al. · 2017 [cited by applicant]
US 20180225082A1 · An et al. · 2018 [cited by applicant]
US 20180306890A1 · Vatcher · 2018 [cited by examiner]
US 20180314689A1 · Wang · 2018 [cited by examiner]
US 20180351523A1 · Lesso · 2018 [cited by examiner]
US 20190028803A1 · Benattar · 2019 [cited by examiner]
US 20190318069A1 · Mitic · 2019 [cited by examiner]
JP 04355549A · 1992 [cited by applicant]
JP 11234176A · 1999 [cited by applicant]
WO 2017039575A1 · 2017 [cited by applicant]
Malcolm Owen “A deep dive into HomePod's adaptive audio, beamforming and why it needs an A8 processor”, Apple Insider, 14 pages (Jan. 26, 2018). Available Online at: https://appleinsider.com/articles/18/01/27/a-deep-div… [cited by applicant]
Extended European Search Report from corresponding European Appl'n No. 19 205 155.5, mailed Dec. 13, 2019. [cited by applicant]
Communication pursuant to Article 94(3) EPC from corresponding European Appl'n No. 19 205 155.5, mailed Apr. 24, 2020. [cited by applicant]
European Summons to Oral Proceedings in Application 19205155.5, mailed Jul. 23, 2020, 7 pages. [cited by applicant]
European Result of Consultation in Application 19205155.5, mailed Jan. 26, 2021, 5 pages. [cited by applicant]