IP Library Granted Patent US 11,605,393
Granted Patent B2
US 11,605,393 · App. 17/158,312 · Granted Mar 14, 2023

Audio cancellation for voice recognition

Inventors: Richard Mitic (Stockholm, SE); Robert Swain (Stockholm, SE); Daniel Bromand (Stockholm, SE); Waqar Sheikh (Malmo, SE); James Robert Stansfield (Flyinge, SE)
Assignee: Spotify AB
G10L21/0232G10L25/51H04R3/00G10L15/20G10L15/22G10L2015/223H04R2420/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,605,393
App. No.
17/158,312
Granted
Mar 14, 2023
Kind
B2
Abstract

An audio cancellation system includes a voice enabled computing system that is connected to an audio output device using a wired or wireless communication network. The voice enabled computing device can provide media content to a user and receive a voice command from the user. The connection between the voice enabled computing system and the audio output device introduces a time delay between the media content being generated at the voice enabled computing device and the media content being reproduced at the audio output device. The system operates to determine a calibration value adapted for the voice enabled computing system and the audio output device. The system uses the calibration value to filter the user's voice command from a recording of ambient sound including the media content, without requiring significant use of memory and computing resources.

Claims (49)

1. A sound system comprising:

a media playback device configured to:

machine-generate an audio cue at a first time;

send an audio cue at a first time to an audio output device;

record sound in an audio buffer using a microphone, the recording including the audio cue recorded at a second time;

detect the audio cue in the recording from the audio buffer over the background noise in the sound environment;

determine a time delay between the first time of the generation of the audio cue and the second time that the audio cue was recorded in the recording in the audio buffer; and

use the time delay to cancel audio from the sound system from subsequent recordings; and

the audio output device configured to:

play media content using a media content signal;

receive the audio cue from the media playback device; and

play the audio cue.

2. The sound system of claim 1 , wherein the audio cue has a first root mean square (RMS) higher than a second RMS associated with the background noise.

3. The sound system of claim 1 , wherein the audio cue has a strong attack of less than 100 milliseconds.

4. The sound system of claim 1 , wherein the audio cue comprises two or more frequencies.

5. The sound system of claim 1 , wherein the audio cue is an emulated sound from a snare drum.

6. The sound system of claim 1 , wherein the background noise is a person talking.

7. The sound system of claim 1 , wherein the background noise is associated with an operation of a motor vehicle or is noise from a room or building where the sound system is located.

8. The sound system of claim 1 , wherein the background noise emanates from an engine, a home appliance, a television, an animal, wind noise, or traffic.

9. The sound system of claim 1 , wherein the audio cue represents a first signal sent at the first time and a second signal sent at a third time, and wherein the first time and the third time are different.

10. The sound system of claim 9 , wherein the media playback device is further configured to:

determine a first time delay associated with the first signal;

determine a second time delay associated with the second signal; and

average the first time delay and the second time delay associated with the first and second signals to determine the time delay.

11. The sound system of claim 1 , wherein the second time at which the audio cue is detected in the recording occurs when a peak-to-RMS ratio crosses a predetermined threshold.

12. The sound system of claim 11 , wherein the predetermined threshold is 30 decibels.

13. A media playback device comprising:

a processor;

a memory storing data instructions that, when executed by the processor, cause the media playback device to:

send an audio cue at a first time;

record sound in an audio buffer using a microphone, the recording including the audio cue recorded at a second time;

detect the audio cue in the recording from the audio buffer over the background noise in the sound environment;

determine a time delay between the first time of the generation of the audio cue and the second time that the audio cue was recorded in the recording in the audio buffer; and

use the time delay to cancel audio from the sound system from subsequent recordings.

14. The media playback device of claim 13 , wherein the audio cue has a first root mean square (RMS) higher than a second RMS associated with the background noise.

15. The media playback device of claim 14 , wherein the audio cue has a strong attack of less than 100 milliseconds.

16. The media playback device of claim 13 , wherein the audio cue comprises two or more frequencies.

17. The media playback device of claim 13 , wherein the audio cue represents a first signal sent at the first time and a second signal sent at a third time, and wherein the first time and the third time are different, and wherein the method further comprises:

determining a first time delay associated with the first signal;

determining a second time delay associated with the second signal; and

averaging the first time delay and the second time delay associated with the first and second signals to determine the time delay.

18. A non-transitory computer readable medium having stored thereon instructions, which when executed by a processor of a computing device, cause the computing device to:

send an audio cue at a first time;

record sound in an audio buffer using a microphone, the recording including the audio cue recorded at a second time;

detect the audio cue in the recording from the audio buffer over the background noise in the sound environment;

determine a time delay between the first time of the generation of the audio cue and the second time that the audio cue was recorded in the recording in the audio buffer; and

use the time delay to cancel audio from the sound system from subsequent recordings.

19. The non-transitory computer readable medium of claim 18 , wherein the audio cue has a first root mean square (RMS) higher than a second RMS associated with the background noise, and wherein the audio cue has a strong attack of less than 100 milliseconds.

20. The non-transitory computer readable medium of claim 18 , wherein the audio cue comprises two or more frequencies.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2021
From: MITIC, RICHARD; SWAIN, ROBERT; BROMAND, DANIEL; SHEIKH, WAQAR; STANSFIELD, JAMES ROBERT
To: SPOTIFY AB
Reel/Frame 055870/0523 →
Priority Claims (1)
EP 18202941 · Oct 26, 2018 · regional
Continuity (3)
Continuation 16593539 · Oct 4, 2019
Provisional Application 62820762 · Mar 19, 2019
Related Publication 20210287693A1 · Sep 16, 2021
Cited By (1)
US 12,254,876