IP Library Granted Patent US 10,943,599
Granted Patent B2
US 10,943,599 · App. 16/593,539 · Granted Mar 9, 2021

Audio cancellation for voice recognition

Inventors: Richard Mitic (Stockholm, SE); Robert Swain (Stockholm, SE); Daniel Bromand (Stockholm, SE); Wagar Sheikh (Malmo, SE); James Robert Stansfield (Flyinge, SE)
Assignee: Spotify AB
G10L21/0232G10L25/51H04R3/00G10L15/20G10L15/22G10L2015/223H04R2420/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,943,599
App. No.
16/593,539
Granted
Mar 9, 2021
Kind
B2
Abstract

An audio cancellation system includes a voice enabled computing system that is connected to an audio output device using a wired or wireless communication network. The voice enabled computing device can provide media content to a user and receive a voice command from the user. The connection between the voice enabled computing system and the audio output device introduces a time delay between the media content being generated at the voice enabled computing device and the media content being reproduced at the audio output device. The system operates to determine a calibration value adapted for the voice enabled computing system and the audio output device. The system uses the calibration value to filter the user's voice command from a recording of ambient sound including the media content, without requiring significant use of memory and computing resources.

Claims (57)

1. A method of audio cancellation comprising:

machine-generating an audio cue at a first time;

playing the audio cue through a sound system in a sound environment, wherein the audio cue is detectable over background noise in the sound environment;

recording sound in an audio buffer using a microphone, the recording including the audio cue recorded at a second time;

detecting the audio cue in the recording from the audio buffer over the background noise in the sound environment;

determining a time delay between the first time of the generation of the audio cue and the second time that the audio cue was recorded in the recording in the audio buffer; and using the time delay to cancel audio from the sound system from subsequent recordings.

2. The method of claim 1 , wherein the audio cue has a first root mean square (RMS) higher than a second RMS associated with the background noise.

3. The method of claim 1 , wherein the audio cue has a strong attack of less than 100 milliseconds.

4. The method of claim 2 , wherein the audio cue comprises two or more frequencies.

5. The method of claim 1 , wherein the audio cue emanates from a snare drum.

6. The method of claim 2 , wherein the background noise is a person talking.

7. The method of claim 2 , wherein the background noise is associated with an operation of a motor vehicle or a room or building where the sound system is located.

8. The method of claim 2 , wherein the background noise emanates from an engine, a home appliance, a television, animal noises, wind noise, or traffic.

9. The method of claim 1 , wherein the audio cue comprises a plurality of signals, each signal played at a different time.

10. The method of claim 9 , wherein the time that the audio cue is detected in the recording occurs is when a peak-to-RMS ratio crosses a predetermined threshold.

11. The method of claim 10 , wherein the predetermined threshold is 30 dB.

12. The method of claim 11 , wherein the audio cue represents two or more signals, and wherein the method further comprises: averaging the time difference associated with the two or more signals.

13. A media playback system comprising:

a sound system including a media playback device and an audio output device, the media playback device operable to generate a media content signal, and the audio output device configured to play media content using the media content signal; and

wherein the sound system is configured to:

generate an audio cue using the media playback device at a first time;

transmit the audio cue to the audio output device;

play the audio cue through the audio output device;

record sound in an audio buffer using the media playback device, the recording including the audio cue recorded at a second time;

detect the audio cue in the recording from the audio buffer by determining that a peak-to-RMS ratio of the audio cue reaches or crosses a threshold;

determine a time delay between the first time of the generation of the audio cue and the second time that the audio cue was recorded in the recording in the audio buffer; and

play audio content through the audio output device;

receive a user command;

record the user command with the audio content;

determine which audio content has been recorded based on the time delay; and

filter the determined audio content to extract the user command from the recording of the user command with the audio content.

14. The media playback system of claim 13 , wherein the media playback device is paired with the audio output device via a wireless communication network, such as BLUETOOTH®.

15. The media playback system of claim 13 , wherein the sound system is configured to:

generate a second audio cue and playing the second audio cue through the sound system;

record sound in the audio buffer using the microphone, the recording including the second audio cue;

detect the second audio cue in the recording from the audio buffer by determining that a second peak-to-RMS ratio of the second audio cue crosses the threshold;

determine a second time delay between the generation of the second audio cue and a time that the second audio cue was recorded in the recording in the audio buffer;

determine a difference between the time delay and the second time delay;

determine whether the difference is within a threshold range;

when determined that the difference is within the threshold range, continue to use the time delay to cancel audio from the sound system from subsequent recordings; and

when determined that the difference is not within the threshold range, use the second time delay to cancel audio from subsequent recordings.

16. A method comprising:

generate an audio cue using a media playback device at a first time;

transmit the audio cue to an audio output device;

play the audio cue through the audio output device;

record sound in an audio buffer using the media playback device, the recording including the audio cue recorded at a second time;

detect the audio cue in the recording from the audio buffer by determining that a difference between a peak amplitude of the audio cue and a RMS of background noise, recorded with the audio cue, crosses a threshold;

determine a time delay between the first time of the generation of the audio cue and the second time that the audio cue was recorded in the recording in the audio buffer;

play audio content through the audio output device;

receive a user command;

record the user command with the audio content being played by the media playback device when the user command was received;

determine which audio content has been recorded based on the time delay; and

filter the determined audio content to extract the user command from the recording of the user command with the audio content.

17. The method of claim 16 , wherein the audio cue has a strong attack of less than 100 milliseconds.

18. The method of claim 16 , wherein the audio cue comprises two or more frequencies.

19. The method of claim 16 , wherein the audio cue mimics a sound that emanates from a snare drum.

20. The method of claim 16 , wherein the background noise is one or more of: a person talking, associated with an operation of a motor vehicle where the media playback device is located, associated with a room where the media playback device is located, associated with a building where the media playback device is located, emanates from an engine, emanates from a home appliance, emanates from a television, emanates from animal noises, emanates from wind noise, and/or emanates from traffic.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2020
From: SWAIN, ROBERT; SHEIKH, WAQAR; STANSFIELD, JAMES ROBERT
To: SPOTIFY AB
Reel/Frame 053743/0986 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2020
From: BROMAND, DANIEL; MITIC, RICHARD
To: SPOTIFY AB
Reel/Frame 051936/0391 →
Priority Claims (1)
EP 18202941 · Oct 26, 2018 · regional
Continuity (2)
Provisional Application 62820762 · Mar 19, 2019
Related Publication 20200135224A1 · Apr 30, 2020
Cited By (2)
US 12,254,876 US 12,334,094