IP Library Granted Patent US 12,273,679
Granted Patent B2
US 12,273,679 · App. 18/642,492 · Granted Apr 8, 2025

Multi-modal audio processing

Inventor: Karl Stahl (Palo Alto, CA)
H04R1/1083G10L15/22G10L21/0316G10L25/06G10L25/51H04R1/08G10L2015/223H04R5/0335H04R2420/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,273,679
App. No.
18/642,492
Granted
Apr 8, 2025
Kind
B2
Abstract

A method for processing an audio signal involves receiving sound waves at a microphone, converting them into a first audio signal, and extracting a second audio signal from an electromagnetic signal received at a receiver. The first audio signal is correlated with the second audio signal to calculate a correlation value. If the correlation value exceeds a threshold, the first audio signal is processed using the second audio signal to reduce unwanted sound contributions, resulting in a processed audio signal. Further processing is then performed on the processed audio signal to determine a characteristic of the desired sound.

Claims (69)

1. A method of processing an audio signal, the method comprising:

receiving a set of sound waves at a microphone of a device, the set of sound waves comprising a sound of interest and other sound;

converting, using the microphone, the set of sound waves into a first audio signal that includes a contribution from the sound of interest and a contribution from the other sound;

receiving, at a receiver of the device, an electromagnetic signal carrying a second audio signal;

extracting the second audio signal from the electromagnetic signal;

correlating the first audio signal with the second audio signal to calculate a correlation value;

in response to the correlation value being larger than a threshold, processing the first audio signal using the second audio signal to reduce the contribution from the other sound in a processed audio signal; and

performing further processing on the processed audio signal to determine a characteristic of the sound of interest.

2. The method of claim 1 , further comprising:

determining an amplitude of a version of the second audio signal that is present within the first audio signal;

scaling the second audio signal obtained from the electromagnetic signal based on the determined amplitude to generate a modified version of the second audio signal; and

subtracting the modified version of the second audio signal from the first audio signal as at least a part of said processing.

3. The method of claim 1 , further comprising:

determining a time difference between a version of the second audio signal that is present within the first audio signal and the second audio signal that is obtained from the electromagnetic signal;

delaying the second audio signal obtained from the electromagnetic signal using the determined time difference to generate a modified version of the second audio signal; and

subtracting the modified version of the second audio signal from the first audio signal as at least a part of said processing.

4. The method of claim 1 , wherein the electromagnetic signal comprises a modulated radio signal and the second audio signal is extracted by demodulating the modulated radio signal.

5. The method of claim 1 , further comprising:

generating, at a speaker device, the electromagnetic signal by modulating a radio signal using the second audio signal;

transmitting, from the speaker device to the device, the electromagnetic signal; and

producing, by the speaker device, at least some of the other sound using the second audio signal.

6. The method of claim 5 , further comprising:

receiving, at the speaker device, an electrical signal through one or more conductors, the electrical signal comprising the second audio signal; and

powering a transmitter for the electromagnetic signal in the speaker device using electrical power extracted from the second audio signal in the electrical signal.

7. The method of claim 1 , wherein the electromagnetic signal comprises a radio-frequency carrier modulated using the second audio signal.

8. The method of claim 1 , further comprising:

determining that both a first version of the second audio signal and a second version of the second audio signal are present within the first audio signal, wherein the first version of the second audio signal and the second version of the second audio signal each have at least one of a different amplitude than, a delay from, or a frequency shift from, the second audio signal and from each other; and

subtracting both the first version of the second audio signal and the second version of the second audio signal from the first audio signal as at least a part of said processing.

9. The method of claim 1 , further comprising:

receiving at least some of the other sound at a transducer of a second device remote from the device;

converting the received at least some of the other sound into the second audio signal;

generating the electromagnetic signal using the second audio signal; and

transmitting the electromagnetic signal from the second device for reception by the device.

10. The method of claim 1 , further comprising:

detecting one or more versions of the second audio signal within the first audio signal;

determining an acoustic transfer function that maps the second audio signal to the detected one or more versions of the second audio signal; and

using the determined acoustic transfer function to remove the one or more versions of the second audio signal from the first audio signal as at least a part of said processing.

11. A device comprising:

a microphone to receive a set of sound waves comprising sound of interest and other sound, and to output a first audio signal that includes a contribution from the sound of interest and a contribution from the other sound;

a receiver to receive an electromagnetic signal and to output a second audio signal obtained from the electromagnetic signal;

a correlator to generate one or more correlation parameters from a correlation of the first audio signal and the second audio signal; and

an audio pre-processor to process the first audio signal using the second audio signal and at least one of the one or more correlation parameters to reduce the contribution from the other sound in a processed audio signal;

wherein the device is configured to provide the processed audio signal to at least one processor to determine a characteristic of the sound of interest.

12. The device of claim 11 , wherein the one or more correlation parameters comprise a time delay and/or a scaling factor.

13. The device of claim 11 , wherein the device is configured to determine signal characteristics for a plurality of copies of the second audio signal that are present within the first audio signal and wherein the audio pre-processor is further configured to process the first audio signal based on the signal characteristics to generate the processed audio signal.

14. The device of claim 11 , wherein the electromagnetic signal is received from a remote device and the second audio signal represents sounds that are captured by a transducer at the remote device.

15. The device of claim 11 , wherein the audio pre-processor is further configured to:

detect one or more versions of the second audio signal within the first audio signal;

determine an acoustic transfer function that maps the second audio signal to the detected one or more versions of the second audio signal; and

use the determined acoustic transfer function to remove the one or more versions of the second audio signal from the first audio signal to generate the processed audio signal.

16. A non-transitory computer-readable storage medium storing instructions which, when executed by at least one processor, program the at least one processor to:

obtain a first audio signal that includes a contribution from sound of interest and a contribution from other sound, the first audio signal derived from a set of sound waves received at a microphone of a device, the set of sound waves comprising the sound of interest and the other sound;

obtain a second audio signal, the second audio signal derived from an electromagnetic signal received at a receiver of the device;

correlate the first audio signal and the second audio signal to generate one or more correlation parameters, the one or more correlation parameters indicating a time delay and/or a scaling factor for the second audio signal;

in response to the scaling factor being larger than a threshold, reduce the contribution from the other sound in the first audio signal using the second audio signal and the one or more correlation parameters to generate a processed audio signal; and

provide the processed audio signal to an audio processor to determine a characteristic of the sound of interest.

17. The storage medium of claim 16 , the at least one processor further programmed to:

obtain a plurality of other audio signals, including the second audio signal, from the electromagnetic signal, the electromagnetic signal comprising one or more electromagnetic signals;

detect one or more of the plurality of other audio signals within the first audio signal; and

process the first audio signal using the detected one or more of the plurality of other audio signals to reduce the contribution from the other sound in the processed audio signal.

18. The storage medium of claim 16 , the at least one processor further programmed to:

obtain a third audio signal from the electromagnetic signal, the electromagnetic signal comprising one or more electromagnetic signals;

correlate the first audio signal with the third audio signal to calculate a correlation value; and

in response to the correlation value being larger than a threshold, further reduce the contribution from the other sound in the first audio signal by using the third audio signal to generate the processed audio signal.

19. The storage medium of claim 16 , wherein the one or more correlation parameters indicate a plurality of time delays and/or scaling factors for the second audio signal due to a plurality of versions of the second audio signal being found in the first audio signal.

20. The storage medium of claim 16 , the at least one processor further programmed to:

determine identifying information for the received electromagnetic signal;

retrieve one or more previously stored characteristics based on the identifying information; and

use the retrieved characteristics with the second audio signal as at least a part of said processing.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: STAHL, KARL
To: SOUNDHOUND, INC.
Reel/Frame 067192/0513 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 067769/0712 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 067769/0768 →
Continuity (4)
Continuation 18194885 · Apr 3, 2023
Continuation 17301308 · Mar 31, 2021
Provisional Application 63004364 · Apr 2, 2020
Related Publication 20240276138A1 · Aug 15, 2024
References Cited (18)
US 10325591B1 · Pogue et al. · 2019 [cited by applicant]
US 10650840B1 · Solbach · 2020 [cited by applicant]
US 11627405B2 · Stahl · 2023 [cited by applicant]
US 20020153778A1 · Oughton · 2002 [cited by applicant]
US 20030118197A1 · Nagayasu et al. · 2003 [cited by applicant]
US 20130058496A1 · Harris · 2013 [cited by applicant]
US 20130077803A1 · Konno et al. · 2013 [cited by applicant]
US 20140270194A1 · Des Jardins · 2014 [cited by examiner]
US 20170034026A1 · Li · 2017 [cited by examiner]
US 20180188351A1 · Jones et al. · 2018 [cited by applicant]
US 20190387312A1 · Liu et al. · 2019 [cited by applicant]
US 20190394567A1 · Moore · 2019 [cited by examiner]
US 20190394603A1 · Moore · 2019 [cited by examiner]
US 20210312920A1 · Stahl · 2021 [cited by applicant]
US 20210375275A1 · Yoon et al. · 2021 [cited by applicant]
Crepaldi, Marco, et al. “An analog-mode impulse radio system for ultra-low power short-range audio streaming.” IEEE Transactions on Circuits and Systems (2015) pp. 2886-2897 (Year: 2015). [cited by applicant]
Flores, Marcel. “Wi-FM: Resolving neighborhood wireless network affairs by listening to music.” 2015 IEEE 23rd International Conference on Network Protocols (ICNP). (2015) pp. 43-53 (Year: 2015). [cited by applicant]
Jeong, Seokhyeon, et al. “Always-on 12-nW acoustic sensing and object recognition microsystem for unattended ground sensor nodes.” IEEE Journal of Solid-State Circuits (2017) pp. 261-274 (Year: 2017). [cited by applicant]