IP Library › Granted Patent US 12,207,074
Granted Patent B2
US 12,207,074 · App. 18/532,988 · Granted Jan 21, 2025

Method and system for detecting sound event liveness using a microphone array

Inventors: Hassan Taherian (Columbus, OH); Jonathan Huang (Pleasanton, CA); Carlos M. Avendano (Campbell, CA)
Assignee: Apple Inc.
H04S7/302H04R3/005H04S2400/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,207,074
App. No.
18/532,988
Granted
Jan 21, 2025
Kind
B2
Abstract

A method performed by an electronic device in a room. The method performs an enrollment process in which a spatial profile of a location of an artificial sound source is created and performs an identification process that determines whether a sound event within the room is produced by the artificial sound source by 1) capturing the sound event using a microphone array and 2) determining a likelihood that the sound event occurred at the location of the artificial sound source.

Claims (44)

1. A method comprising:

capturing, using a microphone array, a sound event produced by a sound source within an environment of the microphone array as a first plurality of microphone signals;

producing, using a machine learning (ML) model, a spatial profile that identifies the sound source and a location within the environment at which the sound source is located based on the first plurality of microphone signals;

capturing, using the microphone array, a subsequent sound event within the environment as a second plurality of microphone signals;

determining whether the subsequent sound event originated from the location within the environment based on the second plurality of microphone signals and the spatial profile; and

responsive to determining that the sound event originated from the location, identifying the sound source based on the second plurality of microphone signals.

2. The method of claim 1 , wherein determining whether the subsequent sound event originated from the location comprises comparing spatial content of the second plurality of microphone signals and the spatial profile.

3. The method of claim 2 , wherein the spatial content comprises a direction of arrival (DoA) of the sound event with respect to the microphone array, wherein determining whether the subsequent sound event originated from the location comprises determining that the DoA matches at least a portion of the spatial profile based on the comparison.

4. The method of claim 1 , wherein identifying the sound source comprises:

extracting a spectral feature of the subsequent sound event from one or more microphone signals of the second plurality of microphone signals; and

comparing the extracted spectral feature with a stored spectral feature associated with the sound source.

5. The method of claim 1 further comprising extracting, from at least one microphone signal of the first plurality of microphone signals, 1) at least one spatial feature of the sound event that indicates the location of the sound source of the sound event with respect to the microphone array, and 2) at least one spectral feature of the sound event, wherein the spatial profile is produced as output of the ML model based on input of the at least one spatial feature and the at least one spectral feature.

6. The method of claim 1 , wherein the capturing of the sound event and the producing of the spatial profile occur during an enrollment process performed by an electronic device for the sound source, and the capturing of the subsequent sound event, determining, and identifying occur during an identification process subsequently performed by the electronic device.

7. The method of claim 1 , wherein the microphone array is a part of a smart speaker.

8. An electronic device, comprising:

a microphone array;

at least one processor; and

memory having instructions stored therein which when executed by the at least one processor causes the electronic device to:

capture, using the microphone array, a sound event produced by a sound source within an environment of the electronic device as a first plurality of microphone signals;

produce, using a machine learning (ML) model, a spatial profile that identifies the sound source and a location within the environment at which the sound source is located based on the first plurality of microphone signals;

capture, using the microphone array, a subsequent sound event within the environment as a second plurality of microphone signals;

determine whether the subsequent sound event originated from the location within the environment based on the second plurality of microphone signals and the spatial profile; and

responsive to determining that the sound event originated from the location, identify the sound source based on the second plurality of microphone signals.

9. The electronic device of claim 8 , wherein the instructions to determine whether the subsequent sound event originated from the location comprises instructions to compare spatial content of the second plurality of microphone signals and the spatial profile.

10. The electronic device of claim 9 , wherein the spatial content comprises a direction of arrival (DoA) of the sound event with respect to the microphone array, wherein the instructions to determine whether the subsequent sound event originated from the location comprises instructions to determine that the DoA matches at least a portion of the spatial profile based on the comparison.

11. The electronic device of claim 8 , wherein the instructions to identify the sound source comprises instructions to:

extract a spectral feature of the subsequent sound event from one or more microphone signals of the second plurality of microphone signals; and

compare the extracted spectral feature with a stored spectral feature associated with the sound source.

12. The electronic device of claim 8 , wherein the memory has further instructions to extract, from at least one microphone signal of the first plurality of microphone signals, 1) one or more spatial features of the sound event that indicates the location of the sound source of the sound event with respect to the microphone array, and 2) one or more spectral features of the sound event, wherein the spatial profile is produced as output of the ML model based on input of the one or more spatial features and the one or more spectral features.

13. The electronic device of claim 8 , wherein the capturing of the sound event and the producing of the spatial profile occur during an enrollment process performed by the electronic device for the sound source, and the capturing of the subsequent sound event, determining, and identifying occur during a subsequent identification process for the sound source.

14. The electronic device of claim 8 is a smart speaker.

15. Processing circuitry of an electronic device that is configured to:

capture, using a microphone array, a sound event produced by a sound source within an environment of the electronic device as a first plurality of microphone signals;

produce, using a machine learning (ML) model, a spatial profile that identifies the sound source and a location within the environment at which the sound source is located based on the first plurality of microphone signals;

capture, using the microphone array, a subsequent sound event within the environment as a second plurality of microphone signals;

determine whether the subsequent sound event originated from the location within the environment based on the second plurality of microphone signals and the spatial profile; and

responsive to determining that the sound event originated from the location, identify the sound source based on the second plurality of microphone signals.

16. The processing circuitry of claim 15 , wherein the processing circuitry determines whether the subsequent sound event originated from the location by comparing spatial content of the second plurality of microphone signals and the spatial profile.

17. The processing circuitry of claim 16 , wherein the spatial content comprises a direction of arrival (DoA) of the sound event with respect to the microphone array, wherein the processing circuitry determines whether the subsequent sound event originated from the location by determining that the DoA matches at least a portion of the spatial profile based on the comparison.

18. The processing circuitry of claim 15 , wherein the processing circuitry identifies the sound source by:

extracting a spectral feature of the subsequent sound event from one or more microphone signals of the second plurality of microphone signals; and

comparing the extracted spectral feature with a stored spectral feature associated with the sound source.

19. The processing circuitry of claim 15 , wherein the processing circuitry captures the sound event and produces the spatial profile during an enrollment process for the sound source, and the processing circuitry captures of the subsequent sound event, determines, and identifies during a subsequent identification process for the sound source.

20. The processing circuitry of claim 15 , wherein the electronic device is a smart speaker.

Continuity (3)
Continuation 18061753 · Dec 5, 2022
Continuation 17326208 · May 20, 2021
Related Publication 20240107254A1 · Mar 28, 2024
References Cited (24)
US 9697248B1 · Ahire · 2017 [cited by applicant]
US 11533577B2 · Taherian · 2022 [cited by examiner]
US 11863961B2 · Taherian · 2024 [cited by examiner]
US 20180122398A1 · Sporer et al. · 2018 [cited by applicant]
US 20190025400A1 · Venalainen · 2019 [cited by examiner]
US 20190132694A1 · Hanes et al. · 2019 [cited by applicant]
US 20190371324A1 · Powell et al. · 2019 [cited by applicant]
US 20200090644A1 · Klingler et al. · 2020 [cited by applicant]
US 20210020018A1 · Kim et al. · 2021 [cited by applicant]
IN 2015014471313 · 2015 [cited by applicant]
WO 2018069774A1 · 2018 [cited by applicant]
Ntalampiras et al., “On Acoustic Surveillance of Hazardous Situations”, 2009 IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19-24, 2009. pp. 165-168. [cited by applicant]
Vinyals et al., “Matching Networks for One Shot Learning”, arXiv: 1606.04080, Dec. 29, 2017, pp. 1-12. [cited by applicant]
Koch et al., “Siamese Neural Networks for One-shot Image Recognition”, Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 2015. JMLR: W&CP vol. 37, 8 pages. [cited by applicant]
Campbell et al., “Support Vector Machines using GMM Supervectors for Speaker Verification”, IEEE Signal Processing Letters, vol. 13, Issue: 5, May 2006, pp. 1-8. [cited by applicant]
Final Office Action of the U.S. Patent Office dated Jan. 13, 2022, for U.S. Appl. No. 16/564,775. [cited by applicant]
Non-Final Office Action of the U.S. Patent Office dated Apr. 6, 2022 for U.S. Appl. No. 16/564,775. [cited by applicant]
Non-Final Office Action of the U.S. Patent Office dated Aug. 13, 2021 for U.S. Appl. No. 16/564,775. [cited by applicant]
Gerhard, David, “Audio Signal Classification: History and Current Techniques”, Technical Report TR-CS Jul. 2003, Nov. 2003, 38 pages. [cited by applicant]
Green, Marc C., et al., “Acoustic Scene Classification Using Spatial Features”, Detection and Classification of Acoustic Scenes and Events 2017, Nov. 16, 2017, 4 pages. [cited by applicant]
“What is the basic difference in perception between acoustic and electronic/synthetic sound?”, Sep. 29, 2015, etrieved from the Internet: <https:/fwww.quora.com/What-is-the-basic-different-in-perception-between-acoustic… [cited by applicant]
Dapayiannis, Constantinos, et al., “Detecting Media Sound Presence in Acoustic Scenes”, Interspeech 2018, Sep. 2, 2018, pp. 1363-1367. [cited by applicant]
“How do we identify whether a sound is live or recorded?”, 2016, retrieved from the Internet: <htlps:/fwww.esearchgate.net/post/How_do_we_identify_whether_a_sound_is_live_or_recorded>, 3 pages. [cited by applicant]
“Can You Tell the Difference Between Natural and Artificial Sounds?”, Oct. 6, 2015, retrieved from the Internet: 9 i::https://www.teacherspayteachers.com/Product/Can-You-Tell-the-Difference-Between-Natural-and Artificia… [cited by applicant]