IP Library › Granted Patent US 11,521,598
Granted Patent B2
US 11,521,598 · App. 16/564,775 · Granted Dec 6, 2022

Systems and methods for classifying sounds

Inventors: Daniel C. Klingler (Sunnyvale, CA); Carlos M. Avendano (Campbell, CA); Hyung-Suk Kim (Santa Clara, CA); Miquel Espi Marques (Cupertino, CA)
Assignee: APPLE INC.
G10L15/16G06N3/08G06N20/00G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,521,598
App. No.
16/564,775
Granted
Dec 6, 2022
Kind
B2
Abstract

An electronic device has one or more microphones that pick up a sound. At least one feature extractor processes the audio signals from the microphones, that contain the picked up the sound, to determine several features for the sound. The electronic device also includes a classifier that has a machine learning model which is configured to determine a sound classification, such as artificial versus natural for the sound, based upon at least one of the determined features. Other aspects are also described and claimed.

Claims (38)

1. An electronic device, comprising:

one or more microphones configured to receive a sound;

a processor and memory having stored therein a plurality of instructions that when executed by the processor implement

at least one feature detector configured to receive one or more audio signals from the one or more microphones that comprise the sound, process the one or more audio signals to determine a plurality of features for the one or more audio signals including i) a directional feature indicating that a sound source has a static location or a dynamic location and ii) a sound class feature indicating that the sound source is producing music or speech; and

a sound classifier including a machine learning model that is configured to receive the directional feature and the sound class feature from the at least one feature detector, and determine a sound classification for the one or more audio signals based upon the directional feature and the sound class feature, the determined sound classification being natural sound which is sound that is generated by a natural sound source in an environment of the electronic device versus artificial sound which is sound that is generated by an artificial sound source being a loudspeaker that is external to the electronic device.

2. The electronic device according to claim 1 , wherein the electronic device performs one or more functions or actions based upon the determined sound classification.

3. The electronic device according to claim 1 , wherein if the sound classifier determines that the sound is a natural sound, then the device is allowed to perform one or more actions or functions based on the sound, and wherein if the sound classifier determines that the sound is an artificial sound, the device is prevented from performing those one or more actions or functions based on the sound.

4. The electronic device according to claim 1 , wherein the plurality of features includes at least three features including a distortion feature.

5. The electronic device according to claim 4 wherein the determined sound classification, being natural sound versus artificial sound, is based on all of the at least three features including the sound class feature indicative of music or speech, the distortion feature, and the directional feature.

6. The electronic device according to claim 1 , wherein the device is a smart phone, a smart speaker, a tablet computer, a laptop computer, or a desktop computer.

7. The electronic device according to claim 1 , wherein the sound classifier accesses a database that stores historical sound data, and wherein the sound classifier determines the classification based upon the historical sound data.

8. The electronic device according to claim 1 , further comprising:

a wireless communications receiver configured to receive a signal from an additional electronic device indicating that the sound originated from a loudspeaker of the additional electronic device, wherein the sound classifier classifies the sound as an artificial sound responsive to the signal.

9. An electronic device, comprising:

a plurality of microphones configured to produce a plurality of audio signals in response to receiving a sound;

at least one feature detector configured to receive the plurality of audio signals from the plurality of microphones, and determine a plurality of features relating to the sound including i) a directional feature indicating that a sound source has a static location or a dynamic location, ii) a sound class feature indicating that the sound source is producing music or speech, and iii) a distortion feature; and

a classifier including a machine learning model that is configured to receive the determined plurality of features, and determine a sound classification for the sound based upon the directional feature and the sound class feature, the sound classification being natural sound which is sound that is generated by a natural sound source in an environment of the electronic device versus artificial sound which is sound that is generated by an artificial sound source being a loudspeaker that is external to the electronic device;

wherein the electronic device performs one or more functions or actions based upon the determined sound classification.

10. The electronic device according to claim 9 , wherein the classifier determines the sound classification, being natural sound versus artificial sound, based on all of the at least three features of the sound class feature indicative of music or speech, the distortion feature, and the directional feature.

11. The electronic device according to claim 9 , wherein the device is a smart phone, a smart speaker, a tablet computer, a laptop computer, or a desktop computer.

12. The electronic device according to claim 9 , wherein the classifier accesses a database storing historical sound data, and wherein the classifier determines the classification based upon the historical sound data.

13. A method performed by a processor of an electronic device for discriminating between two classes of sounds, comprising:

capturing a sound using a plurality of microphones, as a recorded sound;

digitally processing the recorded sound to determine at least two features of the recorded sound that include a directional feature indicating that a sound source has a static location or a dynamic location and a sound class feature indicating that the sound source is producing music or speech;

determining a classification of the sound based on the directional feature and the sound class feature using a machine learning model the classification being natural sound which is sound that is generated by a natural sound source in an environment of the electronic device versus artificial sound which is sound that is generated by an artificial sound source being a loudspeaker that is external to the electronic device; and

performing a virtual assistant action based upon the determined classification.

14. The method according to claim 13 , further comprising processing the recorded sound to determine a third feature being a distortion feature, wherein determining the classification, being natural sound versus artificial sound, is based on all of the at least three features of the sound class feature, including the feature indicative of music or speech, the distortion feature, and the directional feature.

15. The method according to claim 13 , further comprising:

accessing a database that stores historical sound data; and

determining the classification of the recorded sound based upon the historical sound data accessed from the database.

16. The method according to claim 13 , further comprising:

receiving a signal from an additional electronic device wherein the signal indicates that the sound originated from the additional electronic device; and

in response to the signal, determining that the sound is an artificial sound.

17. The method according to claim 13 , further comprising training the machine learning model by:

collecting a data corpus including a variety of labeled natural and artificial sounds;

partitioning the data in the data corpus into a training data set and a testing data set;

calibrating the machine learning model to classify the data using the training data set; and

determining the accuracy of the machine learning model using the testing data set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2019
From: KLINGLER, DANIEL C.; AVENDANO, CARLOS M.; KIM, HYUNG-SUK; MARQUES, MIQUEL ESPI
To: APPLE INC.
Reel/Frame 050562/0598 →
Continuity (2)
Provisional Application 62733026 · Sep 18, 2018
Related Publication 20200090644A1 · Mar 19, 2020