IP Library Patent Application 15647798
Patent Application
App. No. 15/647,798

SOUND-RECOGNITION SYSTEM BASED ON A SOUND LANGUAGE AND ASSOCIATED ANNOTATIONS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
15/647,798
Abstract

The disclosed embodiments provide a system for recognizing a sound event in raw sound. During operation, the system receives the raw sound, wherein the raw sound comprises a sequence of digital samples of sound. Next, the system segments the raw sound into a sequence of tiles, wherein each tile comprises a set of consecutive digital samples. The system then converts the sequence of tiles into a sequence of snips, wherein each snip includes a symbol representing an associated tile in the sequence of tiles. Next, the system generates annotations for the sequence of snips and the raw sound, wherein each annotation specifies a property associated with one or more snips in the sequence of snips or the raw sound. Finally, the system recognizes the sound event based on the generated annotations.

Claims (87)

1 . A method for recognizing a sound event in raw sound, comprising:

receiving the raw sound, wherein the raw sound comprises a sequence of digital samples of sound;

segmenting the raw sound into a sequence of tiles, wherein each tile comprises a set of consecutive digital samples;

converting the sequence of tiles into a sequence of snips, wherein each snip includes a symbol representing an associated tile in the sequence of tiles, wherein each snip takes up less space than an associated tile, wherein each snip is stored in a canonical representation, and wherein the sequence of snips is searchable;

generating annotations for the sequence of snips and the raw sound, wherein each annotation specifies a property associated with one or more snips in the sequence of snips or the raw sound; and

recognizing the sound event based on the generated annotations.

2 . The method of claim 1 , wherein converting the sequence of tiles into the sequence of snips comprises:

identifying tile features for each tile in the sequence of tiles;

performing a clustering operation based on the tile features to identify clusters of tiles and to associate each tile with a cluster;

associating each identified cluster with a unique symbol; and

representing the sequence of tiles as a sequence of symbols representing clusters, wherein the symbols are associated with individual tiles in the sequence of tiles.

3 . The method of claim 1 , wherein the sequence of tiles includes one or more of the following:

overlapping tiles;

non-overlapping tiles;

tiles having variable sizes; and

one or more gaps between tiles in the sequence of tiles, wherein each gap comprises a segment of the raw sound that is not covered by a tile.

4 . The method of claim 1 , wherein annotating the sequence of snips involves:

generating grounded annotations, which are associated with specific segments of raw sound; and

generating higher-level annotations, which are associated with lower-level annotations.

5 . The method of claim 1 , wherein an annotation can include an acoustic annotation, which specifies an acoustic property associated with a sound feature.

6 . The method of claim 1 , wherein an annotation can include a semantic tag.

7 . The method of claim 6 , wherein an annotation can include a higher-level semantic tag, which is associated with one or more lower-level semantic tags.

8 . The method of claim 1 , wherein recognizing the sound event based on the generated annotations additionally involves considering other sensor inputs, which are associated with the raw sound.

9 . The method of claim 1 ,

wherein an annotation for each snip includes a centroid distance parameter, which specifies a distance between a feature vector for a tile associated with the snip and a mean feature vector for all tiles associated with the snip; and

wherein the method further comprises detecting an anomaly in the sequence of snips if the centroid distance for one or more snips in the sequence of snips exceeds a threshold value.

10 . The method of claim 1 ,

wherein an annotation for each snip includes a rareness score that specifies a rareness of the snip; and

wherein the method further comprises detecting an anomaly in the sequence of snips if rareness scores for a proximate set of snips in the sequence of snips exceed a threshold value.

11 . A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for recognizing a sound event in raw sound, the method comprising:

receiving the raw sound, wherein the raw sound comprises a sequence of digital samples of sound;

segmenting the raw sound into a sequence of tiles, wherein each tile comprises a set of consecutive digital samples;

converting the sequence of tiles into a sequence of snips, wherein each snip includes a symbol representing an associated tile in the sequence of tiles wherein each snip takes up less space than an associated tile, wherein each snip is stored in a canonical representation, and wherein the sequence of snips is searchable;

generating annotations for the sequence of snips and the raw sound, wherein each annotation specifies a property associated with one or more snips in the sequence of snips or the raw sound; and

recognizing the sound event based on the generated annotations.

12 . The non-transitory computer-readable storage medium of claim 11 , wherein converting the sequence of tiles into the sequence of snips comprises:

identifying tile features for each tile in the sequence of tiles;

performing a clustering operation based on the tile features to identify clusters of tiles and to associate each tile with a cluster;

associating each identified cluster with a unique symbol; and

representing the sequence of tiles as a sequence of symbols representing clusters, wherein the symbols are associated with individual tiles in the sequence of tiles.

13 . The non-transitory computer-readable storage medium of claim 11 , wherein the sequence of tiles includes one or more of the following:

overlapping tiles;

non-overlapping tiles;

tiles having variable sizes; and

one or more gaps between tiles in the sequence of tiles, wherein each gap comprises a segment of the raw sound that is not covered by a tile.

14 . The non-transitory computer-readable storage medium of claim 11 , wherein annotating the sequence of snips involves:

generating grounded annotations, which are associated with specific segments of raw sound; and

generating higher-level annotations, which are associated with lower-level annotations.

15 . The non-transitory computer-readable storage medium of claim 11 , wherein an annotation can include an acoustic annotation, which specifies an acoustic property associated with a sound feature.

16 . The non-transitory computer-readable storage medium of claim 11 , wherein an annotation can include a semantic tag.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein an annotation can include a higher-level semantic tag, which is associated with one or more lower-level semantic tags.

18 . The non-transitory computer-readable storage medium of claim 11 , wherein recognizing the sound event based on the generated annotations additionally involves considering other sensor inputs, which are associated with the raw sound.

19 . The non-transitory computer-readable storage medium of claim 11 ,

wherein an annotation for each snip includes a centroid distance parameter, which specifies a distance between a feature vector for a tile associated with the snip and a mean feature vector for all tiles associated with the snip; and

wherein the method further comprises detecting an anomaly in the sequence of snips if the centroid distance for one or more snips in the sequence of snips exceeds a threshold value.

20 . The non-transitory computer-readable storage medium of claim 11 ,

wherein an annotation for each snip includes a rareness score that specifies a rareness of the snip; and

wherein the method further comprises detecting an anomaly in the sequence of snips if rareness scores for a proximate set of snips in the sequence of snips exceed a threshold value.

21 . A system that recognizes a sound event in raw sound, comprising:

at least one processor and at least one associated memory; and

a sound-event-recognition mechanism that executes on the at least one processor, wherein during operation, the sound-event-recognition mechanism:

segments the raw sound into a sequence of tiles, wherein each tile comprises a set of consecutive digital samples;

converts the sequence of tiles into a sequence of snips, wherein each snip includes a symbol representing an associated tile in the sequence of tiles, wherein each snip takes up less space than an associated tile, wherein each snip is stored in a canonical representation, and wherein the sequence of snips is searchable;

generates annotations for the sequence of snips and the raw sound, wherein each annotation specifies a property associated with one or more snips in the sequence of snips or the raw sound; and

recognizes the sound event based on the generated annotations.

22 . The system of claim 21 , wherein while converting the sequence of tiles into the sequence of snips, the sound-event-recognition mechanism:

identifies tile features for each tile in the sequence of tiles;

performs a clustering operation based on the tile features to identify clusters of tiles and to associate each tile with a cluster;

associates each identified cluster with a unique symbol; and

represents the sequence of tiles as a sequence of symbols representing clusters, wherein the symbols are associated with individual tiles in the sequence of tiles.

23 . The system of claim 21 , wherein the sequence of tiles includes one or more of the following:

overlapping tiles;

non-overlapping tiles;

tiles having variable sizes; and

one or more gaps between tiles in the sequence of tiles, wherein each gap comprises a segment of the raw sound that is not covered by a tile.

24 . The system of claim 21 , wherein while annotating the sequence of snips, the sound-event-recognition mechanism:

generates grounded annotations, which are associated with specific segments of raw sound; and

generates higher-level annotations, which are associated with lower-level annotations.

25 . The system of claim 21 , wherein an annotation can include an acoustic annotation, which specifies an acoustic property associated with a sound feature.

26 . The system of claim 21 , wherein an annotation can include a semantic tag.

27 . The system of claim 26 , wherein an annotation can include a higher-level semantic tag, which is associated with one or more lower-level semantic tags.

28 . The system of claim 21 , wherein while recognizing the sound event based on the generated annotations, the sound-event-recognition mechanism additionally involves considering other sensor inputs, which are associated with the raw sound.

29 . The system of claim 21 ,

wherein an annotation for each snip includes a centroid distance parameter, which specifies a distance between a feature vector for a tile associated with the snip and a mean feature vector for all tiles associated with the snip; and

wherein the sound-event-recognition mechanism additionally detects an anomaly in the sequence of snips if the centroid distance for one or more snips in the sequence of snips exceeds a threshold value.

30 . The system of claim 21 , wherein an annotation for each snip includes a rareness score that specifies a rareness of the snip; and

wherein the sound-event-recognition mechanism additionally detects an anomaly in the sequence of snips if rareness scores for a proximate set of snips in the sequence of snips exceed a threshold value.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2020
From: OTOSENSE, INC.
To: ANALOG DEVICES, INC.
Reel/Frame 053098/0719 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2017
From: WHALEN, THOR C.; CHRISTIAN, SEBASTIEN J.V.
To: OTOSENSE, INC.
Reel/Frame 043175/0186 →