IP Library Granted Patent US 9,601,119
Granted Patent B2
US 9,601,119 · App. 14/481,918 · Granted Mar 21, 2017

Systems and methods for segmenting and/or classifying an audio signal from transformed audio information

Inventors: David C. Bradley (La Jolla, CA); Robert N. Hilton (San Diego, CA); Daniel S. Goldin (Malibu, CA); Nicholas K. Fisher (Los Angeles, CA); Derrick R. Roos (San Diego, CA); Eric Wiewiora (San Diego, CA)
Assignee: KnuEdge Incorporated
G10L17/02G10L25/51H04R3/00H04R29/00G10L25/84H04R2430/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,601,119
App. No.
14/481,918
Granted
Mar 21, 2017
Kind
B2
Abstract

A system and method may be provided to segment and/or classify an audio signal from transformed audio information. Transformed audio information representing a sound may be obtained. The transformed audio information may specify magnitude of a coefficient related to energy amplitude as a function of frequency for the audio signal and time. Features associated with the audio signal may be obtained from the transformed audio information. Individual ones of the features may be associated with a feature score relative to a predetermined speaker model. An aggregate score may be obtained based on the feature scores according to a weighting scheme. The weighting scheme may be associated with a noise and/or SNR estimation. The aggregate score may be used for segmentation to identify portions of the audio signal containing speech of one or more different speakers. For classification, the aggregate score may be used to determine a likely speaker model to identify a source of the sound in the audio signal.

Claims (27)

1. A system configured for segmenting an audio signal to identify portions of the audio signal containing speech of one or more different speakers, the system comprising:

one or more processors configured by computer readable instructions to:

obtain transformed audio information representing a sound including determining harmonic paths for individual harmonics of the sound based on fractional chirp rate and harmonic number, wherein the transformed audio information specifies magnitude of a coefficient related to energy amplitude as a function of frequency for the audio signal and time;

obtain features associated with the audio signal from the transformed audio information, individual ones of the features being associated with a feature score; and

obtain an aggregate score based on the feature scores, the aggregate score being used for segmentation to identify portions of the audio signal containing speech of one or more different speakers.

2. The system of claim 1 , wherein the segmentation includes dividing sounds represented in the audio signal into groups corresponding to different sources.

3. The system of claim 1 , wherein obtaining the transformed audio information includes determining an amplitude value for individual harmonics at individual time windows.

4. The system of claim 1 , wherein the one or more processors are further configured by computer readable instructions to obtain spectral slope information based on the transformed audio information as a feature associated with the audio signal.

5. The system of claim 1 , wherein the one or more processors are further configured by computer readable instructions to obtain a signal-to-noise ratio estimation as a time-varying quantity associated with the audio signal.

6. The system of claim 1 , wherein obtaining an aggregate score based on the feature scores is performed in accordance with a weighting scheme.

7. The system of claim 6 , wherein the weighting scheme is associated with a noise estimation.

8. The system of claim 6 , wherein the one or more processors are further configured by computer readable instructions to perform training operations on the audio signal to determine characteristics of the audio signal that indicate a set of score weights associated with the weighting scheme.

9. The system of claim 6 , wherein the one or more processors are further configured by computer readable instructions to perform training operations on the audio signal to determine characteristics of conditions pertaining to the recording of the audio signal that indicate a set of score weights associated with the weighting scheme.

10. The system of claim 1 , wherein the one or more processors are further configured by computer readable instructions to determine a likely speaker model to identify a source of the sound in the audio signal based on the aggregate score.

11. A computer-implemented method for segmenting an audio signal to identify portions of the audio signal containing speech of one or more different speakers, the method being implemented in a computer system that includes one or more physical processors, the method comprising:

obtaining, at the one or more physical processors, transformed audio information representing a sound, wherein the transformed audio information specifies magnitude of a coefficient related to energy amplitude as a function of frequency for the audio signal and time;

obtaining, at the one or more physical processors, features associated with the audio signal from the transformed audio information, individual ones of the features being associated with a feature score; and

obtaining, at the one or more physical processors, an aggregate score based on the feature scores in accordance with a weighting scheme, the aggregate score being used for segmentation to identify portions of the audio signal containing speech of one or more different speakers.

12. The method of claim 11 , wherein the segmentation includes dividing sounds represented in the audio signal into groups corresponding to different sources.

13. The method of claim 11 , wherein obtaining the transformed audio information includes determining harmonic paths for individual harmonics of the sound based on fractional chirp rate and harmonic number.

14. The method of claim 11 , wherein obtaining the transformed audio information includes determining an amplitude value for individual harmonics at individual time windows.

15. The method of claim 11 , further comprising obtaining, at the one or more physical processors, spectral slope information based on the transformed audio information as a feature associated with the audio signal.

16. The method of claim 11 , further comprising obtaining, at the one or more physical processors, a signal-to-noise ratio estimation as a time-varying quantity associated with the audio signal.

17. The method of claim 11 , wherein the weighting scheme is associated with a noise estimation.

18. The method of claim 11 , further comprising performing, at the one or more physical processors, training operations on the audio signal to determine characteristics of the audio signal that indicate a set of score weights associated with the weighting scheme.

19. The method of claim 11 , further comprising performing, at the one or more physical processors, training operations on the audio signal to determine characteristics of conditions pertaining to the recording of the audio signal that indicate a set of score weights associated with the weighting scheme.

20. The method of claim 11 , further comprising determining a likely speaker model to identify a source of the sound in the audio signal based on the aggregate score.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2026
From: PATTI, ROBERT S
To: TEATRO, INC.
Reel/Frame 074966/0181 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2018
From: KNUEDGE, INC.
To: FRIDAY HARBOR LLC
Reel/Frame 047156/0582 →
SECURITY INTEREST Recorded Oct 27, 2017
From: KNUEDGE INCORPORATED
To: XL INNOVATE FUND, LP
Reel/Frame 044637/0011 →
SECURITY INTEREST Recorded Nov 11, 2016
From: KNUEDGE INCORPORATED
To: XL INNOVATE FUND, L.P.
Reel/Frame 040601/0917 →
CHANGE OF NAME Recorded Jun 9, 2016
From: THE INTELLISIS CORPORATION
To: KNUEDGE INCORPORATED
Reel/Frame 038926/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2014
From: BRADLEY, DAVID C.; HILTON, ROBERT N.; GOLDIN, DANIEL S.; FISHER, NICHOLAS K.; ROOS, DERRICK R.; WIEWIORA, ERIC
To: THE INTELLISIS CORPORATION
Reel/Frame 033705/0734 →
Continuity (4)
Continuation 13205507 · Aug 8, 2011
Provisional Application 61464493 · Mar 7, 2011
Provisional Application 61454756 · Mar 21, 2011
Related Publication 20140376730A1 · Dec 25, 2014