IP Library Granted Patent US 9,870,785
Granted Patent B2
US 9,870,785 · App. 14/969,029 · Granted Jan 16, 2018

Determining features of harmonic signals

Inventors: David Carlson Bradley (San Diego, CA); Yao Huang Morin (San Diego, CA); Massimo Mascaro (San Diego, CA); Janis I. Intoy (San Diego, CA); Sean Michael O'Connor (San Diego, CA); Ellisha Natalie Marongelli (San Diego, CA); Robert Nicholas Hilton (San Diego, CA)
Assignee: KnuEdge Incorporated
G10L25/90G10L15/02G10L17/02G10L21/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,870,785
App. No.
14/969,029
Granted
Jan 16, 2018
Kind
B2
Abstract

Features that may be computed from a harmonic signal include a fractional chirp rate, a pitch, and amplitudes of the harmonics. A fractional chirp rate may be estimated, for example, by computing scores corresponding to different fractional chirp rates and selecting a highest score. A first pitch may be computed from a frequency representation that is computed using the estimated fractional chirp rate, for example, by using peak-to-peak distances in the frequency distribution. A second pitch may be computed using the first pitch, and a frequency representation of the signal, for example, by using correlations of portions of the frequency representation. Amplitudes of harmonics of the signal may be determined using the estimated fractional chirp rate and second pitch. Any of the estimated fractional chirp rate, second pitch, and harmonic amplitudes may be used for further processing, such as speech recognition, speaker verification, speaker identification, or signal reconstruction.

Claims (63)

1. A computer-implemented method for automatic speaker recognition, the method comprising:

obtaining a first portion of a speech signal;

computing, using one or more processors, a first estimated fractional chirp rate from the first portion of the speech signal;

computing, using the one or more processors, a first frequency representation from the first portion of the speech signal using the first estimated fractional chirp rate;

computing, using the one or more processors, a first pitch estimate from the first portion of the speech signal using a plurality of peak-to-peak distances in the first frequency representation;

computing, using the one or more processors, a second pitch estimate from the first portion of the speech signal using the first pitch estimate and a correlation between a first frequency portion of a second frequency representation of the first portion of the speech signal, and a second frequency portion of the second frequency representation;

obtaining a second portion of the speech signal;

computing, using the one or more processors, a second estimated fractional chirp rate from the second portion of the speech signal;

computing, using the one or more processors, a third frequency representation from the second portion of the speech signal using the second estimated fractional chirp rate;

computing, using the one or more processors, a third pitch estimate from the second portion of the speech signal using a plurality of peak-to-peak distances in the third frequency representation;

computing, using the one or more processors, a fourth pitch estimate from the second portion of the speech signal using the third pitch estimate and a correlation between a first frequency portion of a fourth frequency representation of the second portion of the speech signal, and a second frequency portion of the fourth frequency representation;

generating a sequence of pitch estimates, the sequence of pitch estimates comprising the second pitch estimate and the fourth pitch estimate;

applying the sequence of pitch estimates to recognize a speaker as a source of the speech signal.

2. The method of claim 1 , further comprising computing amplitudes of a plurality of harmonics of the first portion of the speech signal using the first estimated fractional chirp rate and the second pitch estimate.

3. The method of claim 2 , further comprising:

computing a feature vector using the amplitudes of the plurality of harmonics; and

using the feature vector to recognize the speaker.

4. The method of claim 1 , wherein the second frequency representation is substantially equal to the first frequency representation.

5. The method of claim 1 , wherein computing the first estimated fractional chirp rate comprises computing a plurality of scores, wherein the plurality of scores comprise a first score and a second score, the first score is computed using a first fractional chirp rate, the second score is computed using a second fractional chirp rate, and the first estimated fractional chirp rate is computed by selecting one of the first fractional chirp rate or the second fractional chirp rate based on corresponding scores.

6. The method of claim 5 , wherein the first score is computed using an autocorrelation of a frequency representation that is computed using the first fractional chirp rate.

7. The method of claim 1 , wherein the first frequency representation is computed by performing inner products of the first portion of the speech signal with a function of frequency and chirp rate, and wherein the chirp rate of the function increases with the frequency.

8. The method of claim 1 , wherein the first pitch estimate is computed using an estimated cumulative distribution function of the plurality of peak-to-peak distances in the first frequency representation.

9. The method of claim 1 , wherein the first frequency portion of the second frequency representation corresponds to a first multiple of the first pitch estimate and the second frequency portion of the second frequency representation corresponds to a second multiple of the first pitch estimate.

10. The method of claim 1 , wherein the third fractional chirp rate is substantially equal to the first fractional chirp rate.

11. The method of claim 1 , wherein the fourth fractional chirp rate is substantially equal to the second fractional chirp rate.

12. A system for automatic speaker recognition, the system comprising one or more computing devices comprising at least one processor and at least one memory, the one or more computing devices configured to:

obtain a first portion of a speech signal;

compute a first estimated fractional chirp rate from the first portion of the speech signal;

compute a first frequency representation from the first portion of the speech signal using the first estimated fractional chirp rate;

compute a first pitch estimate from the first portion of the speech signal using a plurality of peak-to-peak distances in the first frequency representation;

compute a second pitch estimate from the first portion of the speech signal using the first pitch estimate and a correlation between a first frequency portion of a second frequency representation of the first portion of the speech signal, and a second frequency portion of the second frequency representation;

obtain a second portion of the speech signal;

compute a second estimated fractional chirp rate from the second portion of the speech signal;

compute a third frequency representation from the second portion of the speech signal using the second estimated fractional chirp rate;

compute a third pitch estimate from the second portion of the speech signal using a plurality of peak-to-peak distances in the third frequency representation;

compute a fourth pitch estimate from the second portion of the speech signal using the third pitch estimate and a correlation between a first frequency portion of a fourth frequency representation of the second portion of the speech signal, and a second frequency portion of the fourth frequency representation;

generate a sequence of pitch estimates, the sequence of pitch estimates comprising the second pitch estimate and the fourth pitch estimate; and

apply the sequence of pitch estimates to recognize a speaker as a source of the speech signal.

13. The system of claim 12 , wherein the one or more computing devices are further configured to compute amplitudes of a plurality of harmonics of the first portion of the speech signal using the first estimated fractional chirp rate and the second pitch estimate.

14. The system of claim 12 , wherein the second frequency representation is substantially equal to the first frequency representation.

15. The system of claim 12 , wherein the first frequency representation is computed using a pitch-velocity transformation.

16. The system of claim 12 , wherein the first pitch estimate is computed using a histogram of the plurality of peak-to-peak distances in the first frequency representation.

17. The system of claim 12 , wherein the one or more computing devices are further configured to compute the second pitch estimate by computing a correlation using a reversed version of first frequency portion of the second frequency representation.

18. One or more non-transitory computer-readable media comprising computer executable instructions that, when executed, cause at least one processor to perform actions comprising:

obtaining a first portion of a speech signal;

computing a first estimated fractional chirp rate from the first portion of the speech signal;

computing a first frequency representation from the first portion of the speech signal using the first estimated fractional chirp rate;

computing a first pitch estimate from the first portion of the speech signal using a plurality of peak-to-peak distances in the first frequency representation;

computing a second pitch estimate from the first portion of the speech signal using the first pitch estimate and a correlation between a first frequency portion of a second frequency representation of the first portion of the speech signal, and a second frequency portion of the second frequency representation;

obtaining a second portion of the speech signal;

computing a second estimated fractional chirp rate from the second portion of the speech signal;

computing a third frequency representation from the second portion of the speech signal using the second estimated fractional chirp rate;

computing a third pitch estimate from the second portion of the speech signal using a plurality of peak-to-peak distances in the third frequency representation;

computing a fourth pitch estimate from the second portion of the speech signal using the third pitch estimate and a correlation between a first frequency portion of a fourth frequency representation of the second portion of the speech signal, and a second frequency portion of the fourth frequency representation;

generating a sequence of pitch estimates, the sequence of pitch estimates comprising the second pitch estimate and the fourth pitch estimate;

applying the sequence of pitch estimates to recognize a speaker as a source of the speech signal.

19. The one or more non-transitory computer-readable media of claim 18 , further comprising computer executable instructions that, when executed, cause the at least one processor to perform actions comprising computing amplitudes of a plurality of harmonics of the first portion of the speech signal using the first estimated fractional chirp rate and the second pitch estimate.

20. The one or more non-transitory computer-readable media of claim 19 , further comprising computer executable instructions that, when executed, cause the at least one processor to perform actions comprising:

computing amplitudes of a plurality of harmonics of the first portion of the speech signal using the first estimated fractional chirp rate and the second pitch estimate;

computing a feature vector using the amplitudes;

computing second amplitudes of a second plurality of harmonics of the second portion of the speech signal using the second estimated fractional chirp rate and the fourth pitch estimate;

computing a second feature vector using the second amplitudes; and

using the second feature vector to recognize the speaker.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2026
From: PATTI, ROBERT S
To: TEATRO, INC.
Reel/Frame 074966/0181 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2018
From: KNUEDGE, INC.
To: FRIDAY HARBOR LLC
Reel/Frame 047156/0582 →
SECURITY INTEREST Recorded Oct 27, 2017
From: KNUEDGE INCORPORATED
To: XL INNOVATE FUND, LP
Reel/Frame 044637/0011 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2016
From: BRADLEY, DAVID CARLSON; MORIN, YAO HUANG; MASCARO, MASSIMO; INTOY, JANIS I.; O'CONNOR, SEAN MICHAEL; MARONGELLI, ELLISHA NATALIE; HILTON, ROBERT NICHOLAS
To: KNUEDGE INCORPORATED
Reel/Frame 040681/0445 →
SECURITY INTEREST Recorded Nov 11, 2016
From: KNUEDGE INCORPORATED
To: XL INNOVATE FUND, L.P.
Reel/Frame 040601/0917 →
Continuity (5)
Provisional Application 62112850 · Feb 6, 2015
Provisional Application 62112832 · Feb 6, 2015
Provisional Application 62112796 · Feb 6, 2015
Provisional Application 62112836 · Feb 6, 2015
Related Publication 20160232906A1 · Aug 11, 2016