IP Library › Granted Patent US 9,947,323
Granted Patent B2
US 9,947,323 · App. 15/088,500 · Granted Apr 17, 2018

Synthetic oversampling to enhance speaker identification or verification

Inventors: Narayan Biswal (Folsom, CA); Gokcen Cilingir (Sunnyvale, CA); Barnan Das (San Jose, CA)
Assignee: Intel Corporation
G10L17/06G10L17/02G10L17/04G10L17/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,947,323
App. No.
15/088,500
Granted
Apr 17, 2018
Kind
B2
Abstract

An apparatus for oversampling audio signals is described herein. The apparatus includes one or more microphones to receive audio signals and an extractor to extract a set of feature points from the audio signals. The apparatus also includes a processing unit to determine a distance between each pair of feature points and an oversampling unit to generate a plurality of new feature points based the distance between each pair of feature points.

Claims (47)

1. An apparatus, for oversampling audio signals comprising:

one or more microphones to receive audio signals;

an extractor to extract a set of feature points from the audio signals;

a processing unit to determine a distance between each pair of feature points of the set of feature points; and

an oversampling unit to generate a plurality of new feature points based on the distance between each pair of feature points and to generate a speaker voice model wherein the generation of the plurality of new feature points generates new feature points that number at least 50% of a number of feature points in the set of feature points.

2. The apparatus of claim 1 , wherein a plurality of feature vectors are created using the set of feature points and the plurality of new feature points.

3. The apparatus of claim 2 , wherein the plurality of feature vectors are determined to be valid according to voice activity detection.

4. The apparatus of claim 1 , wherein the distance between each pair of feature points is a Euclidean distance.

5. The apparatus of claim 1 , wherein the oversampling unit is to generate the plurality of new feature points based the distance between each pair of feature points, wherein each pair of feature points are valid feature points according to voice activity detection.

6. The apparatus of claim 1 , wherein the oversampling is a pre-processing function of a voice biometric system.

7. The apparatus of claim 1 , wherein the speaker voice model is based on the set of feature points and the plurality of new feature points.

8. The apparatus of claim 1 , wherein an authentication decision is based on the set of feature points and the plurality of new feature points.

9. The apparatus of claim 1 , wherein the oversampling results in low false rejection and false acceptance rates.

10. The apparatus of claim 1 , wherein an utterance length for audio signal capture is directly proportional to the set of feature points plus the plurality of new feature points.

11. An method for oversampling audio signals, comprising:

capturing audio signals from an utterance;

extracting a set of feature points from the audio signals;

determining a distance between each pair of feature points of the set of feature points;

generating a plurality of new feature points based on the distance between each pair of feature points of the set of feature wherein the generation of the plurality of new feature points generates new feature points that number at least 50% of a number of feature points in the set of feature points; and

identifying a user via recognition decisions based on the plurality of new feature points and the set of feature points.

12. The method of claim 11 , wherein the utterance is a phrase independent utterance.

13. The method of claim 11 , wherein the utterance is a phrase dependent utterance.

14. The method of claim 11 , wherein the set of feature points and the plurality of new feature points are used to train a speaker model.

15. The method of claim 11 , wherein an utterance length for audio signal capture is directly proportional to the set of feature points plus the plurality of new feature points.

16. The method of claim 11 , wherein the plurality of new feature points and the set of feature points are Mel-frequency cepstral coefficients (MFCC).

17. A system for oversampling audio signals comprising, comprising:

a microphone;

a memory that is to store instructions and that is communicatively coupled to the microphone; and

a processor communicatively coupled to the microphone and the memory,

wherein when the processor is to execute the instructions, the processor is to:

receive audio signals from an utterance;

extract a set of feature points from the audio signals;

determine a distance between each pair of feature points of the set of feature points

generating a plurality of new feature points based on the distance between each pair of feature points of the set of feature points, wherein the generation of the plurality of new feature points generates new feature points that number at least 50% of a number of feature points in the set of feature points;

extrapolate a plurality of feature vectors based on the plurality of new feature points and the set of feature points; and

identify a user via recognition decisions based on the plurality of new feature points, the set of feature points, and the plurality of feature vectors.

18. The system of claim 17 , wherein the processor is to build a speaker voice model based on the plurality of new feature points and the set of feature points.

19. The system of claim 17 , wherein the plurality of new feature points and the set of feature points are Mel-frequency cepstral coefficients (MFCC), linear predictive cepstral coefficients (LPCCs), line spectral frequencies (LSFs), and perceptual linear prediction (PLP) coefficients, or any combination thereof.

20. The system of claim 17 , wherein the plurality of feature vectors are determined to be valid according to voice activity detection.

21. A non-transitory computer-readable medium, comprising instructions that, when executed by a processor, direct the processor to

capture audio signals from an utterance;

extract a set of feature points from the audio signals;

determine a distance between each pair of feature points of the set of feature points;

generate a plurality of new feature points based on the distance between each pair of feature points of the set of feature points, wherein the generation of the plurality of new feature points generates new feature points that number at least 50% of a number of feature points in the set of feature points; and

identify a user via recognition decisions based on the plurality of new feature points and the set of feature points.

22. The non-transitory computer-readable medium of claim 21 , wherein the utterance is a phrase independent utterance.

23. The non-transitory computer-readable medium of claim 21 , wherein the utterance is a phrase dependent utterance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2016
From: BISWAL, NARAYAN; CILINGIR, GOKCEN; DAS, BARNAN
To: INTEL CORPORATION
Reel/Frame 038556/0069 →
Continuity (1)
Related Publication 20170287489A1 · Oct 5, 2017