IP Library Granted Patent US 10,446,136
Granted Patent B2
US 10,446,136 · App. 15/592,222 · Granted Oct 15, 2019

Accent invariant speech recognition

Inventors: Ron Fridental (Shoham, IL); Ilya Blayvas (Rehovot, IL); Pavel Nosko (Yavne, IL)
Assignee: ANTS TECHNOLOGY (HK) LIMITED
G10L15/065G10L15/063G10L15/10G10L15/14G10L15/16G10L2015/025G10L2015/0631
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,446,136
App. No.
15/592,222
Granted
Oct 15, 2019
Kind
B2
Abstract

A system and method for accent invariant speech recognition comprising: maintaining a database scoring a set of language units in a given language, and for each of the language units, scoring audio samples of pronunciation variations of the language unit pronounced by a plurality of speakers; extracting and storing m the database a feature vector for locating each of the audio samples in a feature space; identifying pronunciation variation distances, which are distances between locations of audio samples of the same language unit in the feature space, and inter-unit distances, which are distances between locations of audio samples of different language units in the feature space; calculating a transformation applicable on the feature space to reduce the pronunciation variation distances relative to the inter-unit distances; and based on the calculated transformation, training a processor to classify as a same language unit pronunciation variations of the same language unit.

Claims (26)

1. A method for accent invariant speech recognition comprising:

maintaining a database for storing a set of language units in a given language, wherein for each language unit, storing audio samples of pronunciation variations of the language unit pronounced by a plurality of speakers;

extracting and storing in the database a feature vector for locating each of the audio samples in a feature space;

identifying two types of distances: (i) pronunciation variation, which are distances between locations of audio samples of the same language unit with different pronunciations, in the feature space; and (ii) inter-unit distances, which are distances between locations of audio samples of different language units in the feature space;

calculating a transformation applicable on the feature space to reduce the pronunciation variation distances relative to the inter-unit distancesm, the transformation is configured to make various pronunciation variations of the same language unit indistinguishable by a classification processor;

when receiving an input audio:

transforming the received signal to an accent-invariant audio signal by applying the calculated transformation on the input audio signal, wherein language units included in the accent-invariant audio signal are indistinguishable by the classification processor from other pronunciation variations of the same language units; and

recognizing a language unit in said input audio signal, by applying classification by said classification processor.

2. The method of claim 1 , wherein the language units are words or phonemes.

3. The method of claim 1 , wherein recognizing a language unit comprises adjusting classification based on language statistics.

4. The method of claim 1 wherein said method further comprises applying the calculated transformation to the samples of pronunciation variations stored in the database.

5. The method of claim 1 , wherein said calculated transformation comprises a Linear Discriminant Analysis (LDA) transformation.

6. The method of claim 1 , wherein said calculated transformation is performed by an appropriately trained neural network.

7. The method of claim 1 , wherein the stored audio samples are of pronunciation variations of the language unit pronounced by a plurality of speakers of different ethnic groups.

8. A method for accent invariant speech recognition comprising:

maintaining a database storing a set of language units in a given language, and for each language unit, storing audio samples of pronunciation variations of the language unit pronounced by a plurality of speakers with known accents,

wherein the audio samples are indexed according to the language unit and accent integrated in the audio sample;

for each known accent:

identifying two types of distances: (i) pronunciation variation, which are distances between locations of audio samples of the same language unit with different pronunciations, in the feature space; and (ii) inter-unit distances, which are distances between locations of audio samples of different language units in the feature space;

calculating a transformation applicable on the feature space to reduce the pronunciation variation distances relative to the inter-unit distances, the transformation is configured to make various pronunciation variations of the same language unit and accent indistinguishable by a classification processor; and

when receiving an input audio signal,

in case accent of the received audio signal s recognized, applying classification for the recognized accent by said processor, thus recognizing a language unit in said input audio signal; and

in case an accent of the received audio signal is not recognized:

applying a separate classification for each of the known accents, thus recognizing a language unit in said input audio signal for each of the known accents; and

selecting the most probable recognized language unit,

wherein applying classification for the recognized accent comprises transforming the received signal by applying on the input audio signal the corresponding calculated transformation, wherein language units included in the transformed audio signal are indistinguishable by the classification processor from other pronunciation variations of the same language units and accent.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Nov 5, 2024
From: EAST WEST BANK
To: KAMI VISION INCORPORATED
Reel/Frame 070792/0551 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Mar 28, 2022
From: KAMI VISION INCORPORATED
To: EAST WEST BANK
Reel/Frame 059512/0101 →
NUNC PRO TUNC ASSIGNMENT Recorded Mar 8, 2022
From: ANTS TECHNOLOGY (HK) LIMITED
To: SHANGHAI XIAOYI TECHNOLOGY CO., LTD.
Reel/Frame 059194/0736 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: SHANGHAI XIAOYI TECHNOLOGY CO., LTD.
To: KAMI VISION INC.
Reel/Frame 059347/0649 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2017
From: RIDENTAL, RON; BLAYVAS, ILYA; NOSKO, PAVEL
To: ANTS TECHNOLOGY (HK) LIMITED
Reel/Frame 042333/0168 →
Continuity (1)
Related Publication 20180330719A1 · Nov 15, 2018