IP Library › Granted Patent US 11,315,545
Granted Patent B2
US 11,315,545 · App. 16/924,564 · Granted Apr 26, 2022

System and method for language identification in audio data

Inventor: Jonathan C. Wintrode (Annapolis, MD)
Assignee: RAYTHEON APPLIED SIGNAL TECHNOLOGY, INC.
G10L15/005G10L15/08G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,315,545
App. No.
16/924,564
Granted
Apr 26, 2022
Kind
B2
Abstract

A system for identifying a language in audio data includes a feature extraction module for receiving an unknown input audio data stream and dividing the unknown input audio data stream into segments. A similarity module receives the segments and receives known-language audio data models for known languages. For each segment, the similarity module performs comparisons between the segment and the known-language audio data models and generates probability values representative of the probabilities that the segment includes audio data of the known languages. A processor receives the probability values for each segment and computes an entropy value for the probabilities for each segment. If the entropy value for a segment is less than the entropy value for a previous segment, the processor terminates the comparisons prior to completing comparisons for all segments.

Claims (21)

1. A system for identifying a language in audio data, comprising:

a feature extraction module for receiving an unknown input audio data stream and dividing the unknown input audio data stream into a plurality of segments of unknown input audio data;

a similarity module for receiving the plurality of segments of the unknown input audio data and for receiving a plurality of known-language audio data models for a respective plurality of known languages, for each segment of the unknown input audio data, the similarity module performing comparisons between the segment of unknown input audio data and the plurality of known-language audio data models and generating a respective plurality of probability values representative of the probabilities that the segment includes audio data of the known languages; and

a processor for receiving the plurality of probability values for each segment and computing an entropy value for the probabilities for each segment, and identifying the language in the audio data based on the generated probabilities; wherein

if the entropy value for a segment is less than the entropy value for a previous segment, the processor terminates the comparisons prior to completing comparisons for all segments of the unknown input audio data.

2. The system of claim 1 , wherein each segment of unknown audio data comprises an unknown data vector comprising a plurality of data values associated with the segment of unknown input audio data; and each known-language audio data model comprises a known data vector comprising a plurality of data values associate with the known-language audio data model.

3. The system of claim 1 , wherein the feature extraction module comprises a deep neural network.

4. The system of claim 1 , wherein the similarity module performs a probabilistic linear discriminant analysis (PLDA) in generating the plurality of probability values.

5. The system of claim 1 , wherein extents of each segment are defined by a time duration.

6. The system of claim 1 , wherein extents of each segment are defined by a quantity of data in the segment.

7. A method for identifying a language in audio data, comprising:

receiving, at a feature extraction module, an unknown input audio data stream and dividing the unknown input audio data stream into a plurality of segments of unknown input audio data;

receiving, at a similarity module, the plurality of segments of the unknown input audio data and receiving, at the similarity module, a plurality of known-language audio data models for a respective plurality of known languages, for each segment of the unknown input audio data, the similarity module performing comparisons between the segment of unknown input audio data and the plurality of known-language audio data models and generating a respective plurality of probability values representative of the probabilities that the segment includes audio data of the known languages;

computing an entropy value for the probabilities for each segment;

identifying the language in the audio data based on the generated probabilities; and

if the entropy value for a segment is less than the entropy value for a precious segment, terminating the comparisons prior to completing comparisons for all segments of the unknown input audio data.

8. The method of claim 7 , wherein each segment of unknown audio data comprises an unknown data vector comprising a plurality of data values associated with the segment of unknown input audio data; and each known-language audio data model comprises a known data vector comprising a plurality of data values associate with the known-language audio data model.

9. The method of claim 7 , wherein the feature extraction module comprises a deep neural network.

10. The method of claim 7 , wherein the similarity module performs a probabilistic linear discriminant analysis (PLDA) in generating the plurality of probability values.

11. The method of claim 7 , wherein extents of each segment are defined by a time duration.

12. The method of claim 7 , wherein extents of each segment are defined by a quantity of data in the segment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2021
From: WINTRODE, JONATHAN C.
To: RAYTHEON APPLIED SIGNAL TECHNOLOGY, INC.
Reel/Frame 056487/0591 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2020
From: WINTRODE, JONATHAN C.
To: RAYTHEON COMPANY
Reel/Frame 053195/0695 →
Continuity (1)
Related Publication 20220013107A1 · Jan 13, 2022
Cited By (2)
US 12,562,180 US 12,651,598