IP Library Patent Application 16543608
Patent Application
App. No. 16/543,608

SYSTEM AND METHOD FOR PERFORMING MULTI-MODEL AUTOMATIC SPEECH RECOGNITION IN CHALLENGING ACOUSTIC ENVIRONMENTS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/543,608
Abstract

A speech recognition method includes: providing a system having a local computational device, the local computational device having a microphone, processing circuitry, and a non-transitory computer-readable medium; recording a raw audio waveform utilizing the microphone; determining a background noise condition for the raw audio waveform; comparing the background noise condition to a plurality of linguistic models having associated background noise conditions; determining a nearest match between the background noise condition of the raw audio waveform and the associated background noise condition of at least one of the plurality of linguistic models; and performing an automatic speech recognition (ASR) function between the raw audio waveform and the linguistic model having the matching associated background noise condition.

Claims (57)

1 . A speech recognition method comprising:

providing a system having a local computational device, the local computational device having a microphone, processing circuitry, and a non-transitory computer-readable medium;

recording a raw audio waveform utilizing the microphone;

determining a background noise condition for the raw audio waveform;

comparing the background noise condition to a plurality of linguistic models having associated background noise conditions;

determining a nearest match between the background noise condition of the raw audio waveform and the associated background noise condition of at least one of the plurality of linguistic models; and

performing an automatic speech recognition (ASR) function between the raw audio waveform and the linguistic model having the matching associated background noise condition.

2 . The speech recognition method of claim 1 , further comprising:

providing a remote server;

providing a database containing a plurality of linguistic models and associated background noise conditions for each linguistic model;

providing a computerized neural network on the remote server;

wherein the computerized neural network is configured to determine the nearest match between the background noise condition of the raw audio waveform and the associated background noise condition of a particular linguistic model.

3 . The speech recognition method of claim 1 , further comprising:

providing a database containing a plurality of linguistic models and associated background noise conditions for each linguistic model;

providing a computerized neural network on the local device;

wherein the computerized neural network is configured to determine the nearest match between the background noise condition of the raw audio waveform and the associated background noise condition of a particular linguistic model.

4 . The speech recognition method of claim 3 , wherein the computerized neural network is compressed.

5 . The speech recognition method of claim 1 , further comprising:

extracting textual characters representing speech within the raw audio waveform.

6 . The speech recognition method of claim 5 , further comprising:

tracking user interactions with the local computational device;

determining corrections to the textual characters extracted from the raw audio waveform.

7 . The speech recognition method of claim 4 , further comprising:

generating a new linguistic model with an associated background condition based on an original base truth from an original linguistic model and creating a new linguistic model based on an extracted background noise condition;

inserting the new linguistic model into the plurality of linguistic models contained on the database for future consideration by the computerized neural network in future match determination functions.

8 . The speech recognition method of claim 1 , wherein each of the plurality of linguistic models are representative of a single language model recorded in a plurality of associated noise condition, wherein the only variation between each linguistic model is their particular associated noise conditions.

9 . The speech recognition method of claim 1 , wherein the plurality of linguistic models include a plurality of language models, each being recorded in a plurality of associated noise condition, wherein each linguistic model can vary with regard to particular associated noise conditions as well as an underlying language represented thereby.

10 . A speech recognition system, the system comprising:

a local computational system, the local computational system further comprising:

processing circuitry;

a microphone operatively connected to the processing circuitry;

a non-transitory computer-readable media being operatively connected to the processing circuitry;

a remote server configured to receive recorded wavelengths from the local computational system; the remote server having one or more computerized neural networks, wherein the computerized neural networks of the remote server are trained on a plurality of acoustic models, wherein each of the plurality of acoustic models represent a particular linguistic dataset recorded in one or more associated noise predetermined signal to noise ratios;

wherein the non-transitory computer-readable media contains instructions for the processing circuitry to perform:

utilizing the microphone to record raw audio waveforms from an ambient atmosphere;

transmitting the recorded raw audio waveforms to the remote server; and

wherein the computerized neural network is configured to determine the nearest match between the background noise condition of the raw audio waveform and the associated background noise condition of a particular linguistic model.

11 . The speech recognition system of claim 10 , wherein the computerized neural network is further configured to extract textual characters representing speech within the raw audio waveform.

12 . The speech recognition system of claim 11 , wherein the computerized neural network is further configured to track user interactions with the local computational device and determine corrections to the textual characters extracted from the raw audio waveform.

13 . The speech recognition system of claim 12 , wherein the computerized neural network is further configured to generate a new linguistic model with an associated background condition based on an original base truth from an original linguistic model and creating a new linguistic model based on an extracted background noise condition; and insert the new linguistic model into the plurality of linguistic models contained on the database for future consideration by the computerized neural network in future match determination functions.

14 . The speech recognition system of claim 10 , wherein each of the plurality of linguistic models are representative of a single language model recorded in a plurality of associated noise condition, wherein the only variation between each linguistic model is their particular associated noise conditions.

15 . The speech recognition system of claim 10 , wherein the plurality of linguistic models include a plurality of language models, each being recorded in a plurality of associated noise condition, wherein each linguistic model can vary with regard to particular associated noise conditions as well as an underlying language represented thereby.

16 . A robotic apparatus comprising a speech recognition system, the system comprising:

a local computational system, the local computational system further comprising:

processing circuitry;

a microphone operatively connected to the processing circuitry;

a non-transitory computer-readable media being operatively connected to the processing circuitry;

one or more computerized neural networks;

wherein the non-transitory computer-readable media contains instructions for the processing circuitry to perform:

utilizing the microphone to record raw audio waveforms from an ambient atmosphere; and

wherein at least one computerized neural network is configured to wherein the computerized neural network is configured to determine the nearest match between the background noise condition of the raw audio waveform and the associated background noise condition of a particular linguistic model.

17 . The robotic apparatus of claim 16 , wherein the computerized neural network is further configured to extract textual characters representing speech within the raw audio waveform.

18 . The robotic apparatus system of claim 17 ,

wherein the computerized neural network is further configured to track user interactions with the local computational device and determine corrections to the textual characters extracted from the raw audio waveform, and

wherein the computerized neural network is further configured to generate a new linguistic model with an associated background condition based on an original base truth from an original linguistic model and creating a new linguistic model based on an extracted background noise condition; and insert the new linguistic model into the plurality of linguistic models contained on the database for future consideration by the computerized neural network in future match determination functions.

19 . The robotic apparatus system of claim 16 , wherein each of the plurality of linguistic models are representative of a single language model recorded in a plurality of associated noise condition, wherein the only variation between each linguistic model is their particular associated noise conditions.

20 . The robotic apparatus system of claim 16 , wherein the plurality of linguistic models include a plurality of language models, each being recorded in a plurality of associated noise condition, wherein each linguistic model can vary with regard to particular associated noise conditions as well as an underlying language represented thereby.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE'S NAME INSIDE THE ASSIGNMENT DOCUMENT AND ON THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 055556 FRAME: 0131. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 17, 2021
From: CLOUDMINDS TECHNOLOGY, INC.
To: CLOUDMINDS ROBOTICS CO., LTD.
Reel/Frame 056047/0834 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2021
From: CLOUDMINDS TECHNOLOGY, INC.
To: DATHA ROBOT CO., LTD.
Reel/Frame 055556/0131 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2019
From: JANKOWSKI, CHARLES ROBERT, JR.; LIN, RUIXI
To: CLOUDMINDS TECHNOLOGY, INC.
Reel/Frame 050083/0512 →