IP Library Granted Patent US 11,423,889
Granted Patent B2
US 11,423,889 · App. 16/583,688 · Granted Aug 23, 2022

Systems and methods for recognizing a speech of a speaker

Inventor: Ilya Vladimirovish Mikhailov (Saint Petersburg, RU)
Assignee: RingCentral, Inc.
G10L15/22G06N3/08G10L13/00G10L15/04G10L15/16G10L15/24G10L15/30G10L25/90G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,889
App. No.
16/583,688
Granted
Aug 23, 2022
Kind
B2
Abstract

Systems, methods, and computer readable media comprising instructions executable by a processor, for recognizing speech within a received audio signal segment the audio signal to isolate the speech based on a speaker audio profile, determine from the audio signal a command, a first score reflecting confidence in determining the command, and a second score reflecting a potential error in determining the command, and cause the command to be executed if the first score is above a first threshold value and the second score is below a second threshold value.

Claims (58)

1. A method for recognizing speech within a received audio signal, the method comprising:

segmenting, by a processor of a computing device by using a computer-based neural network model, an audio signal to isolate a speech based on a speaker audio profile;

determining, by the processor, from the audio signal:

a command,

a first score reflecting a percentage of confidence in determining the command, and

a second score reflecting a percentage of importance of the command and consequences due to a potential error in determining the command; and

causing, by the processor, the command to be executed if the first score is above a first threshold value and the second score is below a second threshold value,

wherein the determining the first score is performed based on a frequency of using the command by the speaker.

2. The method of claim 1 , wherein the segmentation of the audio signal is based on one of an audio waveform, a cadence of the speech, a pitch of the speech, loudness of the speech, or vocabulary of the speech.

3. The method of claim 1 , further comprising:

receiving video data of the speaker producing the audio signal; and

performing a segmentation of the audio signal based on a correlation of the audio signal and the video data.

4. The method of claim 1 , further comprising:

receiving vibration data of the speaker producing the audio signal; and

performing a segmentation of the audio signal based on a correlation of the audio signal and the vibration data.

5. The method of claim 4 , wherein the vibration data is recorded by a wearable device.

6. The method of claim 1 , wherein the audio data is recorded by at least one microphone.

7. The method of claim 1 , further comprising interacting with the speaker via a synthesized speech.

8. The method of claim 7 , wherein the interacting includes prompting the speaker for a follow-up audio signal containing a command.

9. A system for recognizing a speech of a speaker comprising:

at least one memory device storing instructions; and

at least one processor configured to execute the instructions to perform operations comprising:

receiving an audio signal;

performing a segmentation of the audio signal by using a computer-based neural network model to isolate the speech of the speaker based on a speaker audio profile;

determining from the audio signal:

a command,

a first score reflecting a percentage of confidence in determining the command, and

a second score reflecting a percentage of importance of the command and consequences due to a potential error in determining the command; and

causing the command to be executed if the first score is above a first threshold value and the second score is below a second threshold value,

wherein the determining the first score is performed based on a frequency of using the command by the speaker.

10. The system of claim 9 , wherein the operations further comprise:

transmitting the audio signal to a remote server of a remote computing platform via a network; and

storing the audio signal in a remote database associated with the remote computing platform.

11. The system of claim 10 , wherein transmitting the audio signal is performed in a continuously operating mode.

12. The system of claim 9 , wherein the memory device and the processor may be parts of a mobile computing device comprising one of a wearable electronic device, a smartphone, a tablet, a camera, a laptop or a gaming console.

13. The system of claim 9 , wherein the operations further comprise generating a profile for the speaker at a remote computing platform, the profile comprising speaker related meta-data, audio signals associated with the speaker, and at least one computer-based model for performing the segmentation of the audio signal.

14. The system of claim 13 , wherein the profile further comprises a speech recognition model.

15. The system of claim 9 , wherein the operations further comprise interacting with the speaker via a synthesized speech.

16. A computing platform for recognizing speech comprising:

a server for receiving audio data via a network;

a database configured to store a profile for the speaker, the profile comprising speaker related meta-data, audio signals associated with the speaker, and at least one computer-based model for recognizing the speech of the speaker configured to:

receive an audio signal;

segment the audio signal by using a computer-based neural network model to isolate a speech based on a speaker audio profile;

determine from the audio signal:

a command,

a first score reflecting a percentage of confidence in determining the command, and

a second score reflecting a percentage of importance of the command and consequence due to a potential error in determining the command; and

cause the command to be executed if the first score is above a first threshold value and the second score is below a second threshold value,

wherein the determining the first score is performed based on a frequency of using the command by the speaker.

17. The computing platform of claim 16 , configured to:

receive a request from a computing device for a speaker whose speech requires recognition;

upload the at least one computer-based model to the computing device for recognizing the speech of the speaker if the speaker has the speaker profile stored in the database of the computing system.

18. The computing platform of claim 16 , configured to:

receive a request from a mobile device for selecting a plurality of speakers whose speech requires recognition;

select speakers that have the speaker profiles stored in the database of the computing system, resulting in selected speakers;

upload a plurality of computer-based models corresponding to the selected speakers to a computing device for recognizing the speech of the selected speakers; and

perform a segmentation of the audio signal to isolate at least one speech of at least one of the selected speakers resulting in at least one isolated speech.

19. The computer platform of claim 18 , for facilitating an audio conference for a plurality of speakers, the speakers being participants of the audio conference, further comprising an interface for a participant to select at least one isolated speech of at least one speakers.

Assignments (2)
SECURITY INTEREST Recorded Feb 14, 2023
From: RINGCENTRAL, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 062973/0194 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2019
From: MIKHAILOV, ILYA VLADIMIROVICH
To: RINGCENTRAL, INC.
Reel/Frame 050501/0192 →
Continuity (2)
Continuation PCTRU2018000906 · Dec 28, 2018
Related Publication 20200211544A1 · Jul 2, 2020