IP Library Granted Patent US 9,053,704
Granted Patent B2
US 9,053,704 · App. 14/330,739 · Granted Jun 9, 2015

System and method for standardized speech recognition infrastructure

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,053,704
App. No.
14/330,739
Granted
Jun 9, 2015
Kind
B2
Abstract

Disclosed herein are systems, methods, and computer-readable storage media for selecting a speech recognition model in a standardized speech recognition infrastructure. The system receives speech from a user, and if a user-specific supervised speech model associated with the user is available, retrieves the supervised speech model. If the user-specific supervised speech model is unavailable and if an unsupervised speech model is available, the system retrieves the unsupervised speech model. If the user-specific supervised speech model and the unsupervised speech model are unavailable, the system retrieves a generic speech model associated with the user. Next the system recognizes the received speech from the user with the retrieved model. In one embodiment, the system trains a speech recognition model in a standardized speech recognition infrastructure. In another embodiment, the system handshakes with a remote application in a standardized speech recognition infrastructure.

Claims (55)

1. A computer-implemented method comprising:

initiating a voice call with a remote application;

determining, by processor-executable instructions executed by a processor, whether the remote application can apply standardized speech recognition models;

when the remote application can apply the standardized speech recognition models:

transmitting a speaker specific model to the remote application; and

the processor-executable instructions when executed by the processor instructing the remote application to recognize speech with the transmitted speaker specific model; and

when the remote application can not apply standardized speech recognition models:

the processor-executable instructions when executed by the processor determining whether the remote application can apply transformations of a standard speech recognition model.

2. The computer-implemented method of claim 1 , further comprising:

when the remote application can apply transformations of the standard speech recognition model:

transmitting transformations of a generic model to the remote application; and

instructing the remote application to recognize speech based on the transmitted transformations; and

when the remote application can not apply transformations of the standard speech recognition model, instructing the remote application to recognize speech with the generic model.

3. The computer-implemented method of claim 1 , wherein the speaker specific model is generated on a first device and the transformations are generated on a second device distinct from the first device.

4. The computer-implemented method of claim 1 , further comprising instructing the remote application to perform self-adaptation based on one of the transmitted speaker specific model and the transformations.

5. The computer-implemented method of claim 1 , wherein the transmitting occurs using an edge device in a communications network.

6. The computer-implemented method of claim 1 , wherein the remote application is a game, and wherein the game identifies voice commands in recognized speech and controls elements of the game based on the voice commands.

7. The computer-implemented method of claim 1 , wherein the generic model and the speaker specific model are publicly available.

8. The computer-implemented method of claim 1 , wherein a network client authenticates the speech before instructing the remote application to recognize the speech.

9. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

initiating a voice call with a remote application;

determining if the remote application can apply standardized speech recognition models;

when the remote application can apply the standardized speech recognition models:

transmitting a speaker specific model to the remote application, and

instructing the remote application to recognize speech with the transmitted model;

when the remote application can not apply standardized speech recognition models:

determining if the remote application can apply transformations of a standard speech recognition model.

10. The system of claim 9 , the computer-readable storage medium having additional instructions stored which result in operations comprising:

when the remote application can apply transformations of the standard speech recognition model:

transmitting transformations of a generic model to the remote application, and

instructing the remote application to recognize speech based on the transmitted transformations; and

when the remote application can not apply transformations of the standard speech recognition model, instructing the remote application to recognize speech with the generic model.

11. The system of claim 9 , wherein the speaker specific model is generated on a first device and the transformations are generated on a second device distinct from the first device.

12. The system of claim 9 , the computer-readable storage medium having additional instructions stored which result in operations comprising instructing the remote application to perform self-adaptation based on one of the transmitted speaker specific model and the transformations.

13. The system of claim 9 , wherein the transmitting occurs using an edge device in a communications network.

14. The system of claim 9 , wherein the remote application is a game, and wherein the game identifies voice commands in recognized speech and controls elements of the game based on the voice commands.

15. The system of claim 9 , wherein the generic model and the speaker specific model are publicly available.

16. The system of claim 9 , wherein a network client authenticates the speech before instructing the remote application to recognize the speech.

17. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

initiating a voice call with a remote application;

determining if the remote application can apply standardized speech recognition models;

when the remote application can apply the standardized speech recognition models:

transmitting a speaker specific model to the remote application, and

instructing the remote application to recognize speech with the transmitted model; and

when the remote application can not apply standardized speech recognition models:

determining if the remote application can apply transformations of a standard speech recognition model.

18. The computer-readable storage device of claim 17 , having additional instructions stored which result in operations comprising:

when the remote application can apply transformations of the standard speech recognition model:

transmitting transformations of a generic model to the remote application, and

instructing the remote application to recognize speech based on the transmitted transformations; and

when the remote application can not apply transformations of the standard speech recognition model, instructing the remote application to recognize speech with the generic model.

19. The computer-readable storage device of claim 17 , wherein the speaker specific model is generated on a first device and the transformations are generated on a second device distinct from the first device.

20. The computer-readable storage device of claim 17 , having additional instructions stored which result in operations comprising instructing the remote application to perform self-adaptation based on one of the transmitted speaker specific model and the transformations.

Assignments (18)
RELEASE OF SECURITY INTEREST Recorded Sep 4, 2025
From: RUNWAY GROWTH FINANCE CORP., AS AGENT
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 072802/0931 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 043039/0808 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060557/0636 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 27, 2022
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 060445/0733 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 29, 2017
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 043039/0808 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2014
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034462/0764 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2014
From: LJOLJE, ANDREJ; RENGER, BERNARD S.; TISCHER, STEVEN NEIL
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 033686/0758 →