IP Library › Granted Patent US 11,551,695
Granted Patent B1
US 11,551,695 · App. 15/931,455 · Granted Jan 10, 2023

Model training system for custom speech-to-text models

Inventors: Vivek Govindan (Redmond, WA); Varun Sembium Varadarajan (Bothell, WA); Christian Egon Berkhoff Dossow (Lake Forest Park, WA); Himalay Mohanlal Joriwal (Seattle, WA); Sai Madhuri Bhavirisetty (Bellevue, WA); Abhinav Kumar (Bellevue, WA); Orestis Lykouropoulos (Seattle, WA); Akshay Nalwaya (Seattle, WA); Rahul Gupta (Seattle, WA); Sravan Babu Bodapati (Redmond, WA); Liangwei Guo (Seattle, WA); Julian E. S. Salazar (San Francisco, CA); Yibin Wang (Seattle, WA); K P N V D S Siva Rama (Seattle, WA); Calvin Xuan Li (Seattle, WA); Mohit Narendra Gupta (Seattle, WA); Asem Rustum (Sammamish, WA); Katrin Kirchhoff (Seattle, WA); Pu Zhao (Lynnwood, WA)
Assignee: Amazon Technologies, Inc.
G10L15/26G10L15/063G10L15/07G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,695
App. No.
15/931,455
Filed
May 13, 2020
Granted
Jan 10, 2023
Kind
B1
Examiner
WONG, LINDA
Art Unit
2655
USPC
704/235
Abstract

A transcription service may receive a request from a developer to build a custom speech-to-text model for a specific domain of speech. The custom speech-to-text model for the specific domain may replace a general speech-to-text model or add to a set of one or more speech-to-text models available for transcribing speech. The transcription service may receive a training data and instructions representing tasks. The transcription service may determine respective schedules for executing the instructions based at least in part on dependencies between the tasks. The transcription service may execute the instructions according to the respective schedules to train a speech-to-text model for a specific domain using the training data set. The transcription service may deploy the trained speech-to-text model as part of a network-accessible service for an end user to convert audio in the specific domain into texts.

Claims (46)

1. A system, comprising:

one or more computing devices implementing a network-accessible transcription service of a provider network, wherein to implement the transcription service, the computing devices are configured to:

receive, via a network-accessible interface, a request that specifies a domain of speech and requires to generate a speech-to-text model for the specified domain to be used for transcribing audio data associated with the specified domain;

in response to the request, train a speech-to-text model according to a workflow of tasks and a training data set corresponding to the specified domain to generate a trained speech-to-text model for the specified domain;

receive, via the network-accessible interface, a transcription request to transcribe particular audio data of the specified domain of speech;

identify the trained speech-to-text model for the specified domain of speech instead of a general speech-to-text model to transcribe the particular audio data; and

create a transcription of the particular audio data using the trained speech-to-text model for the specified domain.

2. The system of claim 1 , wherein to generate the trained speech-to-text model for the specified domain, the computing devices are configured to:

determine respective schedules for executing instructions for the tasks based at least in part on the workflow; and

execute the instructions according to the respective schedules to train the speech-to-text model using the training data set to generate the trained speech-to-text model for the specified domain.

3. The system of claim 1 , wherein the workflow represents dependencies between the tasks based at least in part on individual input and output of the tasks.

4. The system of claim 1 , wherein the trained speech-to-text model is stored in one or more data stores that are implemented as part of the transcription service or a data storage service offered by the provider network.

5. A method, comprising:

creating a speech-to-text model for a specific domain of speech, in response to a first request received via an interface of a transcription service of a provider network that is implemented using one or more computing devices, to be added to a set of one or more speech-to-text models available for transcribing speech, wherein the first request specifies the specific domain of speech and requires the creation of the speech-to-text model for the specific domain;

receiving, via the interface, a second request to transcribe an audio file that identifies the speech-to-text model for the specific domain of speech; and

responsive to the second request, selecting the speech-to-text model for the specific domain of speech from the set of speech-to-text models to transcribe the audio file to create a transcription of the audio file.

6. The method of claim 5 , wherein creating the speech-to-text model for the specific domain of speech comprises:

determining respective schedules for executing a plurality of instructions for a plurality of tasks based at least in part on dependencies between the plurality of tasks; and

executing the plurality of instructions according to the respective schedules to train a speech-to-text model using a training data set to create the speech-to-text model for the specific domain.

7. The method of claim 6 , wherein the dependencies between the plurality of tasks include dependencies between individual input and output of the plurality of tasks.

8. The method of claim 6 , wherein executing the plurality of instructions comprises:

obtaining one batch of the plurality of instructions until after execution of another batch of the plurality of instructions; and

executing the batch of the plurality of instructions.

9. The method of claim 6 , wherein executing the plurality of instructions comprises:

provisioning one or more computing resources for respective ones of the plurality of instructions; and

executing the respective ones of the plurality of instructions using the provisioned one or more computing resources.

10. The method of claim 5 , wherein creating the speech-to-text model for the specific domain of speech comprises generating one or more logs associated with training of the speech-to-text model.

11. The method of claim 10 , further comprising:

in response to a request, generating a visual display based at least in part on the one or more logs associated with training of the speech-to-text model.

12. The method of claim 5 , wherein the speech-to-text model for the specific domain is stored in one or more data stores that are implemented as part of the transcription service or a data storage service offered by the provider network.

13. One or more non-transitory, computer readable media, storing program instructions that when executed on or across one or more computing devices, cause the one or more computing devices to implement:

creating a speech-to-text model for a specific domain of speech, in response to a first request received via an interface of a transcription service of a provider network that is implemented using one or more computing devices, to be added to a set of one or more speech-to-text models available for transcribing speech, wherein the first request specifies the specific domain of speech and requires the creation of the speech-to-text model for the specific domain;

receiving, via the interface, a second request to transcribe an audio file that identifies the speech-to-text model for the specific domain of speech; and

responsive to the second request, selecting the speech-to-text model for the specific domain of speech from the set of speech-to-text models to transcribe the audio file to create a transcription of the audio file.

14. The one or more non-transitory, computer readable media of claim 13 , wherein, in creating the speech-to-text model for the specific domain of speech, the program instructions cause the one or more computing devices to implement:

determining respective schedules for executing a plurality of instructions for a plurality of tasks based at least in part on dependencies between the plurality of tasks; and

executing the plurality of instructions according to the respective schedules to train a speech-to-text model using a training data set to create the speech-to-text model for the specific domain.

15. The one or more non-transitory, computer readable media of claim 14 , wherein the dependencies between the plurality of tasks include dependencies between individual input and output of the plurality of tasks.

16. The one or more non-transitory, computer readable media of claim 14 , wherein, in executing the plurality of instructions, the program instructions cause the one or more computing devices to implement:

obtaining one batch of the plurality of instructions until after execution of another batch of the plurality of instructions; and

executing the batch of the plurality of instructions.

17. The one or more non-transitory, computer readable media of claim 14 , wherein the plurality of instructions comprise code in respective container images corresponding to the plurality of instructions.

18. The one or more non-transitory, computer readable media of claim 13 , wherein, in creating the speech-to-text model for the specific domain of speech, the program instructions cause the one or more computing devices to implement generating one or more logs associated with training of the speech-to-text model.

19. The one or more non-transitory, computer readable media of claim 16 , storing further program instructions that when executed on or across the one or more computing devices, cause the one or more computing devices to further implement:

in response to a request, generating a visual display based at least in part on the one or more logs associated with training of the speech-to-text model.

20. The one or more non-transitory, computer readable media of claim 13 , wherein the speech-to-text model for the specific domain is stored in one or more data stores that are implemented as part of the transcription service or a data storage service offered by the provider network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2022
From: GOVINDAN, VIVEK; SEMBIUM VARADARAJAN, VARUN; BERKHOFF DOSSOW, CHRISTIAN EGON; JORIWAL, HIMALAY MOHANLAL; BHAVIRISETTY, SAI MADHURI; KUMAR, ABHINAV; LYKOUROPOULOS, ORESTIS; NALWAYA, AKSHAY; BODAPATI, SRAVAN BABU; GUO, LIANGWEI; SALAZAR, JULIAN E. S.; WANG, YIBIN; K P N V D S SIVA RAMA, .; GUPTA, MOHIT NARENDRA; RUSTUM, ASEM; KIRCHHOFF, KATRIN; ZHAO, PU
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 062045/0707 →
Cited By (2)
US 12,609,102 US 12,711,961