Automated context-specific speech-to-text transcriptions
Disclosed are various approaches for generating a text transcript of a soundtrack. The soundtrack can correspond to an event in a conferencing service. Language models can be trained on data that is specific to organizations, users within the organization, and metadata associated with an agenda for the event. The metadata can include texts, attachments, and other data associated with the event. The language models can be arranged into a convolutional neural network and output a text transcript. The text transcript can be used to retrain the language models for subsequent use.
1. A system, comprising:
a computing device comprising at least one processor and at least one memory; and
machine-readable instructions stored in the at least one memory, wherein the instructions, when executed by the at least one processor, cause the computing device to at least:
identify an event in at least one of a user calendar or a conferencing service;
obtain a soundtrack corresponding to a video input, the video input obtained from the conferencing service, the video input further being associated with a viewer displaying the video input to a user;
identify an organization-specific language model associated with at least one user invited to the event;
identify a user-specific language model associated with the at least one user invited to the event;
identify an agenda-specific language model associated with the event, the agenda-specific language model generated using metadata associated with the event and wherein the organization-specific language model, the user-specific language model, and the agenda-specific language model comprise a convolutional neural network; and
generate a text transcript based upon the soundtrack from the video input using the convolutional neural network.
2. The system of claim 1 , wherein the organization-specific language model further comprises organizational-specific classifiers that specify how to generate the text transcript based upon organization wide rules.
3. The system of claim 1 , wherein the user-specific language model further comprises user-specific classifiers that specify how to generate the text transcript based upon user-specified rules.
4. The system of claim 1 , wherein the metadata associated with the event comprises an event description or at least one document associated with the event.
5. The system of claim 1 , wherein the organization-specific language model comprises a head of the convolutional neural network.
6. The system of claim 1 , wherein the organization-specific language model is trained to expand on a plurality of organization-specific acronyms that are inserted into the text transcript.
7. The system of claim 1 , wherein the convolutional neural network is configurable to generate an abbreviated text transcript or a verbose text transcript.
8. A non-transitory computer-readable medium comprising machine-readable instructions, wherein the instructions, when executed by at least one processor, cause a computing device to at least:
identify an event in at least one of a user calendar or a conferencing service;
obtain a soundtrack corresponding to a video input, the video input obtained from the conferencing service, the video input further being associated with a viewer displaying the video input to a user;
identify an organization-specific language model associated with at least one user invited to the event;
identify a user-specific language model associated with the at least one user invited to the event;
identify an agenda-specific language model associated with the event, the agenda-specific language model generated using metadata associated with the event and wherein the organization-specific language model, the user-specific language model, and the agenda-specific language model comprise a convolutional neural network; and
generate a text transcript based upon the soundtrack from the video input using the convolutional neural network.
9. The non-transitory computer-readable medium of claim 8 , wherein the organization-specific language model further comprises organizational-specific classifiers that specify how to generate the text transcript based upon organization wide rules.
10. The non-transitory computer-readable medium of claim 8 , wherein the user-specific language model further comprises user-specific classifiers that specify how to generate the text transcript based upon user-specified rules.
11. The non-transitory computer-readable medium of claim 8 , wherein the metadata associated with the event comprises an event description or at least one document associated with the event.
12. The non-transitory computer-readable medium of claim 8 , wherein the organization-specific language model comprises a head of the convolutional neural network.
13. The non-transitory computer-readable medium of claim 8 , wherein the organization-specific language model is trained to expand on a plurality of organization-specific acronyms that are inserted into the text transcript.
14. The non-transitory computer-readable medium of claim 8 , wherein the convolutional neural network is configurable to generate an abbreviated text transcript or a verbose text transcript.
15. A method comprising:
identifying an event in at least one of a user calendar or a conferencing service;
obtaining a soundtrack corresponding to a video input, the video input obtained from the conferencing service, the video input further being associated with a viewer displaying the video input to a user;
identifying an organization-specific language model associated with at least one user invited to the event;
identifying a user-specific language model associated with the at least one user invited to the event;
identifying an agenda-specific language model associated with the event, the agenda-specific language model generated using metadata associated with the event and wherein the organization-specific language model, the user-specific language model, and the agenda-specific language model comprise a convolutional neural network; and
generating a text transcript based upon the soundtrack from the video input using the convolutional neural network.
16. The method of claim 15 , wherein the organization-specific language model further comprises organizational-specific classifiers that specify how to generate the text transcript based upon organization wide rules.
17. The method of claim 15 , wherein the user-specific language model further comprises user-specific classifiers that specify how to generate the text transcript based upon user-specified rules.
18. The method of claim 15 , wherein the metadata associated with the event comprises an event description or at least one document associated with the event.
19. The method of claim 15 , wherein the organization-specific language model comprises a head of the convolutional neural network.
20. The method of claim 15 , wherein the organization-specific language model is trained to expand on a plurality of organization-specific acronyms that are inserted into the text transcript.