Automatic generation of standard operating procedures from multimedia content
A procedure generation system obtains multimedia content describing performance of a task and generates a procedure including content for guiding a user through performance of the task. The procedure generation system extracts audio data from the multimedia content and generates a transcription of the audio data through application of a trained model. The transcription includes text corresponding to the audio data and timestamps associated with different text. Based on the transcription, a trained model generates a set of steps, with each step including text corresponding to different time intervals. The procedure generation system identifies portions of the multimedia content corresponding to different steps based on the time intervals and associates identified portions of the multimedia content with corresponding steps to generate the procedure. This generates a procedure with various steps including text and a corresponding portion of the multimedia content.
1 . A method for generating a procedure describing performance of a task, the method comprising:
obtaining multimedia content of performance of the task, the multimedia content including video data and audio data comprising a description of performance of the task;
extracting the audio data from the multimedia content;
generating a transcription of the audio data, the transcription including text corresponding to portions of the audio data and timestamps associated with various text;
generating a set of steps from the transcription of the audio data by applying a trained model to the transcription, each step including a portion of the audio data corresponding to a time interval based on the timestamps;
identifying portions of the video data corresponding to different steps of the set from the multimedia content, an identified portion of the video data for a step including multimedia content occurring during the time interval corresponding to the step;
generating the procedure by associating one or more steps of the set with a corresponding identified portion of the video data for the one or more steps;
receiving, via a client device, an identification of a step of the procedure and an identification of a physical location in a local area, the physical location being associated with the step based on the identification of the step and the identification of the physical location, wherein the physical location is identified through an augmented reality interface of the client device that presents a view of the local area captured by a camera of the client device;
generating a virtual object associated with the physical location and the step; and
storing the procedure in a procedure store for subsequent retrieval such that the virtual object is automatically displayed during the subsequent retrieval based on a viewing device being in proximity to the physical location associated with the step, wherein the viewing device displays information corresponding to the step responsive to interaction with the virtual object via the viewing device.
2 . The method of claim 1 , wherein generating the set of steps from the transcription of the audio data by applying the trained model to the transcription comprises:
generating a prompt for a trained generative model that includes one or more formatting instructions and that includes the transcription having the text and timestamps corresponding to various text; and
applying the trained generative model to the prompt to generate the set of steps from the transcription based on the one or more instructions in the prompt.
3 . The method of claim 1 , wherein a formatting instruction identifies one or more selected from a group consisting of: a language for the steps, characteristics of text to remove from the transcription when generating a step, how to combine text in the step, timing information to include in the step, and any combination thereof.
4 . The method of claim 1 , wherein timestamps associated with various text comprise a timestamp associated with each individual word in the text.
5 . The method of claim 1 , wherein timestamps associated with various text comprise a timestamp associated with different groups of words in the text.
6 . The method of claim 1 , further comprising:
receiving a quiz generation request identifying the procedure;
generating a quiz comprising one or more questions about the procedure by applying a trained quiz generation model to the procedure; and
storing the quiz in the procedure store in association with the procedure.
7 . The method of claim 1 , wherein obtaining multimedia content of performance of the task comprises:
receiving multimedia content of the local area where the task is performed from the client device that captured the multimedia content during performance of the task.
8 . The method of claim 1 , wherein obtaining multimedia content of performance of the task comprises:
receiving an identifier of the multimedia content from the client device; and
retrieving stored multimedia content associated with the identifier.
9 . The method of claim 1 , wherein generating the procedure by associating one or more steps of the set with the corresponding identified portion of the video data for the one or more steps comprises:
storing an association between a point in an environment map of the local area in which the task is performed and a step in response to receiving information from a creating user via the client device identifying the physical location in the local area corresponding to the point.
10 . A non-transitory computer-readable storage medium storing instructions for generating a procedure describing performance of a task, the instructions when executed by one or more processors causing the one or more processors to perform steps comprising:
obtaining multimedia content of performance of the task, the multimedia content including video data and audio data comprising a description of performance of the task;
extracting the audio data from the multimedia content;
generating a transcription of the audio data, the transcription including text corresponding to portions of the audio data and timestamps associated with various text;
generating a set of steps from the transcription of the audio data by applying a trained model to the transcription, each step including a portion of the audio data corresponding to a time interval based on the timestamps;
identifying portions of the video data corresponding to different steps of the set from the multimedia content, an identified portion of the video data for a step including multimedia content occurring during the time interval corresponding to the step;
generating the procedure by associating one or more steps of the set with a corresponding identified portion of the video data for the one or more steps;
receiving, via a client device, an identification of a step of the procedure and an identification of a physical location in a local area, the physical location being associated with the step based on the identification of the step and the identification of the physical location, wherein the physical location is identified through an augmented reality interface of the client device that presents a view of the local area captured by a camera of the client device;
generating a virtual object associated with the physical location and the step; and
storing the procedure in a procedure store for subsequent retrieval such that the virtual object is automatically displayed during the subsequent retrieval based on a viewing device being in proximity to the physical location associated with the step, wherein the viewing device displays information corresponding to the step responsive to interaction with the virtual object via the viewing device.
11 . The non-transitory computer-readable storage medium of claim 10 , wherein a formatting instruction identifies one or more selected from a group consisting of: a language for the steps, characteristics of text to remove from the transcription when generating a step, how to combine text in the step, timing information to include in the step, and any combination thereof.
12 . The non-transitory computer-readable storage medium of claim 10 , wherein timestamps associated with various text comprise a timestamp associated with each individual word in the text.
13 . The non-transitory computer-readable storage medium of claim 10 , further storing instructions that, when executed by the one or more processors causing the one or more processors to perform steps comprising:
generating a quiz comprising one or more questions about the procedure by applying a trained quiz generation model to the procedure; and
storing the quiz in the procedure store in association with the procedure.
14 . The non-transitory computer-readable storage medium of claim 10 , wherein obtaining multimedia content of performance of the task comprises:
receiving multimedia content of the local area where the task is performed from the client device that captured the multimedia content during performance of the task.
15 . The non-transitory computer-readable storage medium of claim 10 , wherein generating the procedure by associating one or more steps of the set with the corresponding identified portion of the video data for the one or more steps comprises:
storing an association between a point in an environment map of the local area in which the task is performed and a step in response to receiving information from a creating user via the client device identifying the physical location in the local area corresponding to the point.
16 . A computer system comprising:
one or more processors; and
a non-transitory computer-readable storage medium storing instructions for generating a procedure describing performance of a task, the instructions when executed by the one or more processors causing the one or more processors to perform steps comprising:
obtaining multimedia content of performance of the task, the multimedia content including video data and audio data comprising a description of performance of the task;
extracting the audio data from the multimedia content;
generating a transcription of the audio data, the transcription including text corresponding to portions of the audio data and timestamps associated with various text;
generating a set of steps from the transcription of the audio data by applying a trained model to the transcription, each step including a portion of the audio data corresponding to a time interval based on the timestamps;
identifying portions of the video data corresponding to different steps of the set from the multimedia content, an identified portion of the video data for a step including multimedia content occurring during the time interval corresponding to the step;
generating the procedure by associating one or more steps of the set with a corresponding identified portion of the video data for the one or more steps;
receiving, via a client device, an identification of a step of the procedure and an identification of a physical location in a local area, the physical location being associated with the step based on the identification of the step and the identification of the physical location, wherein the physical location is identified through an augmented reality interface of the client device that presents a view of the local area captured by a camera of the client device;
generating a virtual object associated with the physical location and the step; and
storing the procedure in a procedure store for subsequent retrieval such that the virtual object is automatically displayed during the subsequent retrieval based on a viewing device being in proximity to the physical location associated with the step, wherein the viewing device displays information corresponding to the step responsive to interaction with the virtual object via the viewing device.
17 . The computer system of claim 16 , wherein a formatting instruction identifies one or more selected from a group consisting of: a language for the steps, characteristics of text to remove from the transcription when generating a step, how to combine text in the step, timing information to include in the step, and any combination thereof.
18 . The computer system of claim 16 , wherein timestamps associated with various text comprise a timestamp associated with each individual word in the text.
19 . The computer system of claim 16 , further storing instructions that, when executed by the one or more processors causing the one or more processors to perform steps comprising:
generating a quiz comprising one or more questions about the procedure by applying a trained quiz generation model to the procedure; and
storing the quiz in the procedure store in association with the procedure.
20 . The computer system of claim 16 , wherein obtaining multimedia content of performance of the task comprises:
receiving multimedia content of the local area where the task is performed from the client device that captured the multimedia content during performance of the task.