Systems and methods configured to generate object vestiges based on audio information conveying natural speech
Systems and methods configured to generate object vestiges based on audio information conveying natural speech are disclosed. Exemplary implementations may: obtain audio information captured by a client computing platform; provide the audio information to a trained intent recognition model, wherein the trained intent recognition model is trained to determine one or more semantic objects based on the audio information and subsequently determine object vestiges based on the one or more semantic objects, wherein the semantic objects indicate entities and have intent types, wherein the intent types include a content type and a directive type, wherein the entities are spoken by the participants or referred to by the participants; obtain, from the trained intent recognition model, the object vestiges; and store the object vestiges to an electronic record of the subject.
1 . A system configured to generate object vestiges based on audio information conveying natural speech, the system comprising:
electronic storage that stores electronic records for subjects and a trained intent recognition model, wherein the trained intent recognition model is trained to determine object vestiges based on semantic objects as input, wherein the semantic objects indicate entities and have intent types, wherein the intent types include a content type and a directive type, and wherein an individual object vestige is a result caused or generated based on an individual semantic object, the result including one or both of a generation of content or a sending of a directive to an external resource; and
one or more processors configured by machine-readable instructions to:
obtain audio information captured by a client computing platform, wherein the audio information conveys conversational speech from participants about a subject;
provide the audio information to the trained intent recognition model, wherein the trained intent recognition model is trained to determine one or more of the semantic objects based on the audio information and subsequently determine the object vestiges based on the one or more semantic objects, wherein the entities are spoken by the participants or referred to by the participants;
obtain, from the trained intent recognition model, the object vestiges, wherein individual ones of the object vestiges include timing information indicating times at which the entities included in the semantic objects corresponding to the object vestiges were spoken, wherein the times are based on the audio information, wherein the individual ones of the object vestiges are one of the intent types, wherein the object vestiges of:
the content type include content related to the subject and the entities, and
the directive type include one or more values for one or more directive parameters related to the entities for external resources to initiate actions for the subject; and
store the object vestiges to an electronic record of the subject, wherein storing the object vestiges includes grouping the object vestiges based on the timing information.
2 . The system of claim 1 , wherein the one or more processors are further configured by the machine-readable instructions to:
receive a request specifying one or more particular subtypes of the content types and/or one or more particular subtypes of the directive types for determining the semantic objects.
3 . The system of claim 2 , wherein the subtypes of the content types include a clinical note, directions, a summary, a follow up note, a letter, wherein the letter includes a referral letter, a thank you letter, an occupational note, and an approval letter.
4 . The system of claim 3 , wherein the subtypes of the directive types include prescription orders, device orders, imaging orders, visit scheduling, prescription modifications, device modifications, and/or scheduled visit modifications.
5 . The system of claim 4 , wherein the object vestiges are based on the electronic record of the subject, the subtypes of the content types, and/or the subtypes of the directive types input to the trained intent recognition model.
6 . The system of claim 1 , wherein the participants include a user of the client computing platform, one or more secondary users of the client computing platform, the subject, and/or a guardian of the subject.
7 . The system of claim 4 , wherein the one or more processors are further configured by the machine-readable instructions to:
effectuate, via the client computing platform of the user, presentation of an ordered list including the subtypes of the content types and the subtypes of the directive types that are selectable by the user, wherein determining the semantic objects is based on the selections, the electronic record of the subject, the user, and/or a purpose of the speech.
8 . A system configured to train an intent recognition model to determine object vestiges based on conversational input, the system comprising:
electronic storage that stores an intent recognition model and training information, wherein the training information includes (i) audio information representing conversational speech and/or transcripts corresponding to the audio information, (ii) entities included in the audio information and/or the transcripts, (iii) intent types, and (iv) final object vestiges, wherein the intent types include a content type and a directive type; and
one or more processors configured by machine-readable instructions to:
obtain the intent recognition model;
obtain the training information;
train the intent recognition model by using the audio information and/or the transcripts, the entities, and the intent types as training input, and the final object vestiges as the training outputs so that the intent recognition model is trained to determine semantic objects that indicate the entities and have one of the intent types and subsequently determine object vestiges based on the semantic objects, wherein individual ones of the object vestiges are one of the intent types and either include content related to the entities or values for directive parameters related to the entities to initiate actions, wherein an individual object vestige is a result caused or generated based on an individual semantic object, the result including one or both of a generation of content or a sending of a directive to an external resource, wherein the individual ones of the object vestiges include timing information indicating times at which the entities included in the semantic objects corresponding to the object vestiges were spoken, wherein the times are based on the audio information, and wherein the object vestiges are stored based on grouping the object vestiges based on the timing information; and
store the trained intent recognition model to the electronic storage.
9 . A method configured to generate object vestiges based on audio information conveying natural speech, the method comprising:
obtaining audio information captured by a client computing platform, wherein the audio information conveys conversational speech from participants about a subject;
providing the audio information to a trained intent recognition model, wherein the trained intent recognition model is trained to determine one or more semantic objects based on the audio information and subsequently determine object vestiges based on the one or more semantic objects, wherein the semantic objects indicate entities and have intent types, wherein the intent types include a content type and a directive type, wherein the entities are spoken by the participants or referred to by the participants, and wherein an individual object vestige is a result caused or generated based on an individual semantic object, the result including one or both of a generation of content or a sending of a directive to an external resource;
obtaining, from the trained intent recognition model, the object vestiges, wherein individual ones of the object vestiges include timing information indicating times at which the entities included in the semantic objects corresponding to the object vestiges were spoken, wherein the times are based on the audio information, wherein the individual ones of the object vestiges are one of the intent types, wherein the object vestiges of the content type include content related to the subject and the entities, and wherein the object vestiges of the directive type include one or more values for one or more directive parameters related to the entities for external resources to initiate actions for the subject; and
storing the object vestiges to an electronic record of the subject, wherein the storing of the object vestiges includes grouping the object vestiges based on the timing information.
10 . The method of claim 9 , further comprising:
receiving a request specifying one or more particular subtypes of the content types and/or one or more particular subtypes of the directive types for determining the semantic objects.
11 . The method of claim 10 , wherein the subtypes of the content types include a clinical note, directions, a summary, a follow up note, a letter, wherein the letter includes a referral letter, a thank you letter, an occupational note, and an approval letter.
12 . The method of claim 11 , wherein the subtypes of the directive types include prescription orders, device orders, imaging orders, visit scheduling, prescription modifications, device modifications, and/or scheduled visit modifications.
13 . The method of claim 12 , wherein the object vestiges are based on the electronic record of the subject, the subtypes of the content types, and/or the subtypes of the directive types input to the trained intent recognition model.
14 . The method of claim 9 , wherein the participants include a user of the client computing platform, one or more secondary users of the client computing platform, the subject, and/or a guardian of the subject.
15 . The method of claim 12 , further comprising:
effectuating, via the client computing platform of the user, presentation of an ordered list including the subtypes of the content types and the subtypes of the directive types that are selectable by the user, wherein determining the semantic objects is based on the selections, the electronic record of the subject, the user, and/or a purpose of the speech.
16 . A method to train an intent recognition model to determine object vestiges based on conversational input, the method comprising:
obtaining an intent recognition model from electronic storage, wherein the electronic storage further stores training information, wherein the training information includes (i) audio information representing conversational speech and/or transcripts corresponding to the audio information, (ii) entities included in the audio information and/or the transcripts, (iii) intent types, and (iv) final object vestiges, wherein the intent types include a content type and a directive type;
obtaining the training information;
training the intent recognition model by using the audio information and/or the transcripts, the entities, and the intent types as training input, and the final object vestiges as the training outputs so that the intent recognition model is trained to determine semantic objects that indicate the entities and have one of the intent types and subsequently determine object vestiges based on the semantic objects, wherein individual ones of the object vestiges are one of the intent types and either include content related to the entities or values for directive parameters related to the entities to initiate actions, and wherein an individual object vestige is a result caused or generated based on an individual semantic object, the result including one or both of a generation of content or a sending of a directive to an external resource, wherein the individual ones of the object vestiges include timing information indicating times at which the entities included in the semantic objects corresponding to the object vestiges were spoken, wherein the times are based on the audio information, and wherein the object vestiges are stored based on grouping the object vestiges based on the timing information; and
storing the trained intent recognition model to the electronic storage.