Systems and methods for processing and presenting conversations
System and method for processing and presenting a conversation. For example, a system includes a sensor configured to capture an audio-form conversation, and a processor configured to automatically transform the audio-form conversation into a transformed conversation. The transformed conversation includes a synchronized text, and the synchronized text is synchronized with the audio-form conversation. Additionally, the system includes a presenter configured to present the transformed conversation including the synchronized text and the audio-form conversation.
1. A system for processing and presenting a conversation, the system comprising:
a sensor configured to, upon receipt of a user instruction, capture a live audio-form conversation;
a processor configured to, upon receiving the live audio-form conversation when the live audio-form conversation is being captured:
automatically transcribe, in real time or in near-real time with the live audio-form conversation, the live audio-form conversation into a live synchronized text, the live synchronized text being synchronized with the live audio-form conversation;
automatically generate, in real time or in near-real time with the live audio-form conversation, one or more segments of the live audio-form conversation and one or more segments of the live synchronized text by at least:
identifying when a speaker change occurs during the live audio-form conversation;
in response to the identifying when a speaker change occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is spoken by only one speaker;
identifying when a natural pause occurs during the live audio-form conversation;
in response to the identifying when a natural pause occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is synchronized with only one segment of the one or more segments of the live synchronized text; and
automatically assign, in real time or in near-real time with the live audio-form conversation, only one speaker label to each segment of the one or more segments of the live synchronized text, each one speaker label representing one speaker; and
a presenter configured to present, in real time or in near-real time with the live audio-form conversation, the labeled live synchronized text and the live audio-form conversation;
wherein in near-real time is a time delay less than one minute.
2. The system of claim 1 , wherein the live audio-form conversation includes a human-to-human conversation in audio form.
3. The system of claim 1 , wherein the live human-to-human conversation includes a meeting conversation.
4. The system of claim 1 , wherein the live human-to-human conversation includes a phone conversation.
5. The system of claim 1 , wherein the presenter is further configured to automatically present, in real time or near-real time with the live audio-form conversation, the speaker-assigned segmented live synchronized text and the corresponding segmented live audio-form conversation.
6. The system of claim 1 , wherein the presenter is further configured to present the labeled live synchronized text to be both navigable and searchable.
7. The system of claim 6 , wherein the presenter is further configured to present one or more matches of a searched text in a first highlighted state, the one or more matches being one or more parts of the synchronized text.
8. The system of claim 7 , wherein the presenter is further configured to highlight the live audio-form conversation at one or more timestamps, the one or more timestamps corresponding to the one or more matches of the searched text respectively.
9. A computer-implemented method for processing and presenting a conversation, the method comprising:
capturing, via a sensor, a live audio-form conversation; and
upon receiving the live audio-form conversation when the live audio-form conversation is being captured:
automatically transcribing, in real time or in near-real time with the live audio-form conversation, the live audio-form conversation into a live synchronized text, the live synchronized text being synchronized with the live audio-form conversation;
automatically generating, in real time or in near-real time with the live audio-form conversation, one or more segments of the live audio-form conversation and one or more segments of the live synchronized text by at least:
identifying when a speaker change occurs during the live audio-form conversation;
in response to the identifying when a speaker change occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is spoken by only one speaker;
identifying when a natural pause occurs during the live audio-form conversation;
in response to the identifying when a natural pause occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is synchronized with only one segment of the one or more segments of the live synchronized text;
automatically assigning, in real time or in near-real time with the live audio-form conversation, only one speaker label to each segment of the one or more segments of the live synchronized text, each one speaker label representing one speaker; and
presenting, in real time or in near-real time with the live audio-form conversation, the labeled live synchronized text and the live audio-form conversation;
wherein in near-real time is a time delay less than one minute.
10. The computer-implemented method of claim 9 , wherein the live audio-form conversation includes a human-to-human conversation in audio form.
11. The computer-implemented method of claim 9 , wherein the live human-to-human conversation includes a meeting conversation.
12. The computer-implemented method of claim 9 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting, in real time or near-real time with the live audio-form conversation, the speaker-assigned segmented live synchronized text and the corresponding segmented live audio-form conversation.
13. The computer-implemented method of claim 9 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting the labeled live synchronized text to be both navigable and searchable.
14. The computer-implemented method of claim 9 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting one or more matches of a searched text in a first highlighted state, the one or more matches being one or more parts of the synchronized text.
15. The computer-implemented method of claim 14 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes highlighting the live audio-form conversation at one or more timestamps, the one or more timestamps corresponding to the one or more matches of the searched text respectively.
16. A non-transitory computer-readable medium with instructions stored thereon, that when executed by a processor, perform the processes comprising:
capturing, via a sensor, a live audio-form conversation; and
upon receiving the live audio-form conversation when the live audio-form conversation is being captured:
automatically transcribing, in real time or in near-real time with the live audio-form conversation, the live audio-form conversation into a live synchronized text, the live synchronized text being synchronized with the live audio-form conversation;
automatically generating, in real time or in near-real time with the live audio-form conversation, one or more segments of the live audio-form conversation and one or more segments of the live synchronized text by at least:
identifying when a speaker change occurs during the live audio-form conversation;
in response to the identifying when a speaker change occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is spoken by only one speaker;
identifying when a natural pause occurs during the live audio-form conversation;
in response to the identifying when a natural pause occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is synchronized with only one segment of the one or more segments of the live synchronized text;
automatically assigning, in real time or in near-real time with the live audio-form conversation, only one speaker label to each segment of the one or more segments of the live synchronized text, each one speaker label representing one speaker; and
presenting, in real time or in near-real time with the live audio-form conversation, the labeled live synchronized text and the live audio-form conversation;
wherein in near-real time is a time delay less than one minute.
17. The non-transitory computer-readable medium of claim 16 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting, in real time or near-real time with the live audio-form conversation, the speaker-assigned segmented live synchronized text and the corresponding segmented live audio-form conversation.
18. The non-transitory computer-readable medium of claim 16 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting the labeled live synchronized text to be both navigable and searchable.
19. The non-transitory computer-readable medium of claim 18 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting one or more matches of a searched text in a first highlighted state, the one or more matches being one or more parts of the synchronized text.
20. The non-transitory computer-readable medium of claim 19 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes highlighting the live audio-form conversation at one or more timestamps, the one or more timestamps corresponding to the one or more matches of the searched text respectively.