IP Library Granted Patent US 11,657,822
Granted Patent B2
US 11,657,822 · App. 17/195,202 · Granted May 23, 2023

Systems and methods for processing and presenting conversations

Inventors: Yun Fu (Cupertino, CA); Simon Lau (San Jose, CA); Fuchun Peng (Cupertino, CA); Kaisuke Nakajima (Sunnyvale, CA); Julius Cheng (Cupertino, CA); Gelei Chen (Mountain View, CA); Sam Song Liang (Palo Alto, CA)
Assignee: Otter.ai, Inc.
G10L15/26G06F16/34G06F16/35G10L15/08G10L17/00H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,657,822
App. No.
17/195,202
Granted
May 23, 2023
Kind
B2
Abstract

System and method for processing and presenting a conversation. For example, a system includes a sensor configured to capture an audio-form conversation, and a processor configured to automatically transform the audio-form conversation into a transformed conversation. The transformed conversation includes a synchronized text, and the synchronized text is synchronized with the audio-form conversation. Additionally, the system includes a presenter configured to present the transformed conversation including the synchronized text and the audio-form conversation.

Claims (53)

1. A system for processing and presenting a conversation, the system comprising:

a sensor configured to, upon receipt of a user instruction, capture a live audio-form conversation;

a processor configured to, upon receiving the live audio-form conversation when the live audio-form conversation is being captured:

automatically transcribe, in real time or in near-real time with the live audio-form conversation, the live audio-form conversation into a live synchronized text, the live synchronized text being synchronized with the live audio-form conversation;

automatically generate, in real time or in near-real time with the live audio-form conversation, one or more segments of the live audio-form conversation and one or more segments of the live synchronized text by at least:

identifying when a speaker change occurs during the live audio-form conversation;

in response to the identifying when a speaker change occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is spoken by only one speaker;

identifying when a natural pause occurs during the live audio-form conversation;

in response to the identifying when a natural pause occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is synchronized with only one segment of the one or more segments of the live synchronized text; and

automatically assign, in real time or in near-real time with the live audio-form conversation, only one speaker label to each segment of the one or more segments of the live synchronized text, each one speaker label representing one speaker; and

a presenter configured to present, in real time or in near-real time with the live audio-form conversation, the labeled live synchronized text and the live audio-form conversation;

wherein in near-real time is a time delay less than one minute.

2. The system of claim 1 , wherein the live audio-form conversation includes a human-to-human conversation in audio form.

3. The system of claim 1 , wherein the live human-to-human conversation includes a meeting conversation.

4. The system of claim 1 , wherein the live human-to-human conversation includes a phone conversation.

5. The system of claim 1 , wherein the presenter is further configured to automatically present, in real time or near-real time with the live audio-form conversation, the speaker-assigned segmented live synchronized text and the corresponding segmented live audio-form conversation.

6. The system of claim 1 , wherein the presenter is further configured to present the labeled live synchronized text to be both navigable and searchable.

7. The system of claim 6 , wherein the presenter is further configured to present one or more matches of a searched text in a first highlighted state, the one or more matches being one or more parts of the synchronized text.

8. The system of claim 7 , wherein the presenter is further configured to highlight the live audio-form conversation at one or more timestamps, the one or more timestamps corresponding to the one or more matches of the searched text respectively.

9. A computer-implemented method for processing and presenting a conversation, the method comprising:

capturing, via a sensor, a live audio-form conversation; and

upon receiving the live audio-form conversation when the live audio-form conversation is being captured:

automatically transcribing, in real time or in near-real time with the live audio-form conversation, the live audio-form conversation into a live synchronized text, the live synchronized text being synchronized with the live audio-form conversation;

automatically generating, in real time or in near-real time with the live audio-form conversation, one or more segments of the live audio-form conversation and one or more segments of the live synchronized text by at least:

identifying when a speaker change occurs during the live audio-form conversation;

in response to the identifying when a speaker change occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is spoken by only one speaker;

identifying when a natural pause occurs during the live audio-form conversation;

in response to the identifying when a natural pause occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is synchronized with only one segment of the one or more segments of the live synchronized text;

automatically assigning, in real time or in near-real time with the live audio-form conversation, only one speaker label to each segment of the one or more segments of the live synchronized text, each one speaker label representing one speaker; and

presenting, in real time or in near-real time with the live audio-form conversation, the labeled live synchronized text and the live audio-form conversation;

wherein in near-real time is a time delay less than one minute.

10. The computer-implemented method of claim 9 , wherein the live audio-form conversation includes a human-to-human conversation in audio form.

11. The computer-implemented method of claim 9 , wherein the live human-to-human conversation includes a meeting conversation.

12. The computer-implemented method of claim 9 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting, in real time or near-real time with the live audio-form conversation, the speaker-assigned segmented live synchronized text and the corresponding segmented live audio-form conversation.

13. The computer-implemented method of claim 9 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting the labeled live synchronized text to be both navigable and searchable.

14. The computer-implemented method of claim 9 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting one or more matches of a searched text in a first highlighted state, the one or more matches being one or more parts of the synchronized text.

15. The computer-implemented method of claim 14 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes highlighting the live audio-form conversation at one or more timestamps, the one or more timestamps corresponding to the one or more matches of the searched text respectively.

16. A non-transitory computer-readable medium with instructions stored thereon, that when executed by a processor, perform the processes comprising:

capturing, via a sensor, a live audio-form conversation; and

upon receiving the live audio-form conversation when the live audio-form conversation is being captured:

automatically transcribing, in real time or in near-real time with the live audio-form conversation, the live audio-form conversation into a live synchronized text, the live synchronized text being synchronized with the live audio-form conversation;

automatically generating, in real time or in near-real time with the live audio-form conversation, one or more segments of the live audio-form conversation and one or more segments of the live synchronized text by at least:

identifying when a speaker change occurs during the live audio-form conversation;

in response to the identifying when a speaker change occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is spoken by only one speaker;

identifying when a natural pause occurs during the live audio-form conversation;

in response to the identifying when a natural pause occurs during the live audio-form conversation, automatically segmenting the live audio-form conversation and the live synchronized text such that each segment of the one or more segments of the live audio-form conversation is synchronized with only one segment of the one or more segments of the live synchronized text;

automatically assigning, in real time or in near-real time with the live audio-form conversation, only one speaker label to each segment of the one or more segments of the live synchronized text, each one speaker label representing one speaker; and

presenting, in real time or in near-real time with the live audio-form conversation, the labeled live synchronized text and the live audio-form conversation;

wherein in near-real time is a time delay less than one minute.

17. The non-transitory computer-readable medium of claim 16 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting, in real time or near-real time with the live audio-form conversation, the speaker-assigned segmented live synchronized text and the corresponding segmented live audio-form conversation.

18. The non-transitory computer-readable medium of claim 16 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting the labeled live synchronized text to be both navigable and searchable.

19. The non-transitory computer-readable medium of claim 18 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes presenting one or more matches of a searched text in a first highlighted state, the one or more matches being one or more parts of the synchronized text.

20. The non-transitory computer-readable medium of claim 19 , wherein the presenting the labeled live synchronized text and the live audio-form conversation includes highlighting the live audio-form conversation at one or more timestamps, the one or more timestamps corresponding to the one or more matches of the searched text respectively.

Assignments (2)
CHANGE OF NAME Recorded Aug 11, 2021
From: AISENSE INC.
To: OTTER.AI, INC.
Reel/Frame 057159/0574 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2021
From: FU, YUN; LAU, SIMON; NAKAJIMA, KAISUKE; CHENG, JULIUS; CHEN, GELEI; LIANG, SAM SONG; PENG, FUCHUN
To: AISENSE, INC.
Reel/Frame 057139/0737 →
Continuity (3)
Continuation 16027511 · Jul 5, 2018
Provisional Application 62530227 · Jul 9, 2017
Related Publication 20210217420A1 · Jul 15, 2021
Cited By (8)
US 12,400,661 US 12,406,672 US 12,406,684 US 12,456,465 US 12,462,808 US 12,494,929 US 12,518,748 US 12,555,581