IP Library Granted Patent US 10,978,073
Granted Patent B1
US 10,978,073 · App. 16/027,511 · Granted Apr 13, 2021

Systems and methods for processing and presenting conversations

Inventors: Yun Fu (Cupertino, CA); Simon Lau (San Jose, CA); Fuchun Peng (Cupertino, CA); Kaisuke Nakajima (Sunnyvale, CA); Julius Cheng (Cupertino, CA); Gelei Chen (Mountain View, CA); Sam Song Liang (Palo Alto, CA)
Assignee: Otter.ai, Inc.
G10L15/26G06F16/34G06F16/35G10L15/08G10L17/00H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,978,073
App. No.
16/027,511
Granted
Apr 13, 2021
Kind
B1
Abstract

System and method for processing and presenting a conversation. For example, a system includes a sensor configured to capture an audio-form conversation, and a processor configured to automatically transform the audio-form conversation into a transformed conversation. The transformed conversation includes a synchronized text, and the synchronized text is synchronized with the audio-form conversation. Additionally, the system includes a presenter configured to present the transformed conversation including the synchronized text and the audio-form conversation.

Claims (44)

1. A system for processing and presenting a conversation, the system comprising:

a sensor configured to capture an audio-form conversation;

a controller configured to switch the sensor between a capturing state and an idling state;

an interface configured to receive a user instruction to instruct the controller to switch the sensor between the capturing state and the idling state;

a processor configured to:

automatically transform the audio-form conversation into a transformed conversation, the transformed conversation including a synchronized text, the synchronized text being synchronized with the audio-form conversation;

automatically generate one or more segments of the audio-form conversation and one or more segments of the synchronized text by at least automatically segmenting the audio-form conversation and the synchronized text when a speaker change occurs or a natural pause occurs such that each segment of the one or more segments of the audio-form conversation is spoken by only one speaker in audio form and is synchronized with only one segment of the one or more segments of the synchronized text; and

automatically assign only one speaker label to each segment of the one or more segments of the synchronized text, each one speaker label representing one speaker; and

a presenter configured to present the transformed conversation including the synchronized text and the audio-form conversation.

2. The system of claim 1 wherein the audio-form conversation includes a human-to-human conversation in audio form.

3. The system of claim 2 wherein the human-to-human conversation includes a meeting conversation.

4. The system of claim 2 wherein the human-to-human conversation includes a phone conversation.

5. The system of claim 1 wherein the presenter is further configured to function as the interface.

6. The system of claim 1 , wherein the presenter is further configured to present the transformed conversation, the transformed conversation including the one or more segments of the audio-form conversation and the one or more segments of the synchronized text.

7. The system of claim 1 , wherein the speaker corresponds to a speaker label.

8. The system of claim 7 wherein the speaker label includes a speaker name of the speaker.

9. The system of claim 7 wherein the speaker label includes a speaker picture of the speaker.

10. The system of claim 1 , wherein the processor is further configured to automatically generate the speaker-assigned segmented synchronized text and the corresponding segmented audio-form conversation.

11. The system of claim 10 wherein the presenter is further configured to present the transformed conversation, the transformed conversation including the speaker-assigned segmented synchronized text and the corresponding segmented audio-form conversation.

12. The system of claim 1 wherein:

the processor is further configured to receive metadata including a date for recording the audio-form conversation, a time for recording the audio-form conversation, a duration for recording the audio-form conversation, and a title for the audio-form conversation; and

the presenter is further configured to present the metadata.

13. The system of claim 1 wherein the presenter is further configured to present the transformed conversation both navigable and searchable.

14. The system of claim 13 wherein the presenter is further configured to present one or more matches of a searched text in a first highlighted state, the one or more matches being one or more parts of the synchronized text.

15. The system of claim 14 wherein the presenter is further configured to highlight the audio-form conversation at one or more timestamps, the one or more timestamps corresponding to the one or more matches of the searched text respectively.

16. The system of claim 15 wherein the presenter is further configured to present a playback text in a second highlighted state, the playback text being at least a part of the synchronized text and corresponding to at least a word recited during playback of the audio-form conversation.

17. A computer-implemented method for processing and presenting a conversation, the method comprising:

receiving a user instruction to switch a sensor from an idling state to a capturing state;

switching the sensor from the idling state to the capturing state;

receiving via the sensor, an audio-form conversation;

automatically transforming the audio-form conversation into a transformed conversation, the transformed conversation including a synchronized text, the synchronized text being synchronized with the audio-form conversation;

automatically generating one or more segments of the audio-form conversation and one or more segments of the synchronized text by at least automatically segmenting the audio-form conversation and the synchronized text when a speaker change occurs or a natural pause occurs such that each segment of the one or more segments of the audio-form conversation is spoken by only one speaker in audio form and is synchronized with only one segment of the one or more segments of the synchronized text;

automatically assigning only one speaker label to each segment of the one or more segments of the synchronized text, each one speaker label representing one speaker; and

presenting the transformed conversation including the synchronized text and the audio-form conversation.

18. A non-transitory computer-readable medium with instructions stored thereon, that when executed by a processor, perform the processes comprising:

receiving a user instruction to switch a sensor from an idling state to a capturing state;

switching the sensor from the idling state to the capturing state;

receiving, via the sensor, an audio-form conversation;

automatically transforming the audio-form conversation into a transformed conversation, the transformed conversation including a synchronized text, the synchronized text being synchronized with the audio-form conversation;

automatically generating one or more segments of the audio-form conversation and one or more segments of the synchronized text by at least automatically segmenting the audio-form conversation and the synchronized text when a speaker change occurs or a natural pause occurs such that each segment of the one or more segments of the audio-form conversation is spoken by only one speaker in audio form and is synchronized with only one segment of the one or more segments of the synchronized text;

automatically assigning only one speaker label to each segment of the one or more segments of the synchronized text, each one speaker label representing one speaker; and

presenting the transformed conversation including the synchronized text and the audio-form conversation.

19. The computer-implemented method of claim 17 , wherein the audio-form conversation includes a human-to-human conversation in audio form.

20. The non-transitory computer-readable medium of claim 18 , wherein the audio-form conversation includes a human-to-human conversation in audio form.

Assignments (2)
CHANGE OF NAME Recorded Oct 7, 2020
From: AISENSE, INC.
To: OTTER.AI, INC.
Reel/Frame 054006/0966 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2020
From: FU, YUN; LAU, SIMON; NAKAJIMA, KAISUKE; CHENG, JULIUS; CHEN, GELEI; LIANG, SAM SONG; PENG, FUCHUN
To: AISENSE, INC.
Reel/Frame 053977/0604 →
Continuity (1)
Provisional Application 62530227 · Jul 9, 2017
Cited By (9)
US 12,400,661 US 12,406,672 US 12,406,684 US 12,456,465 US 12,462,808 US 12,494,929 US 12,518,748 US 12,518,762 US 12,555,581