IP Library Granted Patent US 12,182,502
Granted Patent B1
US 12,182,502 · App. 18/127,343 · Granted Dec 31, 2024

Systems and methods for automatically generating conversation outlines and annotation summaries

Inventors: Kaisuke Nakajima (Sunnyvale, CA); Kean Kheong Chin (Reno, NV); Gregory Kennedy Sell (Bremerton, WA); Cheng Yuan (San Jose, CA); Amro A. Younes (Redwood City, CA); Richard Norman Michael Ward (San Francisco, CA); Robert Firebaugh (Tiburon, CA); Qingyun Mao (Newark, CA); Amanda Song (Seattle, WA); Simon Lau (San Jose, CA); Siddharth Pradeep Sakhadeo (Milpitas, CA); Winfred James Jebasingh (Campbell, CA); Jiankai Xiao (Sunnyvale, CA); Shreyas Aiyar (Sunnyvale, CA); Frazer Hainsworth Kirkman (Sunnyvale, CA); Wen Sun (Sunnyvale, CA); Angus Ka-man Ng (San Francisco, CA); Sam Liang (Palo Alto, CA); Yun Fu (Cupertino, CA)
Assignee: Otter.ai, Inc.
G06F40/169G10L17/14G10L17/18G10L17/22G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,182,502
App. No.
18/127,343
Granted
Dec 31, 2024
Kind
B1
Abstract

Method, system, device, and non-transitory computer-readable medium for presenting a conversation. A computer-implemented method may include: obtaining, via a first virtual participant, a first set of audio data associated with the first conversation while the first conversation occurs; transcribing the first set of audio data into a first set of text data while the first conversation occurs; obtaining a set of annotations associated with the set of text data while the first conversation occurs; identifying one or more topic transitions based at least in part upon the set of text data; generating a conversation summary based at least in part upon the one or more topic transitions; obtaining a first set of visual data associated with the conversation; and presenting the set of annotations, the conversation summary, and the first set of visual data embedded in the first set of text data to the first group of actual participants.

Claims (91)

1. A computer-implemented method for presenting a conversation, the method comprising:

obtaining, via a virtual participant, a set of audio data associated with a conversation while the conversation occurs;

transcribing, via the virtual participant, the set of audio data into a set of text data while the conversation occurs;

obtaining a set of annotations associated with the set of text data while the conversation occurs;

identifying one or more topic transitions based at least in part upon the set of text data;

generating a plurality of headings for a plurality of conversation topics based at least in part on the set of text data and the one or more topic transitions, the generating a plurality of headings including applying a machine learning model to the set of text data to generate at least one of the plurality of headings;

generating a conversation summary including the plurality of headings based at least in part upon the one or more topic transitions;

obtaining a set of visual data associated with the conversation, each visual data of the set of visual data corresponding to a timestamp; and

presenting the set of annotations, the conversation summary, and the set of visual data embedded in the set of text data to a group of actual participants.

2. The computer-implemented method of claim 1 , wherein the generating a plurality of headings comprises generating at least one of the plurality of headings based at least in part on one or more key words.

3. The computer-implemented method of claim 1 , wherein the machine learning model includes a sequence-to-sequence machine learning model or a sequence-to-sequence neural network.

4. The computer-implemented method of claim 1 , further comprising:

receiving an input from a user; and

updating time information for one of the plurality of conversation topics based on the input;

wherein the time information includes at least one selected from a group consisting of a start time, an end time and a transition time.

5. The computer-implemented method of claim 1 , further comprising:

generating metadata associated with each visual data of the set of visual data;

wherein the metadata includes at least one selected from a group consisting of a link to a respective visual data and a time offset for the respective visual data from a beginning of the conversation.

6. The computer-implemented method of claim 5 , wherein the respective visual data is embedded in the set of text data based at least in part on the time offset.

7. The computer-implemented method of claim 1 , wherein the obtaining a set of visual data associated with the conversation comprises automatically obtaining the set of visual data.

8. The computer-implemented method of claim 1 , wherein the obtaining a set of visual data associated with the conversation comprises:

capturing a sequence of visual data at regular time intervals;

receiving an input associated with one visual data in the sequence of visual data;

selecting the one visual data based on the input; and

embedding the one selected visual data in the set of text data.

9. The computer-implemented method of claim 8 , wherein the receiving an input associated with one visual data in the sequence of visual data comprises receiving the input associated with a thumbnail representing the one visual data in the sequence of visual data.

10. The computer-implemented method of claim 1 , further comprising:

embedding at least one of the set of visual data to the set of text data at a time prior to a current time of the conversation.

11. The computer-implemented method of claim 1 , further comprising:

generating an annotation summary including the one or more annotations and one or more corresponding timestamps.

12. The computer-implemented method of claim 11 , further comprising:

receiving a selection of an annotation from the one or more annotations; and

identifying a conversation segment associated with the selected annotation.

13. The computer-implemented method of claim 1 , wherein the identifying one or more topic transitions comprises at least selected from a group consisting of:

identifying a change in speakers;

identifying a change in screenshare;

identifying a change in cue words;

identifying a pause in a conversation; and

identifying a change in semantic meaning of two or more conversation segments.

14. The computer-implemented method of claim 1 , further comprising:

identifying a first speaker associated with a first audio channel; and

identifying a second speaker associated with a second audio channel, the second audio channel being different from the first audio channel;

wherein the identifying one or more topic transitions comprises identifying the one or more topic transitions based at least in part on the identified first speaker and the identified second speaker.

15. A computing system for presenting a conversation, the computing system comprising:

one or more processors; and

a memory storing instructions that, upon execution by the one or more processors, cause the computing system to perform one or more processes comprising:

obtaining, via a virtual participant, a set of audio data associated with a conversation while the conversation occurs;

transcribing, via the virtual participant, the set of audio data into a set of text data while the conversation occurs;

obtaining a set of annotations associated with the set of text data while the conversation occurs;

identifying one or more topic transitions based at least in part upon the set of text data;

generating a plurality of headings for a plurality of conversation topics based at least in part on the set of text data and the one or more topic transitions, the generating a plurality of headings including applying a machine learning model to the set of text data to generate at least one of the plurality of headings;

generating a conversation summary including the plurality of headings based at least in part upon the one or more topic transitions;

obtaining a set of visual data associated with the conversation, each visual data of the set of visual data corresponding to a timestamp; and

presenting the set of annotations, the conversation summary, and the set of visual data embedded in the set of text data to a group of actual participants.

16. The computing system of claim 15 , wherein the generating a plurality of headings comprises generating at least one of the plurality of headings based at least in part on one or more key words.

17. The computing system of claim 15 , wherein the machine learning model includes a sequence-to-sequence machine learning model or a sequence-to-sequence neural network.

18. The computing system of claim 15 , wherein the one or more processes further comprise:

receiving an input from a user; and

updating time information for one of the plurality of conversation topics based on the input;

wherein the time information includes at least one selected from a group consisting of a start time, an end time and a transition time.

19. The computing system of claim 15 , wherein the one or more processes further comprise:

generating metadata associated with each visual data of the set of visual data;

wherein the metadata includes at least one selected from a group consisting of a link to a respective visual data and a time offset for the respective visual data from a beginning of the conversation.

20. The computing system of claim 19 , wherein the respective visual data is embedded in the set of text data based at least in part on the time offset.

21. The computing system of claim 15 , wherein the obtaining a set of visual data associated with the conversation comprises automatically obtaining the set of visual data.

22. The computing system of claim 15 , wherein the obtaining a set of visual data associated with the conversation comprises:

capturing a sequence of visual data at regular time intervals;

receiving an input associated with one visual data in the sequence of visual data;

selecting the one visual data based on the input; and

embedding the one selected visual data in the set of text data.

23. The computing system of claim 22 , wherein the receiving an input associated with one visual data in the sequence of visual data comprises receiving the input associated with a thumbnail representing the one visual data in the sequence of visual data.

24. The computing system of claim 15 , further comprising:

embedding at least one of the set of visual data to the set of text data at a time prior to a current time of the conversation.

25. A non-transitory computer-readable medium storing instructions for presenting a conversation, the instructions upon execution by one or more processors of a computing system, cause the computing system to perform one or more processes including:

obtaining, via a virtual participant, a set of audio data associated with a conversation while the conversation occurs;

transcribing, via the virtual participant, the set of audio data into a set of text data while the conversation occurs;

obtaining a set of annotations associated with the set of text data while the conversation occurs;

identifying one or more topic transitions based at least in part upon the set of text data;

generating a plurality of headings for a plurality of conversation topics based at least in part on the set of text data and the one or more topic transitions, the generating a plurality of headings comprising applying a machine learning model to the set of text data to generate at least one of the plurality of headings;

generating a conversation summary including the plurality of headings based at least in part upon the one or more topic transitions;

obtaining a set of visual data associated with the conversation, each visual data of the set of visual data corresponding to a timestamp; and

presenting the set of annotations, the conversation summary, and the set of visual data embedded in the set of text data to a group of actual participants.

26. A computer-implemented method for presenting a conversation, the method comprising:

obtaining, via a virtual participant, a set of audio data associated with a conversation while the conversation occurs;

transcribing, via the virtual participant, the set of audio data into a set of text data while the conversation occurs;

obtaining a set of annotations associated with the set of text data while the conversation occurs;

identifying one or more topic transitions based at least in part upon the set of text data;

generating a plurality of headings for a plurality of conversation topics based at least in part on the set of text data and the one or more topic transitions by applying a machine learning model to the set of text data;

generating a conversation summary based at least in part upon the one or more topic transitions, the conversation summary including the plurality of headings;

obtaining a set of visual data associated with the conversation, each visual data of the set of visual data corresponding to a timestamp; and

presenting the set of annotations, the conversation summary, and the set of visual data embedded in the set of text data to a group of actual participants.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2024
From: NAKAJIMA, KAISUKE; CHIN, KEAN KHEONG; SELL, GREGORY KENNEDY; YUAN, CHENG; YOUNES, AMRO A.; WARD, RICHARD NORMAN MICHAEL; FIREBAUGH, ROBERT; MAO, QINGYUN; SONG, AMANDA; LAU, SIMON; SAKHADEO, SIDDARTH PRADEEP; JEBASINGH, WINFRED JAMES; XIAO, JIANKAI; AIYAR, SHREYAS; SUN, WEN; NG, ANGUS KA-MAN; LIANG, SAM; FU, YUN
To: OTTER.AI, INC.
Reel/Frame 068532/0096 →
AT-WILL EMPLOYMENT, CONFIDENTIAL INFORMATION, INVENTION ASSIGNMENT, AND ARBITRATION AGREEMENT Recorded Sep 9, 2024
From: KIRKMAN, FRAZER
To: AISENSE INC.
Reel/Frame 068904/0453 →
CHANGE OF NAME Recorded Sep 9, 2024
From: AISENSE INC.
To: OTTER.AI, INC.
Reel/Frame 068906/0340 →
Continuity (1)
Provisional Application 63324490 · Mar 28, 2022
Cited By (7)
US 12,400,661 US 12,406,672 US 12,456,465 US 12,462,808 US 12,494,929 US 12,518,748 US 12,555,581