IP Library Granted Patent US 11,431,517
Granted Patent B1
US 11,431,517 · App. 16/780,630 · Granted Aug 30, 2022

Systems and methods for team cooperation with real-time recording and transcription of conversations and/or speeches

Inventors: Simon Lau (San Jose, CA); Yun Fu (Cupertino, CA); James Mason Altreuter (Belmont, CA); Brian Francis Williams (San Carlos, CA); Xiaoke Huang (Foster City, CA); Tao Xing (San Jose, CA); Wen Sun (San Francisco, CA); Tao Lu (Hayward, CA); Kaisuke Nakajima (Sunnyvale, CA); Kean Kheong Chin (Santa Clara, CA); Hitesh Anand Gupta (Santa Clara, CA); Julius Cheng (Cupertino, CA); Jing Pan (Mountain View, CA); Sam Song Liang (Palo Alto, CA)
Assignee: Otter.ai, Inc.
H04L12/1831G10L15/183G10L15/26H04L12/1822
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,431,517
App. No.
16/780,630
Granted
Aug 30, 2022
Kind
B1
Abstract

Methods and systems for team cooperation with real-time recording of one or more moment-associating elements. For example, a method includes: delivering, in response to an instruction, an invitation to each member of one or more members associated with a workspace; granting, in response to acceptance of the invitation by one or more subscribers of the one or more members, subscription permission to the one or more subscribers; receiving the one or more moment-associating elements; transforming the one or more moment-associating elements into one or more pieces of moment-associating information; and transmitting at least one piece of the one or more pieces of moment-associating information to the one or more subscribers.

Claims (81)

1. A computer-implemented method for team cooperation with real-time recording of one or more moment-associating elements, the method comprising:

delivering, in response to an instruction, an invitation to each member of a group of members, the group of members being dedicated to a shared topic;

granting, in response to acceptance of the invitation by a group of subscribers of the group of members, subscription permission to the group of subscribers;

receiving one or more moment-associating elements associated with an event participated by at least the group of subscribers;

obtaining, from the group of subscribers at least while the moment-associating elements is being received, a set of transcription boosters, the set of transcription boosters including:

a plurality of voiceprints corresponding to the group of subscribers, each voiceprint including a source-specific acoustic model and a source-specific language model,

one or more custom vocabularies associated with the shared topic,

one or more suggested speaker tags,

one or more shared information highlights,

one or more shared information edits, and

one or more shared information comments;

transforming, at least while the moment-associating elements is being received, the one or more moment-associating elements into one or more pieces of moment-associating information based at least in part upon the set of transcription boosters; and

transmitting, at least while the moment-associating elements is being received, the one or more pieces of moment-associating information to the one or more subscribers;

wherein the transforming the one or more moment-associating elements includes:

segmenting the one or more moment-associating elements into a plurality of moment-associating segments based at least in part upon the plurality of voiceprints;

assigning a segment speaker for each segment of the plurality of moment-associating segments based at least in part upon the plurality of voiceprints and the one or more suggested speaker tags;

transcribing the plurality of moment-associating segments into a plurality of transcribed segments based at least in part upon the plurality of voiceprints and the one or more custom vocabularies;

generating a plurality of augmented transcribed segments based at least in part upon the plurality of transcribed segments, the one or more shared information highlights, the one or more shared information edits, and the one or more shared information comments; and

generating the one or more pieces of moment-associating information based at least in part on the plurality of augmented transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.

2. The computer-implemented method of claim 1 , wherein the receiving the one or more moment-associating elements includes assigning a timestamp associated with each element of the one or more moment-associating elements.

3. The computer-implemented method of claim 1 , wherein the receiving the one or more moment-associating elements includes at least one selected from receiving one or more audio elements, receiving one or more visual elements, and receiving one or more environmental elements.

4. The computer-implemented method of claim 3 , wherein the receiving one or more audio elements includes at least one selected from receiving one or more voice elements of one or more voice-generating sources and receiving one or more ambient sound elements.

5. The computer-implemented method of claim 3 , wherein the receiving one or more visual elements includes at least one selected from receiving one or more pictures, receiving one or more images, receiving one or more screenshots, receiving one or more video frames, receiving one or more projections, and receiving one or more holograms.

6. The computer-implemented method of claim 3 , wherein the receiving one or more environmental elements includes at least one selected from receiving one or more global positions, receiving one or more location types, and receiving one or more moment conditions.

7. The computer-implemented method of claim 3 , wherein the receiving one or more environmental elements includes at least one selected from receiving a longitude, receiving a latitude, receiving an altitude, receiving a country, receiving a city, receiving a street, receiving a location type, receiving a temperature, receiving a humidity, receiving a movement, receiving a velocity of a movement, receiving a direction of a movement, receiving an ambient noise level, and receiving one or more echo properties.

8. The computer-implemented method of claim 3 , wherein the transforming the one or more moment-associating elements into one or more pieces of moment-associating information includes:

segmenting the one or more audio elements into a plurality of audio segments;

assigning a segment speaker for each segment of the plurality of audio segments;

transcribing the plurality of audio segments into a plurality of text segments; and

generating the one or more pieces of moment-associating information based at least in part on the plurality of text segments and the segment speaker assigned for each segment of the plurality of audio segments.

9. The computer-implemented method of claim 8 , wherein the transcribing the plurality of audio segments into a plurality of text segments includes transcribing two or more segments of the plurality of audio segments in conjunction with each other.

10. The computer-implemented method of claim 1 , wherein the transcribing the plurality of moment-associating segments into a plurality of transcribed segments includes:

transcribing a first segment of the plurality of moment-associating segments into a first transcribed segment of the plurality of transcribed segments;

transcribing a second segment of the plurality of moment-associating segments into a second transcribed segment of the plurality of transcribed segments; and

correcting the first transcribed segment based at least in part on the second transcribed segment.

11. The computer-implemented method of claim 1 , wherein the segmenting the one or more moment-associating elements into a plurality of moment-associating segments includes:

determining one or more speaker-change timestamps, each timestamp of the one or more speaker-change timestamps corresponding to a timestamp when a speaker change occurs;

determining one or more sentence-change timestamps, each timestamp of the one or more sentence-change timestamps corresponding to a timestamp when a sentence change occurs; and

determining one or more topic-change timestamps, each timestamp of the one or more topic-change timestamps corresponding to a timestamp when a topic change occurs.

12. The computer-implemented method of claim 11 , wherein the segmenting the one or more moment-associating elements into a plurality of moment-associating segments is performed based at least in part on one of:

the one or more speaker-change timestamps;

the one or more sentence-change timestamps; and

the one or more topic-change timestamps.

13. A system for team cooperation with real-time recording of one or more moment-associating elements, the system comprising:

an invitation delivering module configured to deliver, in response to an instruction, an invitation to each member of a group of members the group of members being dedicated to a shared topic;

a permission module configured to grant, in response to acceptance of the invitation by a group of subscribers of the group of members, subscription permission to the group of subscribers;

a receiving module configured to receive one or more moment-associating elements associated with an event participated by at least the group of subscribers;

a booster module configured to obtain, from the group of subscribers at least while the moment-associating elements is being received, a set of transcription boosters, the set of transcription boosters including:

a plurality of voiceprints corresponding to the group of subscribers, each voiceprint including a source-specific acoustic model and a source-specific language model,

one or more custom vocabularies associated with the shared topic,

one or more suggested speaker tags,

one or more shared information highlights,

one or more shared information edits, and

one or more shared information comments;

a transforming module configured to transform, at least while the moment-associating elements is being received, the one or more moment-associating elements into one or more pieces of moment-associating information based at least in part upon the set of transcription boosters; and

a transmitting module configured to transmit, at least while the moment-associating elements is being received, the one or more pieces of moment-associating information to the one or more subscribers;

wherein the transforming module is further configured to:

segment the one or more moment-associating elements into a plurality of moment-associating segments based at least in part upon the plurality of voiceprints;

assign a segment speaker for each segment of the plurality of moment-associating segments based at least in part upon the plurality of voiceprints and the one or more suggested speaker tags;

transcribe the plurality of moment-associating segments into a plurality of transcribed segments based at least in part upon the plurality of voiceprints and the one or more custom vocabularies;

generate a plurality of augmented transcribed segments based at least in part upon the plurality of transcribed segments, the one or more shared information highlights, the one or more shared information edits, and the one or more shared information comments; and

generate the one or more pieces of moment-associating information based at least in part on the plurality of augmented transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.

14. A non-transitory computer-readable medium with instructions stored thereon, that when executed by a processor, perform the processes comprising:

delivering, in response to an instruction, an invitation to each member of a group of members, the group of members being dedicated to a shared topic;

granting, in response to acceptance of the invitation by a group of subscribers of the group of members, subscription permission to the group of subscribers;

receiving one or more moment-associating elements associated with an event participated by at least the group of subscribers;

obtaining, from the group of subscribers at least while the moment-associating elements is being received, a set of transcription boosters, the set of transcription boosters including:

a plurality of voiceprints corresponding to the group of subscribers, each voiceprint including a source-specific acoustic model and a source-specific language model,

one or more custom vocabularies associated with the shared topic,

one or more suggested speaker tags,

one or more shared information highlights,

one or more shared information edits, and

one or more shared information comments;

transforming, at least while the moment-associating elements is being received, the one or more moment-associating elements into one or more pieces of moment-associating information based at least in part upon the set of transcription boosters; and

transmitting, at least while the moment-associating elements is being received, the one or more pieces of moment-associating information to the one or more subscribers;

wherein the transforming the one or more moment-associating elements includes:

segmenting the one or more moment-associating elements into a plurality of moment-associating segments based at least in part upon the plurality of voiceprints;

assigning a segment speaker for each segment of the plurality of moment-associating segments based at least in part upon the plurality of voiceprints and the one or more suggested speaker tags;

transcribing the plurality of moment-associating segments into a plurality of transcribed segments based at least in part upon the plurality of voiceprints and the one or more custom vocabularies;

generating a plurality of augmented transcribed segments based at least in part upon the plurality of transcribed segments, the one or more shared information highlights, the one or more shared information edits, and the one or more shared information comments; and

generating the one or more pieces of moment-associating information based at least in part on the plurality of augmented transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2021
From: AISENSE, INC.
To: OTTER.AI, INC.
Reel/Frame 054914/0840 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2020
From: LAU, SIMON; FU, YUN; ALTREUTER, JAMES MASON; WILLIAMS, BRIAN FRANCIS; HUANG, XIAOKE; XING, TAO; SUN, WEN; LU, TAO; NAKAJIMA, KAISUKE; CHIN, KEAN KHEONG; GUPTA, HITESH ANAND; CHENG, JULIUS; PAN, JING; LIANG, SAM SONG
To: AISENSE, INC.
Reel/Frame 054639/0012 →
Continuity (3)
Continuation 16598820 · Oct 11, 2019
Provisional Application 62802098 · Feb 6, 2019
Provisional Application 62747001 · Oct 17, 2018
Cited By (8)
US 12,400,661 US 12,406,672 US 12,406,684 US 12,456,465 US 12,462,808 US 12,494,929 US 12,518,748 US 12,555,581