IP Library Granted Patent US 11,423,911
Granted Patent B1
US 11,423,911 · App. 16/598,820 · Granted Aug 23, 2022

Systems and methods for live broadcasting of context-aware transcription and/or other elements related to conversations and/or speeches

Inventors: Yun Fu (Cupertino, CA); Tao Xing (San Jose, CA); Kaisuke Nakajima (Sunnyvale, CA); Brian Francis Williams (San Carlos, CA); James Mason Altreuter (Belmont, CA); Xiaoke Huang (Foster City, CA); Simon Lau (San Jose, CA); Sam Song Liang (Palo Alto, CA); Kean Kheong Chin (Santa Clara, CA); Wen Sun (San Francisco, CA); Julius Cheng (Cupertino, CA); Hitesh Anand Gupta (Santa Clara, CA)
Assignee: Otter.ai, Inc.
G10L17/02G10L17/00H04H20/95
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,911
App. No.
16/598,820
Granted
Aug 23, 2022
Kind
B1
Abstract

Computer-implemented method and system for processing and broadcasting one or more moment-associating elements. For example, the computer-implemented method includes granting subscription permission to one or more subscribers; receiving the one or more moment-associating elements; transforming the one or more moment-associating elements into one or more pieces of moment-associating information; and transmitting at least one piece of the one or more pieces of moment-associating information to the one or more subscribers. In certain examples, transforming the one or more moment-associating elements includes: segmenting the one or more moment-associating elements into a plurality of moment-associating segments; assigning a segment speaker for each segment of the plurality of moment-associating segments; transcribing the plurality of moment-associating segments into a plurality of transcribed segments; and generating the one or more pieces of moment-associating information based the plurality of transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.

Claims (67)

1. A computer-implemented method for processing and broadcasting one or more moment-associating elements, the method comprising:

connecting with one or more calendar systems containing event information associated with an event, the event including a plurality of speeches;

receiving event information from the one or more calendar systems, the event information including, for each speech, speaker information of one or more speakers, attendee information of one or more attendees, a speech title, a start time, an end time, a custom vocabulary, and location information;

receiving a plurality of voiceprints corresponding to the speakers and the attendees, each voiceprint of the plurality of voiceprints including a source-specific acoustic model and a source-specific language model;

granting subscription permission to one or more subscribers;

receiving one or more moment-associating elements of the event;

transforming the one or more moment-associating elements into one or more pieces of moment-associating information based at least in part on the event information and the plurality of voiceprints; and

transmitting the one or more pieces of moment-associating information to the one or more subscribers;

wherein the transforming the one or more moment-associating elements includes:

creating a custom language model based at least in part on the event information and the plurality of voiceprints, the custom language model including a context associated with the event;

segmenting the one or more moment-associating elements into a plurality of moment-associating segments based at least in part on the plurality of voiceprints;

assigning a segment speaker for each segment of the plurality of moment-associating segments based at least in part on the plurality of voiceprints;

transcribing the plurality of moment-associating segments into a plurality of transcribed segments based at least in part on the custom language model and the plurality of voiceprints; and

generating the one or more pieces of moment-associating information based at least in part on the plurality of transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.

2. The computer-implemented method of claim 1 , wherein the receiving the one or more moment-associating elements includes assigning a timestamp associated with each element of the one or more moment-associating elements.

3. The computer-implemented method of claim 1 , wherein the receiving the one or more moment-associating elements includes at least one selected from receiving one or more audio elements, receiving one or more visual elements, and receiving one or more environmental elements.

4. The computer-implemented method of claim 3 , wherein the receiving one or more audio elements includes at least one selected from receiving one or more voice elements of one or more voice-generating sources and receiving one or more ambient sound elements.

5. The computer-implemented method of claim 3 , wherein the receiving one or more visual elements includes at least one selected from receiving one or more pictures, receiving one or more images, receiving one or more screenshots, receiving one or more video frames, receiving one or more projections, and receiving one or more holograms.

6. The computer-implemented method of claim 3 , wherein the receiving one or more environmental elements includes at least one selected from receiving one or more global positions, receiving one or more location types, and receiving one or more moment conditions.

7. The computer-implemented method of claim 3 , wherein the receiving one or more environmental elements includes at least one selected from receiving a longitude, receiving a latitude, receiving an altitude, receiving a country, receiving a city, receiving a street, receiving a location type, receiving a temperature, receiving a humidity, receiving a movement, receiving a velocity of a movement, receiving a direction of a movement, receiving an ambient noise level, and receiving one or more echo properties.

8. The computer-implemented method of claim 3 , wherein the transforming the one or more moment-associating elements into one or more pieces of moment-associating information includes:

segmenting the one or more audio elements into a plurality of audio segments;

assigning a segment speaker for each segment of the plurality of audio segments;

transcribing the plurality of audio segments into a plurality of text segments; and

generating the one or more pieces of moment-associating information based at least in part on the plurality of text segments and the segment speaker assigned for each segment of the plurality of audio segments.

9. The computer-implemented method of claim 8 , wherein the transcribing the plurality of audio segments into a plurality of text segments includes transcribing two or more segments of the plurality of audio segments in conjunction with each other.

10. The computer-implemented method of claim 1 , wherein the transcribing the plurality of moment-associating segments into a plurality of transcribed segments includes:

transcribing a first segment of the plurality of moment-associating segments into a first transcribed segment of the plurality of transcribed segments;

transcribing a second segment of the plurality of moment-associating segments into a second transcribed segment of the plurality of transcribed segments; and

correcting the first transcribed segment based at least in part on the second transcribed segment.

11. The computer-implemented method of claim 1 , wherein the segmenting the one or more moment-associating elements into a plurality of moment-associating segments includes:

determining one or more speaker-change timestamps, each timestamp of the one or more speaker-change timestamps corresponding to a timestamp when a speaker change occurs;

determining one or more sentence-change timestamps, each timestamp of the one or more sentence-change timestamps corresponding to a timestamp when a sentence change occurs; and

determining one or more topic-change timestamps, each timestamp of the one or more topic-change timestamps corresponding to a timestamp when a topic change occurs.

12. The computer-implemented method of claim 11 , wherein the segmenting the one or more moment-associating elements into a plurality of moment-associating segments is performed based at least in part on one of:

the one or more speaker-change timestamps;

the one or more sentence-change timestamps; and

the one or more topic-change timestamps.

13. A system for processing and broadcasting one or more moment-associating elements, the system comprising:

an information module configured to:

connect with one or more calendar systems containing event information associated with an event, the event including a plurality of speeches;

receive event information from the one or more calendar systems, the event information including, for each speech, speaker information of one or more speakers, attendee information of one or more attendees, a speech title, a start time, an end time, a custom vocabulary, and location information;

receive a plurality of voiceprints corresponding to the speakers and the attendees, each voiceprint of the plurality of voiceprints including a source-specific acoustic model and a source-specific language model;

a permission module configured to grant subscription permission to one or more subscribers;

a receiving module configured to receive one or more moment-associating elements of the event;

a transforming module configured to transform the one or more moment-associating elements into one or more pieces of moment-associating information based at least in part on the event information and the plurality of voiceprints; and

a transmitting module configured to transmit the one or more pieces of moment-associating information to the one or more subscribers;

wherein the transforming module is further configured to:

create a custom language model based at least in part on the event information and the plurality of voiceprints, the custom language model including a context associated with the event;

segment the one or more moment-associating elements into a plurality of moment-associating segments based at least in part on the plurality of voiceprints;

assign a segment speaker for each segment of the plurality of moment-associating segments based at least in part on the plurality of voiceprints;

transcribe the plurality of moment-associating segments into a plurality of transcribed segments based at least in part on the plurality of voiceprints and the custom language model; and

generate the one or more pieces of moment-associating information based at least in part on the plurality of transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.

14. A non-transitory computer-readable medium with instructions stored thereon, that when executed by a processor, perform the processes comprising:

connecting with one or more calendar systems containing event information associated with an event, the event including a plurality of speeches;

receiving event information from the one or more calendar systems, the event information including, for each speech, speaker information of one or more speakers, attendee information of one or more attendees, a speech title, a start time, an end time, a custom vocabulary, and location information;

receiving a plurality of voiceprints corresponding to the speakers and the attendees, each voiceprint of the plurality of voiceprints including a source-specific acoustic model and a source-specific language model;

granting subscription permission to one or more subscribers;

receiving one or more moment-associating elements of the event;

transforming the one or more moment-associating elements into one or more pieces of moment-associating information based at least in part on the event information and the plurality of voiceprints; and

transmitting the one or more pieces of moment-associating information to the one or more subscribers;

wherein the transforming the one or more moment-associating elements includes:

creating a custom language model based at least in part on the event information and the plurality of voiceprints, the custom language model including a context associated with the event;

segmenting the one or more moment-associating elements into a plurality of moment-associating segments based at least in part on the plurality of voiceprints;

assigning a segment speaker for each segment of the plurality of moment-associating segments based at least in part on the plurality of voiceprints;

transcribing the plurality of moment-associating segments into a plurality of transcribed segments based at least in part on the custom language model and the plurality of voiceprints; and

generating the one or more pieces of moment-associating information based at least in part on the plurality of transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.

Assignments (2)
CHANGE OF NAME Recorded Oct 7, 2020
From: AISENSE, INC.
To: OTTER.AI, INC.
Reel/Frame 054006/0966 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2020
From: FU, YUN; LAU, SIMON; NAKAJIMA, KAISUKE; CHENG, JULIUS; LIANG, SAM SONG; ALTREUTER, JAMES MASON; CHIN, KEAN KHEONG; GUPTA, HITESH ANAND; HUANG, XIAOKE; WILLIAMS, BRIAN FRANCIS; XING, TAO; SUN, WEN
To: AISENSE, INC.
Reel/Frame 053983/0735 →
Continuity (1)
Provisional Application 62747001 · Oct 17, 2018
Cited By (17)
US 12,211,508 US 12,229,313 US 12,300,248 US 12,400,661 US 12,406,672 US 12,406,684 US 12,452,126 US 12,456,465 US 12,462,808 US 12,494,929 US 12,518,748 US 12,518,762 US 12,555,581 US 12,609,115 US 12,625,759 US 12,646,506 US 12,689,534