IP Library Granted Patent US 11,869,508
Granted Patent B2
US 11,869,508 · App. 17/242,465 · Granted Jan 9, 2024

Systems and methods for capturing, processing, and rendering one or more context-aware moment-associating elements

Inventors: Yun Fu (Cupertino, CA); Simon Lau (San Jose, CA); Kaisuke Nakajima (Sunnyvale, CA); Julius Cheng (Cupertino, CA); Sam Song Liang (Palo Alto, CA); James Mason Altreuter (Belmont, CA); Kean Kheong Chin (Santa Clara, CA); Zhenhao Ge (Sunnyvale, CA); Hitesh Anand Gupta (Santa Clara, CA); Xiaoke Huang (Foster City, CA); James Francis McAteer (San Francisco, CA); Brian Francis Williams (San Carlos, CA); Tao Xing (San Jose, CA)
Assignee: Otter.ai, Inc.
G10L15/26G06F16/906G10L15/04G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,869,508
App. No.
17/242,465
Granted
Jan 9, 2024
Kind
B2
Abstract

Computer-implemented method and system for receiving and processing one or more moment-associating elements. For example, the computer-implemented method includes receiving the one or more moment-associating elements, transforming the one or more moment-associating elements into one or more pieces of moment-associating information, and transmitting at least one piece of the one or more pieces of moment-associating information. The transforming the one or more moment-associating elements into one or more pieces of moment-associating information includes segmenting the one or more moment-associating elements into a plurality of moment-associating segments, assigning a segment speaker for each segment of the plurality of moment-associating segments, transcribing the plurality of moment-associating segments into a plurality of transcribed segments, and generating the one or more pieces of moment-associating information based on at least the plurality of transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.

Claims (66)

1. A computer-implemented method for receiving and processing a plurality of moment-associating elements, the method comprising:

receiving the plurality of moment-associating elements, the plurality of moment-associating elements including a plurality of audio elements;

transforming the plurality of moment-associating elements into one or more pieces of moment-associating information by:

segmenting the plurality of audio elements into a plurality of audio segments;

assigning a segment speaker for each audio segment of the plurality of audio segments;

transcribing the plurality of audio segments into a plurality of text elements by:

transcribing a first element of the plurality of audio segments into a first text element of the plurality of text elements;

transcribing a second element of the plurality of audio segments into a second text element of the plurality of text elements;

identifying one or more incorrectly transcribed words in the first transcribed text element using the second transcribed text element; and

updating the first transcribed text element by correcting the one or more incorrectly transcribed words in the first transcribed text element using one or more transcribed words in the second transcribed text element; and

generating the one or more pieces of moment-associating information based on at least the plurality of audio segments, the segment speaker assigned for the each audio segment of the plurality of audio segments, and the plurality of text elements; and

transmitting at least one piece of the one or more pieces of moment-associating information.

2. The computer-implemented method of claim 1 , wherein the receiving the plurality of moment-associating elements includes assigning a timestamp associated with each element of the plurality of moment-associating elements.

3. The computer-implemented method of claim 1 , wherein the receiving the plurality of moment-associating elements includes receiving one or more visual elements or receiving one or more environmental elements.

4. The computer-implemented method of claim 3 , wherein the one or more visual elements includes at least one selected from a group consisting of pictures, images, screenshots, video frames, one or more projections, and holograms.

5. The computer-implemented method of claim 3 , wherein the one or more environmental elements includes at least one selected from a group consisting of global positions, location types, and conditions associated with the one or more environmental elements.

6. The computer-implemented method of claim 3 , wherein the one or more environmental elements includes at least one selected from a group consisting of a longitude, a latitude, an altitude, a country, a city, a street, a location type, a temperature, a humidity, a movement, a velocity of a movement, a direction of a movement, an ambient noise level, and one or more echo properties.

7. The computer-implemented method of claim 1 , wherein the plurality of audio elements includes one or more voice elements of one or more voice-generating sources.

8. The computer-implemented method of claim 1 , further comprising:

receiving one or more voice elements of one or more voice-generating sources; and

receiving one or more voiceprints corresponding to the one or more voice-generating sources respectively.

9. The computer-implemented method of claim 8 , further comprising:

segmenting the plurality of audio elements into a plurality of audio segments includes segmenting the plurality of audio elements into the plurality of audio segments based on at least the one or more voiceprints;

wherein the assigning a segment speaker for each audio segment of the plurality of audio segments includes assigning the segment speaker for the each segment of the plurality of audio segments based on at least the one or more voiceprints; and

wherein the transcribing the plurality of audio segments into a plurality of text elements includes transcribing the plurality of audio segments into the plurality of text elements based on at least the one or more voiceprints.

10. The computer-implemented method of claim 8 , wherein the receiving one or more voiceprints corresponding to the one or more voice-generating sources respectively includes receiving one or more acoustic models and/or language models corresponding to the one or more voice-generating sources respectively.

11. The computer-implemented method of claim 1 , wherein the transcribing the plurality of audio segments into a plurality of text elements includes:

determining one or more speaker-change timestamps, wherein each timestamp of the one or more speaker-change timestamps corresponds to a timestamp when a speaker change occurs, a sentence change occurs, and/or a topic change occurs.

12. The computer-implemented method of claim 11 , wherein the segmenting the plurality of audio elements into a plurality of audio segments includes segmenting the plurality of audio elements into a plurality of audio segments based on at least one selected from a group consisting of the one or more speaker-change timestamps, the one or more sentence-change timestamps, and the one or more topic-change timestamps.

13. The computer-implemented method of claim 1 , further comprising:

establishing one or more anchor points based on at least the plurality of moment-associating elements;

wherein:

the one or more anchor points correspond to one or more timestamps respectively; and

each anchor point of the one or more anchor points is navigable, searchable, or both navigable and searchable.

14. The computer-implemented method of claim 13 , further comprising using the one or more anchor points to navigate the one or more pieces of moment-associating information based on at least the one or more timestamps.

15. The computer-implemented method of claim 13 , wherein the one or more anchor points include at least one selected from a group consisting of a word, a phrase, a photo, and a screenshot.

16. The computer-implemented method of claim 1 , further comprising:

obtaining one or more moment-associating photos, the one or more moment-associating photos being one or more parts of the plurality of moment-associating elements; and

transforming the one or more moment-associating photos into one or more anchor photos, wherein the one or more anchor photos correspond to one or more timestamps respectively.

17. The computer-implemented method of claim 16 wherein each anchor photo of the one or more anchor photos is navigable, searchable, or both navigable and searchable.

18. The computer-implemented method of claim 17 , further comprising:

using the one or more anchor photos to navigate the one or more pieces of moment-associating information based on at least the one or more timestamps.

19. A system for receiving and processing a plurality of moment-associating elements, the system comprising:

a receiving module configured to receive the plurality of moment-associating elements, the plurality of moment-associating elements including a plurality of audio elements;

a transforming module configured to transform the plurality of moment-associating elements into one or more pieces of moment-associating information by:

segmenting the plurality of audio elements into a plurality of audio segments;

assigning a segment speaker for each audio segment of the plurality of audio segments;

transcribing the plurality of audio segments into a plurality of text elements by:

transcribing a first element of the plurality of audio segments into a first text element of the plurality of text elements;

transcribing a second element of the plurality of audio segments into a second text element of the plurality of text elements;

identifying one or more incorrectly transcribed words in the first transcribed text element using the second transcribed text element; and

updating the first transcribed text element by correcting the one or more incorrectly transcribed words in the first transcribed text element using one or more transcribed words in the second transcribed text element; and

generating the one or more pieces of moment-associating information based on at least the plurality of audio segments, the segment speaker assigned for the each audio segment of the plurality of audio segments, and the plurality of text elements; and

a transmitting module configured to transmit at least one piece of the one or more pieces of moment-associating information.

20. A non-transitory computer-readable medium with instructions stored thereon, that when executed by a processor, perform processes comprising:

receiving a plurality of moment-associating elements, the plurality of moment-associating elements including a plurality of audio elements;

transforming the plurality of moment-associating elements into one or more pieces of moment-associating information by:

segmenting the plurality of audio elements into a plurality of audio segments;

assigning a segment speaker for each audio segment of the plurality of audio segments;

transcribing the plurality of audio segments into a plurality of text elements by:

transcribing a first element of the plurality of audio segments into a first text element of the plurality of text elements,

transcribing a second element of the plurality of audio segments into a second text element of the plurality of text elements,

identifying one or more incorrectly transcribed words in the first transcribed text element using the second transcribed text element; and

updating the first transcribed text element by correcting the one or more incorrectly transcribed words in the first transcribed text element using one or more transcribed words in the second transcribed text element; and

generating the one or more pieces of moment-associating information based on at least the plurality of audio segments, the segment speaker assigned for the each audio segment of the plurality of audio segments, and the plurality of text elements; and

transmitting at least one piece of the one or more pieces of moment-associating information.

Assignments (2)
CHANGE OF NAME Recorded Aug 11, 2021
From: AISENSE INC.
To: OTTER.AI, INC.
Reel/Frame 057159/0631 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2021
From: FU, YUN; LAU, SIMON; NAKAJIMA, KAISUKE; CHENG, JULIUS; LIANG, SAM SONG; ALTREUTER, JAMES MASON; CHIN, KEAN KHEONG; GE, ZHENHAO; GUPTA, HITESH ANAND; HUANG, XIAOKE; MCATEER, JAMES FRANCIS; WILLIAMS, BRIAN FRANCIS; XING, TAO
To: AISENSE, INC.
Reel/Frame 057140/0257 →
Continuity (9)
Continuation 16403263 · May 3, 2019
Continuation In Part 16027511 · Jul 5, 2018
Continuation In Part 16276446 · Feb 14, 2019
Continuation In Part 16027511 · Jul 5, 2018
Provisional Application 62668623 · May 8, 2018
Provisional Application 62530227 · Jul 9, 2017
Provisional Application 62710631 · Feb 16, 2018
Provisional Application 62631680 · Feb 17, 2018
Related Publication 20210319797A1 · Oct 14, 2021
Cited By (8)
US 12,400,661 US 12,406,672 US 12,406,684 US 12,456,465 US 12,462,808 US 12,494,929 US 12,518,748 US 12,555,581