IP Library Granted Patent US 11,024,316
Granted Patent B1
US 11,024,316 · App. 16/403,263 · Granted Jun 1, 2021

Systems and methods for capturing, processing, and rendering one or more context-aware moment-associating elements

Inventors: Yun Fu (Cupertino, CA); Simon Lau (San Jose, CA); Kaisuke Nakajima (Sunnyvale, CA); Julius Cheng (Cupertino, CA); Sam Song Liang (Palo Alto, CA); James Mason Altreuter (Belmont, CA); Kean Kheong Chin (Santa Clara, CA); Zhenhao Ge (Sunnyvale, CA); Hitesh Anand Gupta (Santa Clara, CA); Xiaoke Huang (Foster City, CA); James Francis McAteer (San Francisco, CA); Brian Francis Williams (San Carlos, CA); Tao Xing (San Jose, CA)
Assignee: Otter.ai, Inc.
G10L15/26G06F16/906G10L15/04G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,024,316
App. No.
16/403,263
Granted
Jun 1, 2021
Kind
B1
Abstract

Computer-implemented method and system for receiving and processing one or more moment-associating elements. For example, the computer-implemented method includes receiving the one or more moment-associating elements, transforming the one or more moment-associating elements into one or more pieces of moment-associating information, and transmitting at least one piece of the one or more pieces of moment-associating information. The transforming the one or more moment-associating elements into one or more pieces of moment-associating information includes segmenting the one or more moment-associating elements into a plurality of moment-associating segments, assigning a segment speaker for each segment of the plurality of moment-associating segments, transcribing the plurality of moment-associating segments into a plurality of transcribed segments, and generating the one or more pieces of moment-associating information based on at least the plurality of transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.

Claims (80)

1. A computer-implemented method for receiving and processing a plurality of moment-associating elements, the method comprising:

receiving the plurality of moment-associating elements, the plurality of moment-associating elements including a plurality of audio elements;

transforming the plurality of moment-associating elements into one or more pieces of moment-associating information; and

transmitting at least one piece of the one or more pieces of moment-associating information;

wherein the transforming the plurality of moment-associating elements into one or more pieces of moment-associating information includes:

transcribing the plurality of moment-associating elements into a plurality of transcribed elements by at least:

transcribing the plurality of audio elements into a plurality of text elements; and

transcribing two or more audio elements of the plurality of audio elements in conjunction with each other to extrapolate one or more tones corresponding to the plurality of text elements;

segmenting the plurality of moment-associating elements into a plurality of moment-associating segments by at least segmenting the plurality of audio elements into a plurality of audio segments where any change in a text-corresponding tone occurs;

assigning a segment speaker for each segment of the plurality of moment-associating segments by at least assigning a segment speaker for each segment of the plurality of audio segments;

and

generating the one or more pieces of moment-associating information based on at least the plurality of moment-associating segments and the segment speaker assigned for the each segment of the plurality of moment-associating segments.

2. The computer-implemented method of claim 1 wherein the receiving the plurality of moment-associating elements includes assigning a timestamp associated with each element of the plurality of moment-associating elements.

3. The computer-implemented method of claim 1 wherein the receiving the plurality of moment-associating elements includes receiving one or more visual elements or receiving one or more environmental elements.

4. The computer-implemented method of claim 1 wherein the plurality of audio elements includes at least one selected from a group consisting of one or more voice elements of one or more voice-generating sources and one or more ambient sound elements.

5. The computer-implemented method of claim 3 wherein the receiving one or more visual elements includes at least one selected from a group consisting of receiving one or more pictures, receiving one or more images, receiving one or more screenshots, receiving one or more video frames, receiving one or more projections, and receiving one or more holograms.

6. The computer-implemented method of claim 3 wherein the receiving one or more environmental elements includes at least one selected from a group consisting of receiving one or more global positions, receiving one or more location types, and receiving one or more moment conditions.

7. The computer-implemented method of claim 3 wherein the receiving one or more environmental elements includes at least one selected from a group consisting of receiving a longitude, receiving a latitude, receiving an altitude, receiving a country, receiving a city, receiving a street, receiving a location type, receiving a temperature, receiving a humidity, receiving a movement, receiving a velocity of a movement, receiving a direction of a movement, receiving an ambient noise level, and receiving one or more echo properties.

8. The computer-implemented method of claim 1 , and further comprising:

receiving one or more voice elements of one or more voice-generating sources; and

receiving one or more voiceprints corresponding to the one or more voice-generating sources respectively.

9. The computer-implemented method of claim 8 wherein the transforming the plurality of moment-associating elements into one or more pieces of moment-associating information includes at least one selected from a group consisting of:

transcribing the plurality of moment-associating elements into the plurality of transcribed elements based on at least the one or more voiceprints;

segmenting the plurality of moment-associating elements into the plurality of moment-associating segments based on at least the one or more voiceprints; and

assigning the segment speaker for the each segment of the plurality of moment-associating segments based on at least the one or more voiceprints.

10. The computer-implemented method of claim 8 wherein the receiving one or more voiceprints corresponding to the one or more voice-generating sources respectively includes at least one selected from a group consisting of:

receiving one or more acoustic models corresponding to the one or more voice-generating sources respectively; and

receiving one or more language models corresponding to the one or more voice-generating sources respectively.

11. The computer-implemented method of claim 1 , wherein the transcribing the plurality of moment-associating elements into a plurality of transcribed elements includes:

transcribing a first element of the plurality of moment-associating elements into a first transcribed element of the plurality of transcribed elements;

transcribing a second element of the plurality of moment-associating elements into a second transcribed element of the plurality of transcribed elements; and

correcting the first transcribed element based on at least the second transcribed element.

12. The computer-implemented method of claim 1 wherein the transcribing the plurality of moment-associating elements into a plurality of transcribed elements further includes at least one selected from a group consisting of:

determining one or more speaker-change timestamps, each timestamp of the one or more speaker-change timestamps corresponding to a timestamp when a speaker change occurs;

determining one or more sentence-change timestamps, each timestamp of the one or more sentence-change timestamps corresponding to a timestamp when a sentence change occurs; and

determining one or more topic-change timestamps, each timestamp of the one or more topic-change timestamps corresponding to a timestamp when a topic change occurs.

13. The computer-implemented method of claim 12 wherein the segmenting the plurality of moment-associating elements into a plurality of moment-associating segments is performed based on at least one selected from a group consisting of:

the one or more speaker-change timestamps;

the one or more sentence-change timestamps; and

the one or more topic-change timestamps.

14. The computer-implemented method of claim 1 , and further comprising:

establishing one or more anchor points based on at least the plurality of moment-associating elements;

wherein:

the one or more anchor points correspond to one or more timestamps respectively; and

each anchor point of the one or more anchor points is navigable, searchable, or both navigable and searchable.

15. The computer-implemented method of claim 14 , and further comprising:

using the one or more anchor points to navigate the one or more pieces of moment-associating information based on at least the one or more timestamps.

16. The computer-implemented method of claim 14 , wherein the one or more anchor points include at least one selected from a group consisting of a word, a phrase, a photo, and a screenshot.

17. The computer-implemented method of claim 1 , and further comprising:

obtaining one or more moment-associating photos, the one or more moment-associating photos being one or more parts of the plurality of moment-associating elements;

wherein:

the transforming the plurality of moment-associating elements into one or more pieces of moment-associating information includes transforming the one or more moment-associating photos into one or more anchor photos; and

the one or more anchor photos correspond to one or more timestamps respectively.

18. The computer-implemented method of claim 17 wherein each anchor photo of the one or more anchor photos is navigable, searchable, or both navigable and searchable.

19. The computer-implemented method of claim 18 , and further comprising:

using the one or more anchor photos to navigate the one or more pieces of moment-associating information based on at least the one or more timestamps.

20. A system for receiving and processing a plurality of moment-associating elements, the system comprising:

a receiving module configured to receive the plurality of moment-associating elements, the plurality of moment-associating elements including a plurality of audio elements;

a transforming module configured to transform the plurality of moment-associating elements into one or more pieces of moment-associating information; and

a transmitting module configured to transmit at least one piece of the one or more pieces of moment-associating information;

wherein the transforming module is further configured to:

transcribe the plurality of moment-associating elements into a plurality of transcribed elements by at least:

transcribing the plurality of audio elements into a plurality of text elements; and

transcribing two or more audio elements of the plurality of audio elements in conjunction with each other to extrapolate one or more tones corresponding to the plurality of text elements;

segment the plurality of moment-associating elements into a plurality of moment-associating segments by at least segmenting the plurality of audio elements into a plurality of audio segments where any change in a text-corresponding tone occurs;

assign a segment speaker for each segment of the plurality of moment-associating segments by at least assigning a segment speaker for each segment of the plurality of audio segments;

and

generate the one or more pieces of moment-associating information based on at least the plurality of moment-associating segments and the segment speaker assigned for the each segment of the plurality of moment-associating segments.

21. A non-transitory computer-readable medium with instructions stored thereon, that when executed by a processor, perform processes comprising:

receiving the plurality of moment-associating elements, the plurality of moment-associating elements including a plurality of audio elements;

transforming the plurality of moment-associating elements into one or more pieces of moment-associating information; and

transmitting at least one piece of the one or more pieces of moment-associating information;

wherein the transforming the plurality of moment-associating elements into one or more pieces of moment-associating information includes:

transcribing the plurality of moment-associating elements into a plurality of transcribed elements by at least:

transcribing the plurality of audio elements into a plurality of text elements; and

transcribing two or more audio elements of the plurality of audio elements in conjunction with each other to extrapolate one or more tones corresponding to the plurality of text elements;

segmenting the plurality of moment-associating elements into a plurality of moment-associating segments by at least segmenting the plurality of audio elements into a plurality of audio segments where any change in a text-corresponding tone occurs;

assigning a segment speaker for each segment of the plurality of moment-associating segments by at least assigning a segment speaker for each segment of the plurality of audio segments;

and

generating the one or more pieces of moment-associating information based on at least the plurality of moment-associating segments and the segment speaker assigned for the each segment of the plurality of moment-associating segments.

Assignments (2)
CHANGE OF NAME Recorded Oct 7, 2020
From: AISENSE, INC.
To: OTTER.AI, INC.
Reel/Frame 054006/0966 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2020
From: FU, YUN; LAU, SIMON; NAKAJIMA, KAISUKE; CHENG, JULIUS; LIANG, SAM SONG; ALTREUTER, JAMES MASON; CHIN, KEAN KHEONG; GE, ZHENHAO; GUPTA, HITESH ANAND; HUANG, XIAOKE; MCATEER, JAMES FRANCIS; WILLIAMS, BRIAN FRANCIS; XING, TAO
To: AISENSE, INC.
Reel/Frame 053983/0467 →
Continuity (9)
Continuation In Part 16027511 · Jul 5, 2018
Continuation In Part 16276446 · Feb 14, 2019
Continuation In Part 16027511 · Jul 5, 2018
Provisional Application 62668623 · May 8, 2018
Provisional Application 62530227 · Jul 9, 2017
Provisional Application 62710631 · Feb 16, 2018
Provisional Application 62631680 · Feb 17, 2018
Provisional Application 62668623 · May 8, 2018
Provisional Application 62530227 · Jul 9, 2017
Cited By (8)
US 12,400,661 US 12,406,672 US 12,406,684 US 12,456,465 US 12,462,808 US 12,494,929 US 12,518,748 US 12,555,581