IP Library Granted Patent US 11,488,602
Granted Patent B2
US 11,488,602 · App. 15/900,414 · Granted Nov 1, 2022

Meeting transcription using custom lexicons based on document history

Inventors: Timo Mertens (Millbrae, CA); Bradley Neuberg (San Francisco, CA)
Assignee: Dropbox, Inc.
G10L15/26G06F40/166G06F40/169G06Q10/103G10L15/183G10L15/197
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,488,602
App. No.
15/900,414
Granted
Nov 1, 2022
Kind
B2
Abstract

A collaborative content management system allows multiple users to access and modify collaborative documents. When audio data is recorded by or uploaded to the system, the audio data may be transcribed or summarized to improve accessibility and user efficiency. Text transcriptions are associated with portions of the audio data representative of the text, and users can search the text transcription and access the portions of the audio data corresponding to search queries for playback. An outline can be automatically generated based on a text transcription of audio data and embedded as a modifiable object within a collaborative document. The system associates hot words with actions to modify the collaborative document upon identifying the hot words in the audio data. Collaborative content management systems can also generate custom lexicons for users based on documents associated with the user for use in transcribing audio data, ensuring that text transcription is more accurate.

Claims (46)

1. A computer-implemented method comprising:

receiving, by a content creation system, audio data including speech of a plurality of speakers, each speaker of the plurality of speakers included in the audio data, the audio data corresponding to a collaborative document;

responsive to receiving the audio data, accessing a custom lexicon that is generated by:

identifying, by the content creation system, a respective account for each speaker of the plurality of speakers, each respective account associated with one or more documents stored by the content creation system, each document of the one or more documents associated with information identifying a set of accounts as collaborators having accessed the document,

determining, by the content creation system, a plurality of collaborative documents on which each speaker is a collaborator based on the information identifying the set of accounts as collaborators, each collaborative document of the plurality of collaborative documents having been determined for inclusion in the plurality of collaborative documents based on it being stored in association with accounts of all of the plurality of speakers, the plurality of collaborative documents each having one or more content level comments that are annotated in visual association with one or more text spans but displayed in an interface separate from the one or more text spans, and

generating, by the content creation system for the plurality of speakers, the custom lexicon based on the plurality of collaborative documents on which each speaker is a collaborator;

transcribing, by the content creation system, the audio data into text representative of the speech using the custom lexicon; and

modifying, by the content creation system, the collaboration document to include the text representative of the speech.

2. The computer-implemented method of claim 1 , wherein generating the custom lexicon comprises:

identifying, by the content creation system, a set of words or n-grams included within the plurality of collaborative documents; and

modifying, by the content creation system, a default lexicon to include the identified set of words or n-grams to generate the custom lexicon.

3. The computer-implemented method of claim 1 , wherein generating the custom lexicon for the plurality of speakers comprises selecting among a plurality of lexicons associated with each speaker of the plurality of speakers based on a subject matter of the audio data.

4. The computer-implemented method of claim 1 , wherein generating the custom lexicon for the plurality of speakers comprises selecting among a plurality of lexicons associated with each speaker of the plurality of speakers based on one or more characteristics of the one or more speakers.

5. The computer-implemented method of claim 1 , wherein the set of documents is accessible to each of the plurality of speakers.

6. The computer-implemented method of claim 1 , wherein each document of the plurality of collaborative documents is stored within an account of one or more of the plurality of speakers.

7. The computer-implemented method of claim 1 , wherein at least one document of the plurality of collaborative documents comprises the collaborative document.

8. The computer-implemented method of claim 1 , wherein the set of documents is associated with a subject matter of the audio data.

9. The computer-implemented method of claim 1 , wherein the set of documents is selected by a speaker of the plurality of speakers.

10. The computer-implemented method of claim 1 , wherein the audio data is captured during a meeting.

11. The computer-implemented method of claim 10 , wherein the set of documents is selected based on a characteristic of the meeting.

12. The computer-implemented method of claim 1 , wherein a second custom lexicon is accessed in response to the custom lexicon not including text associated with a spoken word.

13. The computer-implemented method of claim 12 , wherein the second custom lexicon is generated based on a second set of documents.

14. A system comprising:

one or more processors; and

a non-transitory computer-readable storage medium storing executable instructions that, when executed by the one or more processors, cause the system to perform steps comprising:

receiving, by a content creation system, audio data including speech of a plurality of speakers, each speaker of the plurality of speakers included in the audio data, the audio data corresponding to a collaborative document;

responsive to receiving the audio data, accessing a custom lexicon that is generated by:

identifying, by the content creation system, a respective account for each speaker of the plurality of speakers, each respective account associated with one or more documents stored by the content creation system, each document of the one or more documents associated with information identifying a set of accounts as collaborators having accessed the document,

determining, by the content creation system, a plurality of collaborative documents on which each speaker is a collaborator based on the information identifying the set of accounts as collaborators, each collaborative document of the plurality of collaborative documents having been determined for inclusion in the plurality of collaborative documents based on it being stored in association with accounts of all of the plurality of speakers, the plurality of collaborative documents each having one or more content level comments that are annotated in visual association with one or more text spans but displayed in an interface separate from the one or more text spans, and

generating, by the content creation system for the plurality of speakers, the custom lexicon based on the plurality of collaborative on which each speaker is a collaborator;

transcribing, by the content creation system, the audio data into text representative of the speech using the custom lexicon; and

modifying, by the content creation system, the collaboration document to include the text representative of the speech.

15. The system of claim 14 , wherein the custom lexicon includes a set of terms included within the plurality of collaborative documents and not included within a default lexicon.

16. A non-transitory computer-readable medium comprising memory with instructions encoded thereon that, when executed, cause one or more processors to perform operations, the instructions comprising instructions to:

receive, by a content creation system, audio data including speech of a plurality of speakers, each speaker of the plurality of speakers included in the audio data, the audio data corresponding to a collaborative document;

responsive to receiving the audio data, access a custom lexicon that is generated by:

identifying, by the content creation system, a respective account for each speaker of the plurality of speakers, each respective account associated with one or more documents stored by the content creation system, each document of the one or more documents associated with information identifying a set of accounts as collaborators having accessed the document,

determining, by the content creation system, a plurality of collaborative documents on which each speaker is a collaborator based on the information identifying the set of accounts as collaborators, each collaborative document of the plurality of collaborative documents having been determined for inclusion in the plurality of collaborative documents based on it being stored in association with accounts of all of the plurality of speakers, the plurality of collaborative documents each having one or more content level comments that are annotated in visual association with one or more text spans but displayed in an interface separate from the one or more text spans, and

generating, by the content creation system for the plurality of speakers, the custom lexicon based on the plurality of collaborative documents on which each speaker is a collaborator;

transcribe, by the content creation system, the audio data into text representative of the speech using the custom lexicon; and

modify, by the content creation system, the collaboration document to include the text representative of the speech.

17. The non-transitory computer-readable medium of claim 16 , wherein the instructions to generate the custom lexicon comprise instructions to:

identify, by the content creation system, a set of words or n-grams included within the plurality of collaborative documents; and

modify, by the content creation system, a default lexicon to include the identified set of words or n-grams to generate the custom lexicon.

18. The non-transitory computer-readable medium of claim 16 , wherein the instructions to generate the custom lexicon for the plurality of speakers comprise instructions to select among a plurality of lexicons associated with each speaker of the plurality of speakers based on a subject matter of the audio data.

19. The non-transitory computer-readable medium of claim 16 , wherein the instructions to generate the custom lexicon for the plurality of speakers comprise instructions to select among a plurality of lexicons associated with each speaker of the plurality of speakers based on one or more characteristics of the one or more speakers.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Dec 13, 2024
From: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
To: DROPBOX, INC.
Reel/Frame 069635/0332 →
SECURITY INTEREST Recorded Dec 12, 2024
From: DROPBOX, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 069604/0611 →
PATENT SECURITY AGREEMENT Recorded Mar 10, 2021
From: DROPBOX, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 055670/0219 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2018
From: MERTENS, TIMO; NEUBERG, BRADLEY
To: DROPBOX, INC.
Reel/Frame 045139/0302 →
Continuity (1)
Related Publication 20190259387A1 · Aug 22, 2019