IP Library Granted Patent US 12,581,037
Granted Patent B2
US 12,581,037 · App. 18/470,885 · Granted Mar 17, 2026

Generating smart topics for video calls using a large language model and a context transformer engine

Inventors: Joseph Grillo (San Francisco, CA); Emir Aydin (Vancouver, CA); Ritu Vincent (Alamo, CA); William Adamowicz (Los Angeles, CA)
Assignee: Dropbox, Inc.
H04N7/155G06F3/0484G10L15/183G10L15/22G10L15/30H04L65/1083H04L65/403H04N7/152
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,581,037
App. No.
18/470,885
Granted
Mar 17, 2026
Kind
B2
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing a context transformer engine, a smart topic agent, and a large language model to generate a smart topic output. In particular, in one or more embodiments, the disclosed systems generate a smart topic output from a transcript of a video call. In some embodiments, the disclosed systems provide a smart topic interface that provides the smart topic output on a client device and receives selections of smart topic elements. In one or more embodiments, the disclosed systems generate a combined smart topic from transcripts of video calls in which client devices that participated are associated with a collaborating user account group.

Claims (69)

1 . A computer-implemented method comprising:

obtaining a transcript of a video call that includes video call data captured from one or more client devices interacting in a video call;

receiving, from a client device of the one or more client devices, a selection of a smart topic element corresponding to the transcript of the video call and presented within a smart topic interface displayed on the client device;

in response to the selection of the smart topic element, generating a smart topic output from the transcript of the video call by utilizing a smart topic agent, a context transformer engine, and a large language model; and

providing the smart topic output for display within the smart topic interface presented on the client device.

2 . The computer-implemented method of claim 1 , wherein generating the smart topic output comprises:

identifying, utilizing the context transformer engine and the smart topic agent, a portion of the transcript of the video call comprising subject matter associated with a smart topic; and

providing the portion of the transcript of the video call to the large language model to generate the smart topic output.

3 . The computer-implemented method of claim 1 , further comprising generating, utilizing the context transformer engine, the smart topic agent as executable computer code for performing a process for a smart topic, wherein the smart topic agent is specific to a particular portion of the transcript of the video call.

4 . The computer-implemented method of claim 1 , wherein obtaining the transcript of the video call further comprises:

capturing the video call data from the client device during the video call; and

generating the transcript of the video call from the video call data.

5 . The computer-implemented method of claim 1 , further comprising:

detecting a video call interface initiating the video call on the client device; and

based on detecting the video call interface initiating the video call, providing a smart topic interface element as an overlay on the video call interface, wherein the smart topic interface element presents an indication that the video call data is being captured from the client device during the video call.

6 . The computer-implemented method of claim 1 , further comprising:

receiving, from the client device during the video call, a selection of a highlight option within a smart topic interface element displayed on the client device;

in response to the selection of the highlight option, generating the smart topic output by generating a video call highlight from portions of the video call data corresponding to the highlight option; and

providing the video call highlight for display within the smart topic interface presented on the client device.

7 . The computer-implemented method of claim 1 , further comprising:

receiving, from the client device, a text input indicating subject matter for the smart topic output; and

generating the smart topic output from the transcript of the video call based on the text input.

8 . The computer-implemented method of claim 1 , further comprising:

detecting a mention of a topic during the video call; and

based on detecting the mention of the topic, providing, within a smart topic interface presented on the client device during the video call, a content item suggestion indicating a content item associated with the topic mentioned during the video call.

9 . The computer-implemented method of claim 1 , further comprising:

identifying that a content item stored for a user account associated with the client device corresponds to the smart topic output; and

providing, within the smart topic interface, an option to view the content item together with the smart topic output.

10 . A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computer system to:

obtain a transcript of a video call that includes video call data captured from one or more client devices interacting in a video call;

obtain application data from an application executed by the one or more client devices separately from a video call application;

receive, from a smart topic interface displayed on a client device of the one or more client devices, a selection of a smart topic element corresponding to the transcript of the video call and the application data;

in response to the selection of the smart topic element, generate a smart topic output from the transcript of the video call and the application data by utilizing a smart topic agent, a context transformer engine, and a large language model; and

provide the smart topic output for display within the smart topic interface presented on the client device.

11 . The non-transitory computer-readable medium of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the computer system to:

receive, from the client device, text input indicating subject matter for the smart topic output;

identify, utilizing the context transformer engine and the smart topic agent, a portion of the transcript comprising subject matter associated with a smart topic; and

provide the portion of the transcript to the large language model to generate the smart topic output.

12 . The non-transitory computer-readable medium of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the computer system to generate, from a portion of the transcript and utilizing the context transformer engine, the smart topic agent as executable computer code specific to generating the smart topic output using one or more computing systems for a smart topic indicated by the portion of the transcript.

13 . The non-transitory computer-readable medium of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the computer system to:

compare the transcript of the video call to the application data;

based on comparing the transcript of the video call to the application data, determine that the transcript of the video call and the application data comprise related subject matter; and

based on determining that the transcript of the video call and the application data comprise related subject matter, generate the smart topic output using the transcript of the video call and the application data.

14 . The non-transitory computer-readable medium of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the computer system to:

obtain an additional transcript of an additional video call that includes additional video call data captured from the client device;

identify that the transcript of the video call and the additional transcript of the additional video call comprise related subject matter; and

generate a suggested smart topic from the related subject matter of the transcript and the additional transcript.

15 . The non-transitory computer-readable medium of claim 14 , further comprising instructions that, when executed by the at least one processor, cause the computer system to:

provide the suggested smart topic for display in the smart topic interface;

receive, from the client device, a user selection of the suggested smart topic; and

generate the smart topic output based on receiving the user selection of the suggested smart topic.

16 . A system comprising:

at least one processor; and

at least one non-transitory computer-readable storage medium storing instructions that, when executed by the at least one processor, cause the system to:

obtain a transcript of a video call that includes video call data captured from one or more client devices interacting in a video call;

receive, from a client device of the one or more client devices, a selection of a smart topic element corresponding to the transcript of the video call and presented within a smart topic interface displayed on the client device;

in response to the selection of the smart topic element, generate a smart topic output from the transcript of the video call by utilizing a smart topic agent, a context transformer engine, and a large language model; and

provide the smart topic output for display within the smart topic interface presented on the client device.

17 . The system of claim 16 , further comprising instructions that, when executed by the at least one processor, cause the system to generate the smart topic output by:

identify, utilizing the context transformer engine and the smart topic agent, a portion of the transcript of the video call comprising subject matter associated with a smart topic; and

provide the portion of the transcript of the video call to the large language model to generate the smart topic output.

18 . The system of claim 16 , further comprising instructions that, when executed by the at least one processor, cause the system to:

identify that a content item stored for a user account associated with the client device corresponds to the smart topic output; and

provide, within the smart topic interface, an option to view the content item with the smart topic output.

19 . The system of claim 16 , further comprising instructions that, when executed by the at least one processor, cause the system to provide the smart topic output by providing, for display within the smart topic interface, portions of the transcript of the video call that relate to the smart topic element.

20 . The system of claim 16 , further comprising instructions that, when executed by the at least one processor, cause the system to:

obtain an additional transcript of an additional video call that includes additional video call data captured from the client device;

generate the smart topic output from the transcript and the additional transcript; and

provide the smart topic output for display by displaying portions of the transcript and the additional transcript.

Assignments (2)
SECURITY INTEREST Recorded Dec 12, 2024
From: DROPBOX, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 069604/0611 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2023
From: GRILLO, JOSEPH; AYDIN, EMIR; VINCENT, RITU; ADAMOWICZ, WILLIAM
To: DROPBOX, INC.
Reel/Frame 064971/0694 →
Continuity (2)
Provisional Application 63519437 · Aug 14, 2023
Related Publication 20250061893A1 · Feb 20, 2025
References Cited (30)
US 10798341B1 · Hegde et al. · 2020 [cited by applicant]
US 11119985B1 · Alagianambi et al. · 2021 [cited by applicant]
US 11232266B1 · Biswas · 2022 [cited by examiner]
US 11275891B2 · Mertens et al. · 2022 [cited by applicant]
US 11595459B1 · Olivieri et al. · 2023 [cited by applicant]
US 11645630B2 · Nelson et al. · 2023 [cited by applicant]
US 12216692B1 · Rogynskyy · 2025 [cited by examiner]
US 20090112623A1 · Schoenberg · 2009 [cited by examiner]
US 20090138317A1 · Schoenberg · 2009 [cited by examiner]
US 20170060917A1 · Marsh · 2017 [cited by examiner]
US 20180089880A1 · Garrido et al. · 2018 [cited by applicant]
US 20200403818A1 · Daredia et al. · 2020 [cited by applicant]
US 20210258424A1 · Brown et al. · 2021 [cited by applicant]
US 20210266402A1 · Nagar et al. · 2021 [cited by applicant]
US 20220107953A1 · Goldstein et al. · 2022 [cited by applicant]
US 20220132090A1 · Deole et al. · 2022 [cited by applicant]
US 20220391591A1 · Ronen et al. · 2022 [cited by applicant]
US 20230004713A1 · Broussard et al. · 2023 [cited by applicant]
US 20230092334A1 · Vendrow · 2023 [cited by applicant]
US 20230115098A1 · Miller et al. · 2023 [cited by applicant]
US 20230117301A1 · Olivieri et al. · 2023 [cited by applicant]
US 20230394854A1 · Polavaram · 2023 [cited by examiner]
US 20240176960A1 · Maurer · 2024 [cited by examiner]
US 20250045313A1 · Rogynskyy et al. · 2025 [cited by applicant]
US 20250063140A1 · Grillo · 2025 [cited by examiner]
US 20250217781A1 · Leaman · 2025 [cited by examiner]
Relnotes, et al., “Automatically Summarize Conversations in Teams-Based Collaboration,” Microsoft Learn, Apr. 28, 2023, 2 pages, Retrieved from the Internet: URL: https://learn.microsoft.com/en-us/dynamics365-release-pl… [cited by applicant]
Zoom, “Enabling Zoom IQ Meeting Summary,” Zoom, Jul. 24, 2023, 3 pages, Retrieved from the Internet: URL: https://support.zoom.us/hc/en-us/articles/15088797434893-Enabling-Zoom-IQ-Meeting-Summary. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/470,929 mailed on Jul. 31, 2025, 11 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/470,929, mailed Oct. 10, 2025, 9 pages. [cited by applicant]