IP Library › Granted Patent US 12,394,443
Granted Patent B2
US 12,394,443 · App. 18/346,727 · Granted Aug 19, 2025

Technical architectures for media content editing using machine learning

Inventors: Fan Chen (Los Angeles, CA); Kin Chung Wong (Los Angeles, CA)
Assignee: Lemon Inc.
G11B27/02G06F16/3329G06F40/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,443
App. No.
18/346,727
Granted
Aug 19, 2025
Kind
B2
Abstract

Examples are provided relating to media content editing architectures utilizing machine learning techniques. One aspect includes a method for media content editing, the method comprising: receiving a media content from a user; receiving an editing request for the media content from the user; and editing the media content based on the editing request to generate edited media content by: retrieving a prompt from a prompt pool, wherein the retrieved prompt is selected based on the editing request; parsing the retrieved prompt and the editing request using a large language model to generate one or more editing actions to be performed on the media content; and performing the one or more editing actions on the media content to generate the edited media content.

Claims (59)

1. A method for media content editing, the method comprising:

receiving a media content from a user;

receiving an editing request for the media content from the user through a dialog-assisted editing interface;

editing the media content based on the editing request to generate edited media content by:

retrieving a prompt from a prompt pool, wherein the retrieved prompt is selected based on the editing request;

parsing the retrieved prompt and the editing request using a large language model to generate one or more editing actions to be performed on the media content; and

performing the one or more editing actions on the media content to generate the edited media content;

storing conversation history between the user and the large language model; and

determining from a conversation history that a plurality of rounds of edits performed by the user on the media content have been performed, and in response prompting the user to publish the edited media content.

2. The method of claim 1 , wherein performing the one or more editing actions comprises performing application programming interface calls provided by a back-end tool service comprising a plurality of editing tools, wherein each application programming interface call corresponds to a respective editing tool in the plurality of editing tools.

3. The method of claim 2 , wherein each editing tool in the plurality of editing tools corresponds to one or more prompts in the prompt pool.

4. The method of claim 2 , wherein the plurality of editing tools is organized into a plurality of groupings, and wherein the prompt pool is generated based at least in part on the plurality of groupings.

5. The method of claim 1 , further comprising rendering and displaying the edited media content to the user; and receiving a second editing request.

6. The method of claim 5 , wherein the second editing request comprises a request to revert the performed one or more editing actions.

7. The method of claim 1 , further comprising storing contextual information relating to the editing of the media content.

8. The method of claim 7 , wherein the contextual information comprises one or more of the conversation history, an editing context, or an editing draft history.

9. The method of claim 8 , further comprising refining the prompt pool based on the contextual information.

10. The method of claim 1 , wherein editing the media content further comprises:

providing a dialog reply to the user, wherein the dialog reply is generated by the large language model in response to the retrieved prompt and the editing request; and

receiving a dialog response from the user in response to the dialog reply.

11. A computing device for media content editing, the computing device comprising:

a processor and memory of a computing device, the processor being configured to execute a program using portions of the memory to:

receive a media content from a user:

receive an editing request for the media content from the user through a dialog-assisted editing interface;

edit the media content based on the editing request to generate edited media content by:

retrieving a prompt from a prompt pool, wherein the retrieved prompt is selected based on the editing request;

parsing the retrieved prompt and the editing request using a large language model to generate one or more editing actions to be performed on the media content; and

performing the one or more editing actions on the media content to generate the edited media content;

store conversation history between the user and the large language model; and

determine from a conversation history that a plurality of rounds of edits performed by the user on the media content have been performed, and in response prompt the user to publish the edited media content.

12. The computing device of claim 11 , wherein performing the one or more editing actions comprises performing application programming interface calls provided by a back-end tool service comprising a plurality of editing tools, wherein each application programming interface call corresponds to a respective editing tool in the plurality of editing tools.

13. The computing device of claim 12 , wherein:

each editing tool in the plurality of editing tools corresponds to one or more prompts in the prompt pool;

the plurality of editing tools is organized into a plurality of groupings; and

the prompt pool is generated based at least in part on the plurality of groupings.

14. The computing device of claim 11 , wherein the processor is further configured to store contextual information relating to the editing of the media content, wherein the contextual information comprises one or more of the conversation history, an editing context, or an editing draft history.

15. The computing device of claim 11 , wherein editing the media content further comprises:

providing a dialog reply to the user, wherein the dialog reply is generated by the large language model in response to the retrieved prompt and the editing request; and

receiving a dialog response from the user in response to the dialog reply.

16. A computing system for media content editing, the computing system comprising:

a display;

a back-end tool service comprising a prompt pool, a plurality of editing tools, and a plurality of application programming interfaces, each application programming interface corresponding to an editing tool in the plurality of editing tools;

a processor and memory of a computing device, the processor being configured to execute a program using portions of the memory to:

receive a media content from a user:

receive an editing request for the media content from the user through a dialog-assisted editing interface;

edit the media content based on the editing request to generate edited media content by:

retrieving a prompt from the prompt pool, wherein the retrieved prompt is selected based on the editing request;

parsing the retrieved prompt and the editing request using one or more large language models to generate one or more editing actions to be performed on the media content; and

performing the one or more editing actions on the media content by calling at least one application programming interface in the plurality of application programming interfaces to generate the edited media content;

render and display the edited media content using the display through the dialog-assisted editing interface;

store conversation history between the user and the large language model; and

determine from a conversation history that a plurality of rounds of edits performed by the user on the media content have been performed, and in response prompt the user to publish the edited media content.

17. The computing system of claim 16 , wherein the one or more large language models comprises a plurality of large language models, each trained for at least one task, and wherein the processor is configured to select a large language model from the plurality of large language models to parse the retrieved prompt and the editing request.

18. The computing system of claim 16 , wherein:

each editing tool in the plurality of editing tools corresponds to one or more prompts in the prompt pool;

the plurality of editing tools is organized into a plurality of groupings; and

prompts in the prompt pool are generated based at least in part on the plurality of groupings.

19. The computing system of claim 16 , wherein the processor is further configured to store contextual information relating to the editing of the media content, wherein the contextual information comprises one or more of the conversation history, an editing context, or an editing draft history.

20. A non-transitory computer readable medium for media content editing, the non-transitory computer readable medium comprising instructions that, when executed by a computing device, cause the computing device to implement the method of claim 1 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2025
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 071211/0037 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2025
From: CHEN, FAN; WONG, KIN CHUNG
To: BYTEDANCE INC.
Reel/Frame 071358/0905 →
Continuity (1)
Related Publication 20250014605A1 · Jan 9, 2025
References Cited (46)
US 11960500B1 · Sboychakova et al. · 2024 [cited by applicant]
US 20080134282A1 · Fridman et al. · 2008 [cited by applicant]
US 20170228384A1 · Caruso et al. · 2017 [cited by applicant]
US 20190108276A1 · Kovács et al. · 2019 [cited by applicant]
US 20190163728A1 · Koren et al. · 2019 [cited by applicant]
US 20190179895A1 · Bhatt et al. · 2019 [cited by applicant]
US 20190196698A1 · Cohen et al. · 2019 [cited by applicant]
US 20190259293A1 · Hellman · 2019 [cited by examiner]
US 20200005157A1 · Rosenstein · 2020 [cited by examiner]
US 20200066261A1 · Ganz et al. · 2020 [cited by applicant]
US 20210027065A1 · Chung et al. · 2021 [cited by applicant]
US 20210157618A1 · Moon · 2021 [cited by examiner]
US 20210272599A1 · Patterson et al. · 2021 [cited by applicant]
US 20220108726A1 · Li et al. · 2022 [cited by applicant]
US 20220229832A1 · Li · 2022 [cited by examiner]
US 20230042221A1 · Xu et al. · 2023 [cited by applicant]
US 20230055241A1 · Zionpour · 2023 [cited by examiner]
US 20230069133A1 · Matsuoka · 2023 [cited by examiner]
US 20230074406A1 · Baeuml et al. · 2023 [cited by applicant]
US 20230135179A1 · Mielke · 2023 [cited by examiner]
US 20230141807A1 · Groenewegen et al. · 2023 [cited by applicant]
US 20230244506A1 · Geller · 2023 [cited by examiner]
US 20240038226A1 · Nouri · 2024 [cited by examiner]
US 20240095077A1 · Singh · 2024 [cited by examiner]
US 20240126997A1 · Bent, III et al. · 2024 [cited by applicant]
US 20240177739A1 · Su et al. · 2024 [cited by applicant]
US 20240179380A1 · Hannan et al. · 2024 [cited by applicant]
US 20240211439A1 · Shah et al. · 2024 [cited by applicant]
US 20240256773A1 · Correia Ribeiro et al. · 2024 [cited by applicant]
US 20240394502A1 · Sami et al. · 2024 [cited by applicant]
US 20250005051A1 · Khosla et al. · 2025 [cited by applicant]
CN 107685824A · 2018 [cited by applicant]
CN 110389796A · 2019 [cited by applicant]
CN 111726676A · 2020 [cited by applicant]
CN 114430499A · 2022 [cited by applicant]
CN 115129212A · 2022 [cited by applicant]
CN 116187282A · 2023 [cited by applicant]
EP 4354887A1 · 2024 [cited by applicant]
JP 2019109106A · 2019 [cited by applicant]
WO 2022260188A1 · 2022 [cited by applicant]
Chang, Y. et al., “A Survey on Evaluation of Large Language Models,” ACM Transactions on Intelligent Systems and Technology, vol. 15, No. 3, Mar. 29, 2024, 45 pages. [cited by applicant]
ISA Intellectual Property Office of Singapore, International Search Report Issued in Application No. PCT/SG2024/050423, Sep. 24, 2024, WIPO, 4 pages. [cited by applicant]
ISA Intellectual Property Office of Singapore, International Search Report Issued in Application No. PCT/SG2024/050424, Sep. 20, 2024, WIPO, 3 pages. [cited by applicant]
ISA Intellectual Property Office of Singapore, International Search Report Issued in Application No. PCT/SG2024/050425, Jul. 11, 2024, WIPO, 3 pages. [cited by applicant]
ISA Intellectual Property Office of Singapore, International Search Report Issued in Application No. PCT/SG2024/050427, Aug. 23, 2024, WIPO, 4 pages. [cited by applicant]
Teubner, T. et al., “Welcome to the Era of ChatGPT et al.: The Prospects of Large Language Models,” Business & Information Systems Engineering, vol. 65, No. 2, Mar. 13, 2023, 7 pages. [cited by applicant]