IP Library › Granted Patent US 12,277,400
Granted Patent B1
US 12,277,400 · App. 18/590,498 · Granted Apr 15, 2025

Multimedia content management for large language model(s) and/or other generative model(s)

Inventors: Sanil Jain (Sunnyvale, CA); Wei Yu (Mountain View, CA); Ágoston Weisz (Zurich, CH); Michael Andrew Goodman (Oakland, CA); Diana Avram (Zurich, CH); Amin Ghafouri (San Francisco, CA); Golnaz Ghiasi (Mountain View, CA); Igor Petrovski (Zurich, CH); Khyatti Gupta (Zurich, CH); Oscar Akerlund (Zurich, CH); Evgeny Sluzhaev (Zurich, CH); Rakesh Shivanna (Sunnyvale, CA); Thang Luong (Santa Clara, CA); Komal Singh (Kitchener, CA); Yifeng Lu (Mountain View, CA); Vikas Peswani (Mountain View, CA)
Assignee: GOOGLE LLC
G06F40/40G06V10/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,400
App. No.
18/590,498
Granted
Apr 15, 2025
Kind
B1
Abstract

Implementations relate to managing multimedia content that is obtained by large language model(s) (LLM(s)) and/or generated by other generative model(s). Processor(s) of a system can: receive natural language (NL) based input that requests multimedia content, generate a response that is responsive to the NL based input, and cause the response to be rendered. In some implementations, and in generating the response, the processor(s) can process, using a LLM, LLM input to generate LLM output, and determine, based on the LLM output, at least multimedia content to be included in the response. Further, the processor(s) can evaluate the multimedia content to determine whether it should be included in the response. In response to determining that the multimedia content should not be included in the response, the processor(s) can cause the response, including alternative multimedia content or other textual content, to be rendered.

Claims (58)

1. A method implemented by one or more processors, the method comprising:

receiving natural language (NL) based input associated with a client device of a user, the NL based input requesting multimedia content;

generating a response that is responsive to the NL based input, wherein generating the response that is responsive to the NL based input comprises:

processing, using a large language model (LLM), LLM input to generate LLM output, the LLM input including at least the NL based input;

determining, based on the LLM output, textual content and multimedia content to be included in the response that is responsive to the NL based input;

initiating obtaining of the multimedia content to be included in the response that is response to the NL based input;

while obtaining the multimedia content to be included in the response that is responsive to the NL based input:

determining, based on one or more signals, whether to continue obtaining the multimedia content to be included in the response that is responsive to the NL based input; and

in response to determining to refrain from continuing to obtain the multimedia content to be included in the response that is responsive to the NL based input:

disengaging obtaining of the multimedia content; and

determining canned textual content or other textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input; and

causing the response, including the canned textual content or the other textual content, to be rendered at the client device of the user.

2. The method of claim 1 , wherein the one or more signals comprises one or more of: an NL based input context given the NL based input with respect to the multimedia content requested by the NL based input, a multimedia content context given the multimedia content requested by the NL based input with respect to the NL based input, or a response context given the multimedia content requested by the NL based input with respect to the textual content.

3. The method of claim 2 , wherein the one or more signals comprise the NL based input context given the NL based input with respect to the multimedia content requested by the NL based input, and wherein the NL based input context indicates whether the multimedia content requested by the NL based input should be rendered given the NL based input.

4. The method of claim 2 , wherein the one or more signals comprise the multimedia content context given the multimedia content requested by the NL based input with respect to the NL based input, and wherein the multimedia content context indicates whether the multimedia content requested by the NL based input should be rendered given a generative multimedia content prompt to generate the multimedia content or given a non-generative multimedia content query to obtain the multimedia content.

5. The method of claim 2 , wherein the one or more signals comprise the response context given the multimedia content requested by the NL based input with respect to the textual content, and wherein the response context indicates whether the multimedia content requested by the NL based input should be rendered given the textual content that is determined to be included in the response.

6. The method of claim 1 , wherein the multimedia content is non-generative multimedia content, and wherein disengaging obtaining of the multimedia content comprises:

canceling a non-generative multimedia content query to obtain the multimedia content.

7. The method of claim 1 , wherein the multimedia content is generative multimedia content, and wherein disengaging obtaining of the multimedia content comprises:

canceling processing, by a generative multimedia content model, of a generative multimedia content prompt to generate the multimedia content.

8. The method of claim 1 , wherein determining the canned textual content or the other textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input comprises:

determining the canned textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input, wherein the canned textual content indicates that the multimedia content cannot be rendered.

9. The method of claim 8 , wherein the canned textual content further indicates a certain reason for why the multimedia content cannot be rendered.

10. The method of claim 1 , wherein determining the canned textual content or the other textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input comprises:

determining the other textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input, wherein the other textual content is determined based on the LLM output.

11. A method implemented by one or more processors, the method comprising:

receiving natural language (NL) based input associated with a client device of a user, the NL based input requesting multimedia content;

determining, based on one or more terms included in the NL based input, to refrain from including the multimedia content a response that is responsive to the multimedia content;

generating the response that is responsive to the NL based input, wherein generating the response that is responsive to the NL based input comprises:

processing, using a large language model (LLM), LLM input to generate LLM output, the LLM input including at least the NL based input;

determining, based on the LLM output, textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input; and

causing the response, including the textual content, to be rendered at the client device of the user.

12. A system comprising:

one or more processors; and

memory storing instructions that, when executed, cause the one or more processors to be operable to:

receive natural language (NL) based input associated with a client device of a user, the NL based input requesting multimedia content;

generate a response that is responsive to the NL based input, wherein, in generating the response that is responsive to the NL based input, the one or more processors are operable to:

process, using a large language model (LLM), LLM input to generate LLM output, the LLM input including at least the NL based input;

determine, based on the LLM output, textual content and multimedia content to be included in the response that is responsive to the NL based input;

initiate obtaining of the multimedia content to be included in the response that is response to the NL based input;

while obtaining the multimedia content to be included in the response that is responsive to the NL based input:

determine, based on one or more signals, whether to continue obtaining the multimedia content to be included in the response that is responsive to the NL based input; and

in response to determining to refrain from continuing to obtain the multimedia content to be included in the response that is responsive to the NL based input:

disengage obtaining of the multimedia content; and

determine canned textual content or other textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input; and

cause the response, including the canned textual content or the other textual content, to be rendered at the client device of the user.

13. The system of claim 12 , wherein the one or more signals comprises one or more of: an NL based input context given the NL based input with respect to the multimedia content requested by the NL based input, a multimedia content context given the multimedia content requested by the NL based input with respect to the NL based input, or a response context given the multimedia content requested by the NL based input with respect to the textual content.

14. The system of claim 13 , wherein the one or more signals comprise the NL based input context given the NL based input with respect to the multimedia content requested by the NL based input, and wherein the NL based input context indicates whether the multimedia content requested by the NL based input should be rendered given the NL based input.

15. The system of claim 13 , wherein the one or more signals comprise the multimedia content context given the multimedia content requested by the NL based input with respect to the NL based input, and wherein the multimedia content context indicates whether the multimedia content requested by the NL based input should be rendered given a generative multimedia content prompt to generate the multimedia content or given a non-generative multimedia content query to obtain the multimedia content.

16. The system of claim 13 , wherein the one or more signals comprise the response context given the multimedia content requested by the NL based input with respect to the textual content, and wherein the response context indicates whether the multimedia content requested by the NL based input should be rendered given the textual content that is determined to be included in the response.

17. The system of claim 12 , wherein the multimedia content is non-generative multimedia content, and wherein, in disengaging obtaining of the multimedia content, the one or more processors are operable to:

cancel a non-generative multimedia content query to obtain the multimedia content.

18. The system of claim 12 , wherein the multimedia content is generative multimedia content, and wherein, in disengaging obtaining of the multimedia content, the one or more processors are operable to:

cancel processing, by a generative multimedia content model, of a generative multimedia content prompt to generate the multimedia content.

19. The system of claim 12 , wherein, in determining the canned textual content or the other textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input, the one or more processors are operable to:

determine the canned textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input, wherein the canned textual content indicates that the multimedia content cannot be rendered and/or a certain reason for why the multimedia content cannot be rendered.

20. The system of claim 12 , wherein, in determining the canned textual content or the other textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input, the one or more processors are operable to:

determine the other textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input, wherein the other textual content is determined based on the LLM output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2024
From: JAIN, SANIL; YU, WEI; WEISZ, ÁGOSTON; GOODMAN, MICHAEL ANDREW; AVRAM, DIANA; GHAFOURI, AMIN; GHIASI, GOLNAZ; PETROVSKI, IGOR; GUPTA, KHYATTI; AKERLUND, OSCAR; SLUZHAEV, EVGENY; SHIVANNA, RAKESH; LUONG, THANG; SINGH, KOMAL; LU, YIFENG; PESWANI, VIKAS
To: GOOGLE LLC
Reel/Frame 066788/0763 →
Continuity (1)
Continuation 18520218 · Nov 27, 2023
References Cited (7)
US 11358063B2 · Ashoori · 2022 [cited by applicant]
US 11769017B1 · Gray · 2023 [cited by applicant]
US 11875240B1 · Bosnjakovic · 2024 [cited by applicant]
US 11947923B1 · Jain · 2024 [cited by applicant]
Du, H. et al., “Spear or Shield: Leveraging Generative AI to Tackle Security Threats of Intelligent Network Services”; arXiv.org, Cornell University; arXiv:2306.02384; 9 pages; dated Jun. 4, 2023. [cited by applicant]
Chen, X. et al., “Next Steps for Human-Centered Generative AI: A Technical Perspective”; arXiv, Cornell University; arXiv:2306.15774; 34 pages; dated Jun. 27, 2023. [cited by applicant]
European Patent Office; International Search Report and Written Opinion issued in Application No. PCT/US2024/050037; 10 pages; dated Jan. 28, 2025. [cited by applicant]