IP Library › Granted Patent US 11,947,923
Granted Patent B1
US 11,947,923 · App. 18/520,218 · Granted Apr 2, 2024

Multimedia content management for large language model(s) and/or other generative model(s)

Inventors: Sanil Jain (Sunnyvale, CA); Wei Yu (Mountain View, CA); Ágoston Weisz (Zurich, CH); Michael Andrew Goodman (Oakland, CA); Diana Avram (Zurich, CH); Amin Ghafouri (San Francisco, CA); Golnaz Ghiasi (Mountain View, CA); Igor Petrovski (Zurich, CH); Khyatti Gupta (Zurich, CH); Oscar Akerlund (Zurich, CH); Evgeny Sluzhaev (Zurich, CH); Rakesh Shivanna (Sunnyvale, CA); Thang Luong (Santa Clara, CA); Komal Singh (Kitchener, CA); Yifeng Lu (Mountain View, CA); Vikas Peswani (Mountain View, CA)
Assignee: GOOGLE LLC
G06F40/40G06V10/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,947,923
App. No.
18/520,218
Filed
Nov 27, 2023
Granted
Apr 2, 2024
Kind
B1
Art Unit
2654
USPC
704/9
Abstract

Implementations relate to managing multimedia content that is obtained by large language model(s) (LLM(s)) and/or generated by other generative model(s). Processor(s) of a system can: receive natural language (NL) based input that requests multimedia content, generate a response that is responsive to the NL based input, and cause the response to be rendered. In some implementations, and in generating the response, the processor(s) can process, using a LLM, LLM input to generate LLM output, and determine, based on the LLM output, at least multimedia content to be included in the response. Further, the processor(s) can evaluate the multimedia content to determine whether it should be included in the response. In response to determining that the multimedia content should not be included in the response, the processor(s) can cause the response, including alternative multimedia content or other textual content, to be rendered.

Claims (78)

1. A method implemented by one or more processors, the method comprising:

receiving natural language (NL) based input associated with a client device of a user, the NL based input requesting multimedia content;

generating a response that is responsive to the NL based input, wherein generating the response that is responsive to the NL based input comprises:

processing, using a large language model (LLM), LLM input to generate LLM output, the LLM input including at least the NL based input;

determining, based on the LLM output, multimedia content to be included in the response that is responsive to the NL based input;

obtaining the multimedia content to be included in the response that is responsive to the NL based input;

determining, based on processing the NL based input and the multimedia content that is to be included in the response that is responsive to the NL based input, whether to cause the response, including the multimedia content, to be rendered at the client device; and

in response to determining to refrain from causing the response to be rendered at the client device:

determining alternative multimedia content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input; and

causing the response, including the alternative multimedia content, to be rendered at the client device of the user.

2. The method of claim 1 , wherein processing the NL based input and the multimedia content that is to be included in the response that is responsive to the NL based input comprises:

processing, using a visual language model (VLM), the NL based input and the multimedia content that is to be included in the response that is responsive to the NL based input to generate VLM output; and

determining, based on the VLM output, whether to cause the response, including the multimedia content, to be rendered at the client device.

3. The method of claim 2 , further comprising:

processing, using the VLM, and along with the NL based input and the multimedia content that, a prompt to generate the VLM output, wherein the prompt includes a request for the VLM to determine whether the multimedia content should be rendered at the client device of the user and given the NL based input.

4. The method of claim 3 , wherein the VLM output includes an indication of whether the multimedia content should be rendered at the client device of the user and given the NL based input.

5. The method of claim 4 , wherein the VLM output further includes a reason for whether the multimedia content should be rendered at the client device of the user and given the NL based input.

6. The method of claim 4 , wherein determining to refrain from causing the response to be rendered at the client device is based on the VLM output including an indication that the multimedia content should not be rendered at the client device of the user.

7. The method of claim 6 , wherein determining the alternative multimedia content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input comprises:

processing, using the LLM, additional LLM input to generate additional LLM output, the additional LLM input including at least the NL based input and at least a portion of the VLM output; and

determining, based on the additional LLM output, the alternative multimedia content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input.

8. The method of claim 1 , wherein processing the NL based input and the multimedia content that is to be included in the response that is responsive to the NL based input comprises:

processing, using a captioning model, the multimedia content to generate one or more corresponding captions for the multimedia content;

processing, using the LLM, the NL based input and one or more corresponding captions for the multimedia content to generate evaluation LLM output; and

determining, based on the evaluation LLM output, whether to cause the response, including the multimedia content, to be rendered at the client device.

9. The method of claim 8 , further comprising:

processing, using the LLM, and along with the NL based input and one or more corresponding captions for the multimedia content, a prompt to generate the evaluation LLM output, wherein the prompt includes a request for the LLM to determine whether the multimedia content should be rendered at the client device of the user and given the NL based input.

10. The method of claim 9 , wherein the evaluation LLM output includes an indication of whether the multimedia content should be rendered at the client device of the user and given the NL based input.

11. The method of claim 10 , wherein the evaluation LLM output further includes a reason for whether the multimedia content should be rendered at the client device of the user and given the NL based input.

12. The method of claim 10 , wherein determining to refrain from causing the response to be rendered at the client device is based on the evaluation LLM output including an indication that the multimedia content should not be rendered at the client device of the user.

13. The method of claim 12 , wherein determining the alternative multimedia content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input comprises:

processing, using the LLM, additional LLM input to generate additional LLM output, the additional LLM input including at least the NL based input and at least a portion of the evaluation LLM output; and

determining, based on the additional LLM output, the alternative multimedia content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input.

14. The method of claim 8 , wherein the LLM output and the evaluation LLM output are generated using a single call to the LLM.

15. The method of claim 1 , wherein processing the NL based input and the multimedia content that is to be included in the response that is responsive to the NL based input comprises:

obtaining one or more corresponding multimedia content tags that are stored in association with the multimedia content;

processing, using the LLM, the NL based input and one or more of the corresponding multimedia content tags to generate evaluation LLM output; and

determining, based on the evaluation LLM output, whether to cause the response, including the multimedia content, to be rendered at the client device.

16. The method of claim 15 , further comprising:

processing, using the LLM, and along with the NL based input and one or more of the corresponding multimedia content tags for the multimedia content, a prompt to generate the evaluation LLM output, wherein the prompt includes a request for the LLM to determine whether the multimedia content should be rendered at the client device of the user and given the NL based input.

17. The method of claim 16 , wherein the evaluation LLM output includes an indication of whether the multimedia content should be rendered at the client device of the user and given the NL based input.

18. The method of claim 17 , wherein the evaluation LLM output further includes a reason for whether the multimedia content should be rendered at the client device of the user and given the NL based input.

19. The method of claim 17 , wherein determining to refrain from causing the response to be rendered at the client device is based on the evaluation LLM output including an indication that the multimedia content should not be rendered at the client device of the user.

20. The method of claim 19 , wherein determining the alternative multimedia content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input comprises:

processing, using the LLM, additional LLM input to generate additional LLM output, the additional LLM input including at least the NL based input and at least a portion of the evaluation LLM output; and

determining, based on the additional LLM output, the alternative multimedia content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input.

21. A method implemented by one or more processors, the method comprising:

receiving natural language (NL) based input associated with a client device of a user, the NL based input requesting multimedia content;

generating a response that is responsive to the NL based input, wherein generating the response that is responsive to the NL based input comprises:

processing, using a large language model (LLM), LLM input to generate LLM output, the LLM input including at least the NL based input;

determining, based on the LLM output, multimedia content to be included in the response that is responsive to the NL based input;

obtaining the multimedia content to be included in the response that is responsive to the NL based input;

determining, based on processing the NL based input and the multimedia content that is to be included in the response that is responsive to the NL based input, whether to cause the response, including the multimedia content, to be rendered at the client device; and

in response to determining to refrain from causing the response to be rendered at the client device:

determining canned textual content and/or other textual content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input; and

causing the response, including the canned textual content, to be rendered at the client device of the user.

22. A system comprising:

one or more processors; and

memory storing instructions that, when executed, cause the one or more processors to be operable to:

receive natural language (NL) based input associated with a client device of a user, the NL based input requesting multimedia content;

generate a response that is responsive to the NL based input, wherein, in generating the response that is responsive to the NL based input, the one or more processors are operable to:

process, using a large language model (LLM), LLM input to generate LLM output, the LLM input including at least the NL based input;

determine, based on the LLM output, multimedia content to be included in the response that is responsive to the NL based input;

obtain the multimedia content to be included in the response that is responsive to the NL based input;

determine, based on processing the NL based input and the multimedia content that is to be included in the response that is responsive to the NL based input, whether to cause the response, including the multimedia content, to be rendered at the client device; and

in response to determining to refrain from causing the response to be rendered at the client device:

determine alternative multimedia content, canned textual content, and/or other textual content, to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input; and

cause the response to be rendered at the client device of the user.

23. The system of claim 22 , wherein, in processing the NL based input and the multimedia content that is to be included in the response that is responsive to the NL based input, the one or more processors are operable to:

process, using a visual language model (VLM), the NL based input and the multimedia content that is to be included in the response that is responsive to the NL based input to generate VLM output; and

determine, based on the VLM output, whether to cause the response, including the multimedia content, to be rendered at the client device.

24. The system of claim 23 , wherein the one or more processors are further operable to:

process, using the VLM, and along with the NL based input and the multimedia content that, a prompt to generate the VLM output, wherein the prompt includes a request for the VLM to determine whether the multimedia content should be rendered at the client device of the user and given the NL based input.

25. The system of claim 24 , wherein the VLM output includes an indication of whether the multimedia content should be rendered at the client device of the user and given the NL based input, and/or wherein the VLM output includes a reason for whether the multimedia content should be rendered at the client device of the user and given the NL based input.

26. The system of claim 25 , wherein determining to refrain from causing the response to be rendered at the client device is based on the VLM output including an indication that the multimedia content should not be rendered at the client device of the user.

27. The system of claim 26 , wherein, in determining the alternative multimedia content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input, the one or more processors are further operable to:

process, using the LLM, additional LLM input to generate additional LLM output, the additional LLM input including at least the NL based input and at least a portion of the VLM output; and

determine, based on the additional LLM output, the alternative multimedia content to be included in the response, and in lieu of the multimedia content, that is responsive to the NL based input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2023
From: JAIN, SANIL; YU, WEI; WEISZ, ÁGOSTON; GOODMAN, MICHAEL ANDREW; AVRAM, DIANA; GHAFOURI, AMIN; GHIASI, GOLNAZ; PETROVSKI, IGOR; GUPTA, KHYATTI; AKERLUND, OSCAR; SLUZHAEV, EVGENY; SHIVANNA, RAKESH; LUONG, THANG; SINGH, KOMAL; LU, YIFENG; PESWANI, VIKAS
To: GOOGLE LLC
Reel/Frame 065731/0465 →
Cited By (4)
US 12,277,400 US 12,368,931 US 12,699,727 US 12,724,785