IP Library Granted Patent US 12694231
Granted Patent B2
US 12694231 · App. 18/413,495 · Granted Jul 28, 2026

Generating multi-modal response(s) through utilization of large language model(s)

Inventors: Oscar Akerlund (Zurich, CH); Evgeny Sluzhaev (Zurich, CH); Golnaz Ghiasi (Mountain View, CA); Thang Luong (Santa Clara, CA); Yifeng Lu (Mountain View, CA); Igor Petrovski (Zurich, CH); Ágoston Weisz (Zurich, CH); Wei Yu (Mountain View, CA); Rakesh Shivanna (Sunnyvale, CA); Michael Andrew Goodman (Oakland, CA); Apoorv Kulshreshtha (Mountain View, CA); Yu Du (Sunnyvale, CA); Amin Ghafouri (San Francisco, CA); Sanil Jain (Sunnyvale, CA); Dustin Tran (San Francisco, CA); Vikas Peswani (Mountain View, CA); YaGuang Li (Sunnyvale, CA)
Assignee: GOOGLE LLC
G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694231
App. No.
18/413,495
Granted
Jul 28, 2026
Kind
B2
Abstract

Implementations relate to generating multi-modal response(s) through utilization of large language model(s) (LLM(s)). Processor(s) of a system can: receive natural language (NL) based input, generate a multi-modal response that is responsive to the NL based output, and cause the multi-modal response to be rendered. In some implementations, and in generating the multi-modal response, the processor(s) can process, using a LLM, LLM input (e.g., that includes at least the NL based input) to generate LLM output, and determine, based on the LLM output, textual content for inclusion in the multi-modal response and multimedia content for inclusion in the multi-modal response. In some implementations, the multimedia content can be obtained based on a multimedia content tag that is included in the LLM output and that is indicative of the multimedia content. In various implementations, the multimedia content can be interleaved between segments of the textual content.

Claims (45)

1 . A method implemented by one or more processors, the method comprising:

obtaining a plurality of training instances to be utilized in fine-tuning a large language model (LLM), wherein each training instance, of the plurality of training instances, includes:

a corresponding natural language (NL) based input, and

a corresponding multi-modal response that is responsive to the corresponding NL based input, the corresponding multi-modal response including corresponding textual content and a corresponding multimedia content tag that is indicative of corresponding multimedia content that is to be included in the corresponding multi-modal response;

fine-tuning, based on the plurality of training instances, the LLM; and

causing the LLM to be deployed for utilization in generating subsequent multi-modal responses that are responsive to subsequent NL based inputs that are associated with client devices of users.

2 . The method of claim 1 , wherein the corresponding NL based input and the corresponding textual content of the corresponding multi-modal response, for one or more of the plurality of training instances, are obtained from conversation logs between users and the LLM.

3 . The method of claim 2 , wherein one or more of the plurality of training instances are curated by a developer that is associated with the LLM.

4 . The method of claim 3 , further comprising:

receiving input from the developer that indicates where the corresponding multimedia content tag, that is indicative of where the corresponding multimedia content, is to be included in the corresponding multi-modal response.

5 . The method of claim 2 , wherein one or more of the plurality of training instances are automatically generated.

6 . The method of claim 5 , further comprising:

automatically inserting the corresponding multimedia content tag, that is indicative of where the corresponding multimedia content is to be included, in the corresponding multi-modal response.

7 . The method of claim 1 , wherein fine-tuning the LLM based on a given training instance, from among the plurality of training instances, comprises:

processing, using the LLM, the corresponding NL based input, of the given training instance, and the corresponding multi-modal response of the given training instance.

8 . The method of claim 7 , wherein fine-tuning the LLM based on the given training instance causes the LLM to determine when to include the corresponding multimedia content tag that is indicative of the corresponding multimedia content.

9 . The method of claim 7 , wherein fine-tuning the LLM based on the given training instance causes the LLM to determine where to include the corresponding multimedia content tag that is indicative of the corresponding multimedia content and with respect to the corresponding textual content of the corresponding multi-modal response.

10 . The method of claim 1 , wherein causing the LLM to be deployed for utilization in generating the subsequent multi-modal responses is in response to determining one or more conditions are satisfied.

11 . The method of claim 10 , wherein the one or more conditions comprise one or more of: the LLM being fine-tuned based based on a threshold quantity of training instances, the LLM being fine-tuned for a threshold duration of time, or performance of the LLM satisfying a threshold level of performance.

12 . A system comprising:

at least one processor; and

memory storing instructions that, when executed, cause the at least one processor to be operable to:

obtain a plurality of training instances to be utilized in fine-tuning a large language model (LLM), wherein each training instance, of the plurality of training instances, includes:

a corresponding natural language (NL) based input, and

a corresponding multi-modal response that is responsive to the corresponding NL based input, the corresponding multi-modal response including corresponding textual content and a corresponding multimedia content tag that is indicative of corresponding multimedia content that is to be included in the corresponding multi-modal response;

fine-tuning, based on the plurality of training instances, the LLM; and

cause the LLM to be deployed for utilization in generating subsequent multi-modal responses that are responsive to subsequent NL based inputs that are associated with client devices of users.

13 . The system of claim 12 , wherein the corresponding NL based input and the corresponding textual content of the corresponding multi-modal response, for one or more of the plurality of training instances, are obtained from conversation logs between users and the LLM.

14 . The system of claim 13 , wherein one or more of the plurality of training instances are curated by a developer that is associated with the LLM.

15 . The system of claim 13 , wherein the at least one processor is further operable to:

receive input from the developer that indicates where the corresponding multimedia content tag, that is indicative of where the corresponding multimedia content, is to be included in the corresponding multi-modal response.

16 . The system of claim 13 , wherein one or more of the plurality of training instances are automatically generated.

17 . The system of claim 16 , wherein the at least one processor is further operable to:

automatically insert the corresponding multimedia content tag, that is indicative of where the corresponding multimedia content is to be included, in the corresponding multi-modal response.

18 . The system of claim 12 , wherein fine-tuning the LLM based on a given training instance, from among the plurality of training instances, comprises:

processing, using the LLM, the corresponding NL based input, of the given training instance, and the corresponding multi-modal response of the given training instance.

19 . The system of claim 18 ,

wherein fine-tuning the LLM based on the given training instance causes the LLM to determine when to include the corresponding multimedia content tag that is indicative of the corresponding multimedia content, and

wherein fine-tuning the LL M based on the given training instance causes the LLM to determine where to include the corresponding multimedia content tag that is indicative of the corresponding multimedia content and with respect to the corresponding textual content of the corresponding multi-modal response.

20 . A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor to perform operations, the operations comprising:

obtaining a plurality of training instances to be utilized in fine-tuning a large language model (LLM), wherein each training instance, of the plurality of training instances, includes:

a corresponding natural language (NL) based input, and

a corresponding multi-modal response that is responsive to the corresponding NL based input, the corresponding multi-modal response including corresponding textual content and a corresponding multimedia content tag that is indicative of corresponding multimedia content that is to be included in the corresponding multi-modal response;

fine-tuning, based on the plurality of training instances, the LLM; and

causing the LLM to be deployed for utilization in generating subsequent multi-modal responses that are responsive to subsequent NL based inputs that are associated with client devices of users.