Multimodal embeddings
Implementations relate to generating and using multimodal embeddings. In various implementations, first modality data may be obtained and encoded into first modality embedding(s) using a trained first modality encoder that is stored in memory of edge-based client device(s). Second modality data may be obtained and encoded into second modality embedding(s) using a trained second modality encoder that is also stored in the memory of the edge-based client device(s). The first and second modality embeddings may be processed using an edge-based multimodal LLM that is also stored locally in memory of the edge-based client device(s) to generate a multimodal contextual embedding, which may be provided to a remote server that hosts a central LLM, e.g., in conjunction with a natural language input provided by the user. Information generated using the central LLM, responsive to the natural language input, may be received from the remote server.
1 . A method implemented using one or more processors of one or more edge-based client devices and comprising:
contemporaneously with receipt of a natural language input from a user, obtaining a first digital image captured in an environment using a front-facing camera of a given edge-based client device;
encoding the first digital image into one or more first modality embeddings using a trained first modality encoder that is stored in memory of the given edge-based client device;
contemporaneously with receipt of the natural language input, obtaining a second digital image captured in the environment, wherein the second digital image comprises a screenshot captured by the given edge-based client device or is captured using a rear-facing camera of the given edge-based client device;
encoding the second digital image into one or more second modality embeddings using a trained second modality encoder that is stored in memory of one or more of the edge-based client devices;
processing one or more of the first modality embeddings and one or more of the second modality embeddings using an edge-based multimodal large language model (LLM) that is stored locally in memory of the given edge-based client device to generate a multimodal contextual embedding;
providing, to a remote server that hosts a central LLM, data indicative of the multimodal contextual embedding and the natural language input provided by the user; and
receiving, from the remote server, information generated using the central LLM that is responsive to the natural language input provided by the user.
2 . The method of claim 1 , wherein the second digital image comprises the screenshot captured by the given edge-based client device.
3 . The method of claim 1 , wherein the first digital image captures a facial expression of the user, and one or more of the first modality embeddings numerically represents the captured facial expression.
4 . An edge-based system comprising one or more edge processors and memory storing instructions that, in response to execution by the one or more edge processors, cause the one or more edge processors to:
contemporaneously with receipt of a natural language input from a user, obtain a first digital image captured in an environment using a front-facing camera of a given edge-based client device;
encode the first digital image into one or more first modality embeddings using a trained first modality encoder that is stored in memory of the given edge-based client device;
contemporaneously with receipt of the natural language input, obtain a second digital image captured in the environment, wherein the second digital image comprises a screenshot captured by the given edge-based client device or is captured using a rear-facing camera of the given edge-based client device;
encode the second digital image into one or more second modality embeddings using a trained second modality encoder that is stored in memory of the given edge-based client device;
process one or more of the first modality embeddings and one or more of the second modality embeddings using an edge-based multimodal large language model (LLM) that is stored locally in memory of one or more of the edge-based client devices to generate a multimodal contextual embedding;
provide, to a remote server that hosts a central LLM, data indicative of the multimodal contextual embedding and the natural language input provided by the user; and
receive, from the remote server, information generated using the central LLM that is responsive to the natural language input provided by the user.
5 . The system of claim 4 , wherein the first digital image captures a facial expression of the user, and one or more of the first modality embeddings numerically represents the captured facial expression.
6 . At least one non-transitory computer-readable medium comprising instructions that, in response to execution by one or more edge processors, cause the one or more edge processors to:
contemporaneously with receipt of a natural language input from a user, obtain a first digital image captured in an environment using a front-facing camera of a given edge-based client device;
encode the first digital image into one or more first modality embeddings using a trained first modality encoder that is stored in memory of the given edge-based client device;
contemporaneously with receipt of the natural language input, obtain a second digital image captured in the environment, wherein the second digital image comprises a screenshot captured by the given edge-based client device or is captured using a rear-facing camera of the given edge-based client device;
encode the second digital image into one or more second modality embeddings using a trained second modality encoder that is stored in memory of the given edge-based client device;
process one or more of the first modality embeddings and one or more of the second modality embeddings using an edge-based multimodal large language model (LLM) that is stored locally in memory of one or more of the edge-based client devices to generate a multimodal contextual embedding;
provide, to a remote server that hosts a central LLM, data indicative of the multimodal contextual embedding and the natural language input provided by the user; and
receive, from the remote server, information generated using the central LLM that is responsive to the natural language input provided by the user.