IP Library Granted Patent US 12688524
Granted Patent B2
US 12688524 · App. 18/542,097 · Granted Jul 21, 2026

Intelligent recommendations based on multimodal inputs

Inventors: Landon Chambers (Austin, TX); Travers Humble (Austin, TX); Anush Kumar (Bellevue, WA); Nehemias Luna (Seattle, WA); Raisa Meneses Guzman (Seattle, WA); Marc Silbey (Mercer Island, WA); Justin Simonelli (Seattle, WA); Phalguna Yadlapati (Toronto, CA)
Assignee: Expedia, Inc.
G06Q30/0631G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12688524
App. No.
18/542,097
Granted
Jul 21, 2026
Kind
B2
Abstract

A computing system includes a memory device and a processor structured to: receive, from a user device, a media input; extract, using a machine-learning model, at least one feature of the media input; identify, using the machine-learning model, at least one intent; determine, using the machine-learning model, at least one action policy; generate a user interface comprising the media input, the at least one extracted feature of the media input, and the at least one action policy; provide the user interface to the user device; receive, via an input to the user interface, an indication of a selection of the at least one action policy displayed on the user interface; generate a second user interface comprising a plurality of options associated with the selected at least one action policy; and provide the second user interface to the user device.

Claims (83)

1 . A computing system, comprising:

a network interface configured to communicate with a user device;

a processing circuit comprising at least one processor and at least one memory, the at least one memory structured to store instructions that are executable to cause the at least one processor to:

receive, via the network interface and from the user device, a media input;

extract, using a machine-learning model, at least one feature of the media input;

receive a user preference associated with a user of the user device;

receive contextual information associated with the user device;

identify, using the machine-learning model and based at least partly on the user preference and the contextual information, at least one intent associated with the at least one extracted feature of the media input;

determine, using the machine-learning model, at least one action policy based on the at least one identified intent;

generate a user interface comprising the media input, the at least one extracted feature of the media input, and the at least one action policy;

provide, via the network interface, the user interface to the user device;

receive, via the network interface and via an input to the user interface, an indication of a selection of the at least one action policy displayed on the user interface;

generate a second user interface comprising a plurality of options associated with the selected at least one action policy; and

provide, via the network interface, the second user interface to the user device.

2 . The computing system of claim 1 , wherein the instructions are executable to further cause the at least one processor to:

train the machine-learning model using predetermined training data;

wherein the predetermined training data includes one or more of image data, caption data, travel metadata, intent data, or action data.

3 . The computing system of claim 1 , wherein the at least one action policy comprises at least one of a travel itinerary, a recommendation to book travel, or a recommended location to visit.

4 . The computing system of claim 1 , wherein the instructions are executable to further cause the at least one processor to:

pull, from a database, historical information of a previously received media input from the user device;

generate a prompt regarding the pulled historical information;

provide, via the network interface, the prompt to the user device;

receive, via the network interface and via an input to the user interface, an indication of a selection of the prompt, the selection of the prompt confirming the historical information displayed on the user interface; and

determine, using the machine-learning model, the at least one action policy based on the at least one identified intent and the confirmed historical information.

5 . The computing system of claim 1 , wherein the at least one feature of the media input comprises a metadata tag, and wherein the instructions are executable to further cause the at least one processor to:

pull, from a database, a plurality of predetermined tags;

compare, using the machine-learning model, the metadata tag with the plurality of predetermined tags;

determine, using the machine-learning model and based on the comparison, information associated with the metadata tag; and

identify, using the machine-learning model, the at least one intent based on the information associated with the metadata tag.

6 . The computing system of claim 1 , wherein receiving the media input comprises retrieving the media input from a stored collection of media inputs on the user device.

7 . The computing system of claim 1 , wherein receiving the media input comprises receiving the media input via a camera device of the user device.

8 . The computing system of claim 1 , wherein receiving the media input comprises receiving the media input via a third-party application programming interface (API).

9 . The computing system of claim 1 , wherein the machine-learning model comprises at least one pre-processing model and at least one prediction model.

10 . A computer-implemented method, comprising:

receiving, by a computing system and from a user device communicably coupled to the computing system, a media input;

extracting, by the computing system using a machine-learning model stored in the computing system, at least one feature of the media input;

receiving, by the computing system, a user preference associated with a user of the user device;

receiving, by the computing system, contextual information associated with the user device;

identifying, by the computing system using the machine-learning model and based at least partly on the user preference and the contextual information, at least one intent associated with the at least one extracted feature of the media input;

determining, by the computing system using the machine-learning model, at least one action policy based on the at least one identified intent;

generating, by the computing system, a user interface comprising the media input, the at least one extracted feature of the media input, and the at least one action policy;

providing, by the computing system, the user interface to the user device;

receiving, by the computing system and via an input to the user interface of the user device, an indication of a selection of the at least one action policy displayed on the user interface;

generating, by the computing system, a second user interface comprising a plurality of options associated with the selected at least one action policy; and

providing, by the computing system, the second user interface to the user device.

11 . The method of claim 10 , further comprising:

training, by the computing system, the machine-learning model using predetermined training data;

wherein the predetermined training data includes one or more of image data, caption data, travel metadata, intent data, or action data.

12 . The method of claim 10 , wherein the at least one action policy comprises at least one of a travel itinerary, a recommendation to book travel, or a recommended location to visit.

13 . The method of claim 10 , further comprising:

pulling, by the computing system and from a database, historical information of a previously received media input from the user device;

generating, by the computing system, a prompt regarding the pulled historical information;

providing, by the computing system, the prompt to the user device;

receiving, by the computing system and via an input to the user interface, an indication of a selection of the prompt, the indication of the selection of the prompt confirming the historical information displayed on the user interface; and

determining, by the computing system using the machine-learning model, the at least one action policy based on the at least one identified intent and the confirmed historical information.

14 . The method of claim 10 , wherein the at least one feature of the media input comprises a metadata tag, and wherein the method further comprises:

pulling, by the computing system and from a database, a plurality of predetermined tags;

comparing, by the computing system using the machine-learning model, the metadata tag with the plurality of predetermined tags;

determining, by the computing system using the machine-learning model and based on the comparison, information associated with the metadata tag; and

identifying, by the computing system and using the machine-learning model, the at least one intent based on the information associated with the metadata tag.

15 . The method of claim 10 , wherein receiving the media input comprises retrieving the media input from a stored collection of media inputs on the user device.

16 . The method of claim 10 , wherein receiving the media input comprises receiving the media input via a camera device of the user device.

17 . The method of claim 10 , wherein receiving the media input comprises receiving the media input via a third-party application programming interface (API).

18 . A non-transitory computer-readable media having computer-executable instructions embodied therein that, when executed by at least one processor of a provider computing system, cause the provider computing system to perform operations comprising:

receiving, from a user device, a media input;

extracting, using a machine-learning model, at least one feature of the media input;

receiving a user preference associated with a user of the user device;

receiving contextual information associated with the user device;

identifying, using the machine-learning model and based at least partly on the user preference and the contextual information, at least one intent associated with the at least one extracted feature of the media input;

determining, using the machine-learning model, at least one action policy based on the at least one identified intent;

generating a user interface comprising the media input, the at least one extracted feature of the media input, and the at least one action policy;

providing the user interface to the user device;

receiving, via an input to the user interface of the user device, an indication of a selection of the at least one action policy displayed on the user interface;

generating a second user interface comprising a plurality of options associated with the selected at least one action policy; and

providing the second user interface to the user device.

19 . The non-transitory computer-readable media of claim 18 , wherein the operations further comprise:

training the machine-learning model using predetermined training data;

wherein the predetermined training data includes one or more of image data, caption data, travel metadata, intent data, or action data.

20 . The non-transitory computer-readable media of claim 18 , wherein the at least one feature of the media input comprises a metadata tag, and wherein the operations further comprise:

pulling, from a database, a plurality of predetermined tags;

comparing, using the machine-learning model, the metadata tag with the plurality of predetermined tags;

determining, using the machine-learning model, based on the comparison, information associated with the metadata tag; and

identifying, using the machine-learning model, the at least one intent based on the information associated with the metadata tag.