IP Library Granted Patent US 12705994
Granted Patent B2
US 12705994 · App. 18/944,276 · Granted Aug 11, 2026

System and method for audio guide

Inventors: Alvaro Javier Lazaro Aguilar (Tampa, FL); Weiyi He (New York, NY); Ambar Aballo Ruiz (Miami, FL); Paige Lynette Reiter (Clarkston, MI); Thomas Owen Williams (Orlando, FL); Howard Bruce Mall (Winter Springs, FL)
Assignee: Universal City Studios LLC
G09B5/04G06F40/40G06V10/74G06V20/50G10L13/047H04N23/66
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705994
App. No.
18/944,276
Granted
Aug 11, 2026
Kind
B2
Abstract

A method of providing audio descriptions of landmarks includes causing a user device to capture an image via an imaging sensor of the user device, comparing the captured image to a reference image to identify a landmark that appears in the captured image, providing a prompt requesting a description associated with the identified landmark to a large language model (LLM), receiving an audio file of the description associated with the identified landmark, and providing the audio file of the description associated with the identified landmark to the one or more user devices.

Claims (66)

1 . An audio guide system, comprising:

a portable device associated with a guest, wherein the portable device comprises an imaging sensor configured to capture an image;

a beacon configured to:

detect a presence of the portable device in an area;

query the portable device based on detecting the presence of the portable device in the area; and

cause the imaging sensor of the portable device to capture the image in response to the query; and

a computing device comprising:

processing circuitry; and

memory, accessible by the processing circuitry and storing instructions that, when executed by the processing circuitry, cause the processing circuitry to perform operations comprising:

receiving the captured image;

comparing the captured image to a reference image to identify a landmark that appears in the captured image;

generating a prompt requesting a description associated with the identified landmark;

providing the prompt to a large language model (LLM);

receiving, from the LLM, the description associated with the identified landmark;

providing the description associated with the identified landmark to a text-to-speech model;

receiving an audio file of the description associated with the identified landmark from the text-to-speech model; and

providing the audio file of the description associated with the identified landmark to the portable device, wherein the portable device is configured to play the audio file in response to receipt of the audio file.

2 . The audio guide system of claim 1 , wherein the area comprises a portion of an amusement park.

3 . The audio guide system of claim 1 , wherein the portable device comprises:

a wearable device comprising the imaging sensor, wherein the wearable device is configured to be affixed to clothing of the guest; and

a handheld device configured to play the audio file for the guest.

4 . The audio guide system of claim 3 , wherein the handheld device comprises a speaker, wherein the handheld device is configured to play the audio file for the guest via the speaker.

5 . The audio guide system of claim 3 , wherein the handheld device comprises a headphone port configured to couple the handheld device to one or more headphones, wherein the handheld device is configured to play the audio file for the guest via the one or more headphones.

6 . The audio guide system of claim 1 , wherein the portable device comprises a mobile device.

7 . The audio guide system of claim 1 , wherein the reference image is retrieved from a landmark images database.

8 . A method of providing audio descriptions associated with landmarks, the method comprising:

detecting, via a beacon, a presence of a user device in an area;

querying, via the beacon, the user device based on detecting the presence of the user device in the area;

causing, via the beacon, the user device to capture an image via an imaging sensor of the user device in response to the query;

comparing the captured image to a reference image to identify a landmark that appears in the captured image;

providing, to a large language model (LLM), a prompt requesting a description associated with the identified landmark;

receiving an audio file of the description associated with the identified landmark; and

providing the audio file of the description associated with the identified landmark to the user device.

9 . The method of claim 8 , wherein the audio file is generated by the LLM.

10 . The method of claim 8 , comprising:

receiving, from the LLM, the description associated with the identified landmark; and

providing the description associated with the identified landmark to a text-to-speech model, wherein the audio file of the description associated with the identified landmark is generated by the text-to-speech model.

11 . The method of claim 8 , comprising:

receiving, from the user device, an input requesting additional description associated with the identified landmark;

providing, to the LLM, an additional prompt requesting the additional description associated with the identified landmark;

receiving an additional audio file of the additional description associated with the identified landmark; and

providing the additional audio file of the additional description associated with the identified landmark to the user device.

12 . The method of claim 8 , comprising providing one or more pieces of contextual data to the LLM, wherein the contextual data comprises an interest profile associated with a guest.

13 . The method of claim 12 , wherein the contextual data is comprises one or more types of landmarks in which the guest has demonstrated interest or disinterest.

14 . The method of claim 12 , wherein the contextual data is comprises a level of detail preferred by the guest.

15 . The method of claim 12 , comprising training the LLM based on contextual data.

16 . The method of claim 8 , comprising:

causing an additional user device to capture an additional image;

identifying an additional landmark that appears in the captured additional image;

providing, to the LLM, an additional prompt requesting an additional description associated with the identified additional landmark;

receiving an additional audio file of the additional description associated with the identified additional landmark; and

providing the additional audio file of the additional description associated with the identified additional landmark to the additional user device.

17 . A non-transitory computer readable medium storing instructions that, when executed by processing circuitry, cause the processing circuitry to perform operations comprising:

detecting, via a beacon, a presence of a portable device in an area;

querying, via the beacon, the portable device based on detecting the presence of the portable device in the area;

causing, via the beacon, the imaging sensor of the portable device to capture the image in response to the query;

receiving the captured image from the portable device;

comparing the captured image to a reference image to identify a landmark that appears in the captured image;

providing, to a large language model (LLM), a prompt requesting a description associated with the identified landmark;

receiving, from the LLM, the description associated with the identified landmark;

providing the description associated with the identified landmark to a text-to-speech model;

receiving an audio file of the description associated with the identified landmark from the text-to-speech model; and

providing the audio file of the description associated with the identified landmark to a user device, wherein the user device is configured to play the audio file in response to receipt of the audio file.

18 . The non-transitory computer readable medium of claim 17 , wherein the LLM is trained based on a database of amusement park documents.

19 . The non-transitory computer readable medium of claim 17 , wherein comparing the captured image to the reference image to identify the landmark that appears in the captured image is performed via a feature matching model.

20 . The non-transitory computer readable medium of claim 17 , wherein-the portable device comprises the user device.