IP Library Granted Patent US 12,666,223
Granted Patent B2
US 12,666,223 · App. 18/387,398 · Granted Jun 23, 2026

Location-aware assistant

Inventor: Dongeek Shin (San Jose, CA)
Assignee: GOOGLE LLC
H04W4/021G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,666,223
App. No.
18/387,398
Granted
Jun 23, 2026
Kind
B2
Abstract

Methods, systems, apparatus, including computer programs encoded on a computer storage medium, for spatialized audio feedback from automated assistants. In one aspect, the method includes actions of determining that a user input has been received at a client device, identifying, based on sensor data, one or more points of interest of an environment in which the client device is located and an orientation of a user of the client device relative to the one or more points of interest, identifying, based on processing the user input and the sensor data, a natural language response providing information relevant corresponds to a particular point of interest, determining, based on the orientation of the user of the client device relative to the particular point of interest, one or more spatial audio parameters to be used to provision the natural language response to the user, and causing the natural language response to be audibly rendered at the client device using one or more of the spatial audio parameters.

Claims (102)

1 . A computer-implemented method, the method comprising:

determining that a user input has been received at a client device;

identifying, based on sensor data available to the client device from one or more sensors:

one or more points of interest of an environment in which the client device is located, wherein identifying the one or more points of interest of the environment comprises cross referencing a geolocation of the client device with one or more landmarks contained in an electronic map of the environment, and

the points of interest are associated with dynamic information corresponding to the landmarks,

wherein the dynamic information comprises merchandise information; and

identifying an orientation of a user of the client device relative to the one or more points of interest,

wherein identifying the orientation of the user comprises determining the geolocation of the client device;

identifying, based on processing the user input and the sensor data, a natural language response to the user input, wherein the natural language response provides information relevant to a particular point of interest, of the points of interest, of the environment;

determining, based on the orientation of the user of the client device relative to the particular point of interest, one or more spatial audio parameters to be used to provision the natural language response to the user; and

causing the natural language response to be audibly rendered at the client device using one or more of the spatial audio parameters.

2 . The method of claim 1 , wherein the dynamic information is retrieved from a website, database, and/or application programming interface (API) created for the landmarks.

3 . The method of claim 1 , wherein the one or more spatial audio parameters comprise binauralized directional audio.

4 . The method of claim 1 , further comprising:

prior to identifying the one or more points of interest of the environment and orientation of the user relative to the points of interest:

determining the user input is directed towards the environment.

5 . The method of claim 1 , wherein the one or more spatial audio parameters further comprise at least one of volume, reverb, or echo parameters adjusted based on the orientation and a distance between the user and the particular point of interest.

6 . A computer-implemented method, the method comprising:

determining that a user input has been received at a client device;

identifying, based on sensor data available to the client device from one or more sensors:

one or more points of interest of an environment in which the client device is located, wherein identifying the one or more points of interest of the environment comprises cross referencing a geolocation of the client device with one or more landmarks contained in an electronic map of the environment, and

the points of interest are associated with dynamic information corresponding to the landmarks

wherein the dynamic information comprises accessibility information; and

identifying an orientation of a user of the client device relative to the one or more points of interest,

wherein identifying the orientation of the user comprises determining the geolocation of the client device;

identifying, based on processing the user input and the sensor data, a natural language response to the user input, wherein the natural language response provides information relevant to a particular point of interest, of the points of interest, of the environment;

determining, based on the orientation of the user of the client device relative to the particular point of interest, one or more spatial audio parameters to be used to provision the natural language response to the user; and

causing the natural language response to be audibly rendered at the client device using one or more of the spatial audio parameters.

7 . The method of claim 6 , wherein the dynamic information is retrieved from a website, database, and/or application programming interface (API) created for the landmarks.

8 . The method of claim 6 , wherein the one or more spatial audio parameters comprise binauralized directional audio.

9 . The method of claim 6 , further comprising:

prior to identifying the one or more points of interest of the environment and orientation of the user relative to the points of interest:

determining the user input is directed towards the environment.

10 . A computer-implemented method, the method comprising:

determining that a user input has been received at a client device;

identifying, based on sensor data available to the client device from one or more sensors:

one or more points of interest of an environment in which the client device is located, wherein identifying the one or more points of interest of the environment comprises cross referencing a geolocation of the client device with one or more landmarks contained in an electronic map of the environment, and

the points of interest are associated with dynamic information corresponding to the landmarks

wherein the dynamic information comprises historical information; and

identifying an orientation of a user of the client device relative to the one or more points of interest,

wherein identifying the orientation of the user comprises determining the geolocation of the client device;

identifying, based on processing the user input and the sensor data, a natural language response to the user input, wherein the natural language response provides information relevant to a particular point of interest, of the points of interest, of the environment;

determining, based on the orientation of the user of the client device relative to the particular point of interest, one or more spatial audio parameters to be used to provision the natural language response to the user; and

causing the natural language response to be audibly rendered at the client device using one or more of the spatial audio parameters.

11 . The method of claim 10 , wherein the sensors comprise one or more of a geolocational position sensor, a gyroscopic sensor, and/or an accelerometer.

12 . The method of claim 10 , wherein the dynamic information is retrieved from a website, database, and/or application programming interface (API) created for the landmarks.

13 . The method of claim 10 , wherein the one or more spatial audio parameters comprise binauralized directional audio.

14 . The method of claim 10 , further comprising:

prior to identifying the one or more points of interest of the environment and orientation of the user relative to the points of interest:

determining the user input is directed towards the environment.

15 . A computer-implemented method, the method comprising:

determining that a user input has been received at a client device;

identifying, based on sensor data available to the client device from one or more sensors:

one or more points of interest of an environment in which the client device is located, and

an orientation of a user of the client device relative to the one or more points of interest,

wherein the sensor data comprises a distance between the user and the points of interest;

identifying, based on processing the user input and the sensor data, a natural language response to the user input, wherein the natural language response provides information relevant to a particular point of interest, of the points of interest, of the environment and

wherein identifying the natural language response comprises:

identifying the natural language response and one or more other natural language responses providing information relevant to other particular points of interest;

ranking, based on the distance between the user and the particular point of interest and the other particular points of interest, the natural language response and the other natural language responses; and

determining, based on the ranking, the natural language response satisfies criterion;

determining, based on the orientation of the user of the client device relative to the particular point of interest, one or more spatial audio parameters to be used to provision the natural language response to the user; and

causing the natural language response to be audibly rendered at the client device using one or more of the spatial audio parameters.

16 . A computer-implemented method, the method comprising:

determining that a user input has been received at a client device;

identifying, based on sensor data available to the client device from one or more sensors:

one or more points of interest of an environment in which the client device is located, and

an orientation of a user of the client device relative to the one or more points of interest;

identifying, based on processing the user input and the sensor data, a natural language response to the user input, wherein the natural language response provides information relevant to a particular point of interest, of the points of interest, of the environment;

determining, based on the orientation of the user of the client device relative to the particular point of interest, one or more spatial audio parameters to be used to provision the natural language response to the user,

wherein determining the spatial audio parameters comprises:

identifying preconfigured settings modified by the user corresponding to the spatial audio parameters; and

determining the spatial audio parameters at least partially based on the preconfigured settings; and

causing the natural language response to be audibly rendered at the client device using one or more of the spatial audio parameters.

17 . A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to:

determine that a user input has been received at a client device;

identify, based on sensor data available to the client device from one or more sensors:

one or more points of interest of an environment in which the client device is located,

wherein identifying the one or more points of interest of the environment comprises cross referencing a geolocation of the client device with one or more landmarks contained in an electronic map of the environment, and

the points of interest are associated with dynamic information corresponding to

the landmarks,

wherein the dynamic information comprises merchandise information; and

identify an orientation of a user of the client device relative to the one or more points of interest,

wherein identifying the orientation of the user comprises determining the geolocation of the client device;

process the user input and the sensor data to identify a natural language response to the user input, wherein the natural language response provides information relevant to a particular point of interest, of the points of interest, of the environment;

determining, based on the orientation of the user of the client device relative to the particular point of interest, one or more spatial audio parameters to be used to provision the natural language response to the user; and

causing the natural language response to be audibly rendered at the client device using one or more of the spatial audio parameters.

18 . The system of claim 17 , wherein the sensors comprise one or more of a geolocational position sensor, a gyroscopic sensor, and/or an accelerometer.

19 . A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to:

determine that a user input has been received at a client device;

identify, based on sensor data available to the client device from one or more sensors:

one or more points of interest of an environment in which the client device is located,

wherein identifying the one or more points of interest of the environment comprises cross referencing a geolocation of the client device with one or more landmarks contained in an electronic map of the environment, and

the points of interest are associated with dynamic information corresponding to

the landmarks,

wherein the dynamic information comprises accessibility information; and

identify an orientation of a user of the client device relative to the one or more points of interest,

wherein identifying the orientation of the user comprises determining the geolocation of the client device;

process the user input and the sensor data to identify a natural language response to the user input, wherein the natural language response provides information relevant to a particular point of interest, of the points of interest, of the environment;

determining, based on the orientation of the user of the client device relative to the particular point of interest, one or more spatial audio parameters to be used to provision the natural language response to the user; and

causing the natural language response to be audibly rendered at the client device using one or more of the spatial audio parameters.

20 . The system of claim 19 , wherein the sensors comprise one or more of a geolocational position sensor, a gyroscopic sensor, and/or an accelerometer.