IP Library Granted Patent US 12,601,607
Granted Patent B2
US 12,601,607 · App. 18/740,870 · Granted Apr 14, 2026

Content-aware navigation instructions

Inventors: Victor Carbune (Zurich, CH); Matthew Sharifi (Zurich, CH)
Assignee: Google LLC
G01C21/3629G01C21/3655G10L15/05G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,601,607
App. No.
18/740,870
Granted
Apr 14, 2026
Kind
B2
Abstract

To provide content-aware audio navigation instructions, a client device executing a mapping application obtains one or more audio navigation directions for traversing from a starting location to a destination location along a route. The client device also identifies electronic media content playing from a source different from the mapping application which is executing at the client device or in proximity to the client device. The client device determines characteristics of the electronic media content and adjusts the audio navigation directions in accordance with the characteristics of the electronic media content. Then the client device presents the adjusted audio navigation directions to a user.

Claims (84)

1 . A method for training a machine learning model to generate content-aware navigation instructions, the method comprising:

obtaining, by one or more processors, a plurality of sets of audio navigation instructions previously provided to users;

obtaining, by the one or more processors, characteristics of electronic media content playing when the plurality of sets of audio navigation instructions were presented;

obtaining, by the one or more processors, at least one of: indications of adjustments made by users to the plurality of sets of audio navigation instructions, or indications regarding the users' satisfaction with the plurality of sets of audio navigation instructions;

training, by the one or more processors, a machine learning model for adjusting audio navigation instructions based on electronic media content using (i) the plurality of sets of audio navigation instructions previously provided to the users, (ii) the characteristics of the electronic media content playing when the plurality of sets of audio navigation instructions were presented, and (iii) at least one of: the indications of adjustments made by the users to the plurality of sets of audio navigation instructions, or the indications regarding the users' satisfaction with the plurality of sets of audio navigation instructions; and

providing, by the one or more processors, the trained machine learning model for adjusting a set of audio navigation instructions.

2 . The method of claim 1 , wherein training the machine learning model for adjusting audio navigation instructions includes training the machine learning model for adjusting at least one of:

a timing in which the audio navigation instructions are presented,

a language in which the audio navigation instructions are presented,

a speed at which the audio navigation instructions are presented, or

a recommendation for a point of interest (POI) along a route for the audio navigation instructions.

3 . The method of claim 1 , wherein obtaining characteristics of the electronic media content includes at least one of:

obtaining a speed in which the electronic media content is presented,

obtaining a language in which the electronic media content is presented,

obtaining a pause in the electronic media content, or

obtaining a point of interest (POI) or geographical topic which is included in the electronic media content.

4 . The method of claim 1 , wherein the characteristics of the electronic media content are obtained from sources different from mapping applications that presented the plurality of sets of audio navigation instructions.

5 . The method of claim 1 , wherein the indications of adjustments made by users to the plurality of sets of audio navigation instructions include at least one of:

a change to a language for one of the plurality of sets of audio navigation instructions;

a change to a rate of speech for one of the plurality of sets of audio navigation instructions;

a change to a volume for one of the plurality of sets of audio navigation instructions; or

a request to repeat an audio navigation instruction in one of the plurality of sets of audio navigation instructions.

6 . The method of claim 1 , wherein training the machine learning model for adjusting audio navigation instructions includes:

training, by the one or more processors, a first machine learning model for determining a language for the audio navigation instructions;

training, by the one or more processors, a second machine learning model for determining a timing for providing the audio navigation instructions; and

training, by the one or more processors, a third machine learning model for determining a rate of speech for the audio navigation instructions.

7 . The method of claim 1 , wherein training the machine learning model using the plurality of sets of audio navigation instructions previously provided to the users including:

training, by the one or more processors, the machine learning model using audio navigation instruction parameters for the plurality of sets of audio navigation instructions including at least one of:

a maneuver type for maneuvers in the plurality of sets of audio navigation instructions,

a location of the maneuvers in the plurality of sets of audio navigation instructions,

a complexity level of the maneuvers in the plurality of sets of audio navigation instructions, or

an urgency level for playing the plurality of sets of audio navigation instructions.

8 . A method for generating content-aware navigation instructions, the method comprising:

obtaining, by one or more processors in a client device, one or more audio navigation directions for traversing from a starting location to a destination location along a route;

identifying, by the one or more processors, electronic media content playing;

determining, by the one or more processors, characteristics of the electronic media content;

adjusting, by the one or more processors, at least one of the one or more audio navigation directions by applying the characteristics of the electronic media content to a trained machine learning model for adjusting audio navigation instructions based on electronic media content; and

presenting, by the one or more processors, the at least one adjusted audio navigation direction to the user.

9 . The method of claim 8 , wherein adjusting at least one of the one or more audio navigation directions includes at least one of:

adjusting a timing in which the at least one audio navigation direction is presented,

adjusting a language in which the at least one audio navigation direction is presented,

adjusting a speed at which the at least one audio navigation direction is presented, or

providing a recommendation for a point of interest (POI) along the route.

10 . The method of claim 8 , wherein determining characteristics of the electronic media content includes at least one of:

determining a speed in which the electronic media content is presented,

determining a language in which the electronic media content is presented,

identifying a pause in the electronic media content, or

identifying a point of interest (POI) or geographical topic which is included in the electronic media content.

11 . The method of claim 8 , wherein:

the one or more audio navigation directions are obtained by the client device via a mapping application, and

identifying electronic media content includes identifying, by the one or more processors, the electronic media content playing from a source different from the mapping application, the source executing at the client device or in proximity with the client device.

12 . The method of claim 11 , wherein identifying electronic media content playing from a source different from the mapping application includes at least one of:

obtaining, by the one or more processors, audio playback data from an audio application executing on the client device which is different from the mapping application;

obtaining, by the one or more processors, audio playback data from a device communicatively coupled to the client device; or

comparing, by the one or more processors, ambient audio fingerprints to one or more audio fingerprints of predetermined media content.

13 . The method of claim 8 , wherein the trained machine learning model is trained using (i) a plurality of sets of audio navigation instructions previously provided to users, (ii) characteristics of electronic media content playing when the plurality of sets of audio navigation instructions were presented, and (iii) at least one of: indications of adjustments made by the users to the plurality of sets of audio navigation instructions, or indications regarding the users' satisfaction with the plurality of sets of audio navigation instructions.

14 . The method of claim 13 , wherein the trained machine learning model includes a first machine learning model for determining a language for the audio navigation instructions, a second machine learning model for determining a timing for providing the audio navigation instructions, and a third machine learning model for determining a rate of speech for the audio navigation instructions.

15 . A client device for generating content-aware navigation instructions, the client device comprising:

a speaker;

one or more processors; and

a non-transitory computer-readable memory coupled to the one or more processors and the speaker and storing instructions thereon that, when executed by the one or more processors, cause the client device to:

obtain one or more audio navigation directions for traversing from a starting location to a destination location along a route;

identify electronic media content playing;

determine characteristics of the electronic media content;

adjust at least one of the one or more audio navigation directions by applying the characteristics of the electronic media content to a trained machine learning model for adjusting audio navigation instructions based on electronic media content; and

present, via the speaker, the at least one adjusted audio navigation direction to the user.

16 . The client device of claim 15 , wherein to adjust the at least one audio navigation direction, the instructions cause the client device to at least one of:

adjust a timing in which the at least one audio navigation direction is presented,

adjust a language in which the at least one audio navigation direction is presented,

adjust a speed at which the at least one audio navigation direction is presented, or

provide a recommendation for a point of interest (POI) along the route.

17 . The client device of claim 15 , wherein to determine characteristics of the electronic media content, the instructions cause the client device to at least one of:

determine a speed in which the electronic media content is presented,

determine a language in which the electronic media content is presented,

identify a pause in the electronic media content, or

identify a point of interest (POI) or geographical topic which is included in the electronic media content.

18 . The client device of claim 15 , wherein:

the one or more audio navigation directions are obtained by the client device via a mapping application, and

the electronic media content is playing from a source different from the mapping application, the source executing at the client device or in proximity with the client device.

19 . The client device of claim 18 , wherein to identify electronic media content playing from a source different from the mapping application, the instructions cause the client device to at least one of:

obtain audio playback data from an audio application executing on the client device which is different from the mapping application;

obtain audio playback data from a device communicatively coupled to the client device; or

compare ambient audio fingerprints to one or more audio fingerprints of predetermined media content.

20 . The client device of claim 15 , wherein the trained machine learning model is trained using (i) a plurality of sets of audio navigation instructions previously provided to users, (ii) characteristics of electronic media content playing when the plurality of sets of audio navigation instructions were presented, and (iii) at least one of: indications of adjustments made by the users to the plurality of sets of audio navigation instructions, or indications regarding the users' satisfaction with the plurality of sets of audio navigation instructions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2024
From: CARBUNE, VICTOR; SHARIFI, MATTHEW
To: GOOGLE LLC
Reel/Frame 067740/0821 →
Continuity (2)
Continuation 17273672
Related Publication 20240328808A1 · Oct 3, 2024
References Cited (35)
US 9087508B1 · Dzik · 2015 [cited by applicant]
US 9354074B2 · Azose et al. · 2016 [cited by applicant]
US 9576574B2 · van Os · 2017 [cited by applicant]
US 10137902B2 · Juneja et al. · 2018 [cited by applicant]
US 10261564B1 · Gollakota · 2019 [cited by applicant]
US 10701510B2 · Norris et al. · 2020 [cited by applicant]
US 11725957B2 · Padegimaite · 2023 [cited by examiner]
US 20110301728A1 · Hamilton et al. · 2011 [cited by applicant]
US 20150106012A1 · Kandangath et al. · 2015 [cited by applicant]
US 20150276421A1 · Beaurepaire et al. · 2015 [cited by applicant]
US 20160259027A1 · Said · 2016 [cited by applicant]
US 20180283891A1 · Andrew et al. · 2018 [cited by applicant]
US 20180335312A1 · Bennett et al. · 2018 [cited by applicant]
US 20200044757A1 · Modi · 2020 [cited by applicant]
US 20200329333A1 · Norris et al. · 2020 [cited by applicant]
US 20210404833A1 · Padegimaite et al. · 2021 [cited by applicant]
US 20230168101A1 · Carbune et al. · 2023 [cited by applicant]
US 20240102816A1 · Sharifi · 2024 [cited by examiner]
CN 107036614A · 2017 [cited by applicant]
CN 110140332A · 2019 [cited by applicant]
CN 110785630A · 2020 [cited by applicant]
CN 111693064A · 2020 [cited by applicant]
JP 2010117476A · 2010 [cited by applicant]
JP 2011095142A · 2011 [cited by applicant]
JP 2019117324A · 2019 [cited by applicant]
JP 2020138314A · 2020 [cited by applicant]
WO WO2013184473A2 · 2013 [cited by applicant]
WO WO2015103457A2 · 2015 [cited by applicant]
WO WO2020091806A1 · 2020 [cited by applicant]
Katz et al., B.F. Navig: Augmented Reality Guidance System for the Visually Impaired: Combining object localization, GNSS, and spatial audio, Google Scholar, Virtual Reality, vol. 16, Jun. 2012, pp. 253-269. (Year: 2012… [cited by examiner]
Office Action for European Patent Application No. 20807581.2 dated Oct. 6, 2025. 7 pages. [cited by applicant]
First Office Action for Chinese Application No. 202080106401.9, dated Sep. 25, 2024. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2020/056756, dated Jun. 23, 2021. [cited by applicant]
Japanese Patent Application No. 2023-519747, Notice of Reasons for Rejection, mailing date of Apr. 22, 2024. [cited by applicant]
Korean Office Action for Application No. 10-2023-7012688, dated Mar. 12, 2025. [cited by applicant]