IP Library Granted Patent US 12,518,749
Granted Patent B2
US 12,518,749 · App. 18/446,798 · Granted Jan 6, 2026

Adaptive sending or rendering of audio with text messages sent via automated assistant

Inventors: Victor Carbune (Zurich, CH); Matthew Sharifi (Kilchberg, CH)
Assignee: GOOGLE LLC
G10L15/19G10L15/22G10L21/028G10L25/51G10L15/063G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,749
App. No.
18/446,798
Granted
Jan 6, 2026
Kind
B2
Abstract

Implementations set forth herein relate to an automated assistant that can selectively communicate audio data to a recipient when a user solicits the automated assistant to send a text message to the recipient. The audio data can include a snippet of audio that characterizes content of the text message, and the automated assistant can communicate the audio data to the recipient when score data for a speech recognition hypothesis does not satisfy a confidence threshold. The score data can correspond to an entirety of content of a text message and/or speech recognition hypothesis, and/or less than an entirety of the content. A recipient device can optionally re-process the audio data using a model that is associated with the recipient device. This can provide more accurate transcripts in some instances, thereby improving accuracy of communications and decreasing a number of corrective messages sent between users.

Claims (70)

1 . A method implemented by one or more processors, the method comprising:

receiving, at a recipient computing device, textual data and audio data corresponding to a message communicated from a user of an automated assistant,

wherein the user provided a spoken utterance to the automated assistant in furtherance of causing the message to be communicated to the recipient;

processing the textual data and the audio data in furtherance of generating additional textual data,

wherein the textual data characterizes content of the spoken utterance and the additional textual data characterizes other content of the audio data;

determining, based on the textual data and the additional textual data, that the additional textual data characterizes the other content with a higher confidence than the textual data characterizes the content of the spoken utterance;

causing, based on the textual data and the additional textual data, the additional textual data to be rendered at the recipient computing device;

determining contextual data associated with the user,

wherein the contextual data indicates speech of a separate person included in the audio data is relevant to the textual data or the additional textual data; and

causing the audio data to be rendered at the recipient computing device based on the contextual data.

2 . The method of claim 1 ,

wherein the contextual data further indicates a level of background noise apparent at the recipient computing device.

3 . The method of claim 1 , further comprising:

receiving score data that indicates a degree of confidence that the textual data characterizes the content of the spoken utterance; and

generating, based on the additional textual data, additional score data that characterizes a separate degree of confidence that the additional textual data characterizes the other content of the audio data,

wherein the higher confidence is determined based at least on the score data and the additional score data.

4 . A method implemented by one or more processors, the method comprising:

receiving, by an automated assistant, a spoken utterance from a user in furtherance of directing the automated assistant to send a text message to a separate device,

wherein a computing device that provides access to the automated assistant captures audio data corresponding to the spoken utterance;

generating, based on the audio data captured by the computing device, a speech recognition hypothesis that characterizes textual content of the spoken utterance;

generating, based on the audio data, speech diarization data that characterizes additional information associated with the spoken utterance;

determining, based on the speech diarization data, an amount of text of the speech recognition hypothesis to include in the text message to the separate device;

determining, based on the speech diarization data, a duration of audio of the spoken utterance to send, as audio data, to the separate device; and

causing, by the automated assistant, the text message and the audio data to be communicated to the separate device,

wherein the amount of text included with the text message is less than an entirety of content of the spoken utterance, and the duration of the audio of the spoken utterance is less than an entire duration of the spoken utterance received by the automated assistant.

5 . The method of claim 4 ,

wherein the additional information indicates that an intonation of the user changed during a portion of the spoken utterance, and

wherein the amount of text does not include other content of the spoken utterance during a change in the intonation of the user.

6 . The method of claim 4 ,

wherein the additional information indicates that a separate person was speaking during a portion of the spoken utterance, and

wherein the duration of the audio includes other content of the spoken utterance when the separate person was speaking.

7 . The method of claim 4 ,

wherein the additional information indicates a degree of confidence that a particular portion of the speech recognition hypothesis characterizes a corresponding portion of the spoken utterance, and

wherein the amount of text does not include other content of the particular portion of the speech recognition hypothesis.

8 . The method of claim 7 , wherein the duration of the audio includes other audio content of the spoken utterance corresponding to the particular portion of the speech recognition hypothesis.

9 . The method of claim 8 , further comprising:

causing the other audio content to be re-processed at the separate device in furtherance of generating other textual content to replace the particular portion of the speech recognition hypothesis,

wherein the separate device renders the text message with the other textual content.

10 . The method of claim 9 , further comprising:

causing the separate device to generate score data that indicates a degree of confidence that the other textual content accurately characterizes natural language content of the other audio content,

wherein the separate device renders the text message with the other textual content based on the score data satisfying a threshold confidence score.

11 . The method of claim 10 , wherein the score data is generated using a speech processing model that is trained from training data that is at least partially based on interactions between another user and the separate device.

12 . A system comprising:

memory storing instructions;

one or more processors operable to execute the instructions to:

receive, by an automated assistant, a spoken utterance from a user in furtherance of directing the automated assistant to send a text message to a separate device,

wherein a computing device that provides access to the automated assistant captures audio data corresponding to the spoken utterance;

generate, based on the audio data captured by the computing device, a speech recognition hypothesis that characterizes textual content of the spoken utterance;

generate, based on the audio data, speech diarization data that characterizes additional information associated with the spoken utterance;

determine, based on the speech diarization data, an amount of text of the speech recognition hypothesis to include in the text message to the separate device;

determine, based on the speech diarization data, a duration of audio of the spoken utterance to send, as audio data, to the separate device; and

cause, by the automated assistant, the text message and the audio data to be communicated to the separate device,

wherein the amount of text included with the text message is less than an entirety of content of the spoken utterance, and the duration of the audio of the spoken utterance is less than an entire duration of the spoken utterance received by the automated assistant.

13 . The system of claim 12 ,

wherein the additional information indicates that an intonation of the user changed during a portion of the spoken utterance, and

wherein the amount of text does not include other content of the spoken utterance during a change in the intonation of the user.

14 . The system of claim 12 ,

wherein the additional information indicates that a separate person was speaking during a portion of the spoken utterance, and

wherein the duration of the audio includes other content of the spoken utterance when the separate person was speaking.

15 . The system of claim 12 ,

wherein the additional information indicates a degree of confidence that a particular portion of the speech recognition hypothesis characterizes a corresponding portion of the spoken utterance, and

wherein the amount of text does not include other content of the particular portion of the speech recognition hypothesis.

16 . The system of claim 15 , wherein the duration of the audio includes other audio content of the spoken utterance corresponding to the particular portion of the speech recognition hypothesis.

17 . The system of claim 16 , wherein one or more of the processors are further operable to execute the instructions to:

cause the other audio content to be re-processed at the separate device in furtherance of generating other textual content to replace the particular portion of the speech recognition hypothesis,

wherein the separate device renders the text message with the other textual content.

18 . The system of claim 17 , wherein one or more of the processors are further operable to execute the instructions to:

cause the separate device to generate score data that indicates a degree of confidence that the other textual content accurately characterizes natural language content of the other audio content,

wherein the separate device renders the text message with the other textual content based on the score data satisfying a threshold confidence score.

19 . The system of claim 18 , wherein the score data is generated using a speech processing model that is trained from training data that is at least partially based on interactions between another user and the separate device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2023
From: CARBUNE, VICTOR; SHARIFI, MATTHEW
To: GOOGLE LLC
Reel/Frame 064560/0460 →
Continuity (1)
Related Publication 20250054495A1 · Feb 13, 2025
References Cited (11)
US 11636851B2 · Mahmood · 2023 [cited by examiner]
US 11783828B2 · Sharifi · 2023 [cited by examiner]
US 11977816B1 · Sepasi Ahoei · 2024 [cited by examiner]
US 12211517B1 · Maas · 2025 [cited by examiner]
US 20140245140A1 · Brown · 2014 [cited by examiner]
US 20170263248A1 · Gruber · 2017 [cited by examiner]
US 20170270927A1 · Brown · 2017 [cited by examiner]
US 20210258641A1 · Konzelmann · 2021 [cited by examiner]
US 20220293109A1 · Sharifi · 2022 [cited by examiner]
US 20220319219A1 · Tsibulevskiy · 2022 [cited by examiner]
US 20230025709A1 · Sharifi · 2023 [cited by examiner]