IP Library Granted Patent US 12,223,944
Granted Patent B2
US 12,223,944 · App. 17/744,440 · Granted Feb 11, 2025

Dynamically adapting given assistant output based on a given persona assigned to an automated assistant

Inventors: Martin Baeuml (Zurich, CH); Thushan Amarasiriwardena (Alameda, CA); Roberto Pieraccini (Zurich, CH); Gianluca Martini (Zurich, CH)
Assignee: GOOGLE LLC
G10L13/10G06F40/169G06T7/20G06V20/40G06V40/20G10L13/02G10L13/08G10L15/063G10L15/1815G10L15/183G10L15/22G10L25/57H04N5/04G06T2207/10016G06T2207/30196G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,944
App. No.
17/744,440
Granted
Feb 11, 2025
Kind
B2
Abstract

Implementations relate to dynamically adapting a given assistant output based on a given persona, from among a plurality of disparate personas, assigned to an automated assistant. In some implementations, the given assistant output can be generated and subsequently adapted based on the given persona assigned to the automated assistant. In other implementations, the given assistant output can be generated specific to the given persona and without having to subsequently adapt the given assistant output to the given persona. Notably, the given assistant output can include a stream of textual content to be synthesized for audible presentation to the user, and a stream of visual cues utilized in controlling a display of a client device and/or in controlling a visualized representation of the automated assistant. Various implementations utilize large language models (LLMs), or output previously generated utilizing LLMs, to reflect the given persona in the given assistant output.

Claims (62)

1. A method implemented by one or more processors, the method comprising:

obtaining, from an online multimedia repository, video content that includes a stream of audio data for audible content of the video content and a stream of vision data for visual content of the video content;

processing, using an automatic speech recognition model, the stream of audio data for the audible content of the video content to generate a stream of textual content corresponding to one or more spoken utterances captured in the stream of audio data for the audible content of the video content;

processing, using one or more movement tracking machine learning models, the stream of vision data for the visual content of the video content to generate a stream of visual cues corresponding to one or more movements captured in the stream of vision data for the visual content of the video content;

generating, based on processing the stream of audio data and based on processing the stream of video data, a given persona training data instance to be utilized in further training an instance of a given large language model (LLM) that is specific to a given persona, from among of a plurality of disparate personas, embodied in the video content;

training the instance of the given LLM based on at least the given persona training instance; and

causing the instance of the given LLM to be utilized in subsequently processing additional streams of audio data capturing additional spoken utterances directed to an instance of an automated assistant that is assigned the given persona.

2. The method of claim 1 , wherein the online multimedia repository is a public online multimedia repository that enables a plurality of users to upload and share corresponding video content with other users.

3. The method of claim 1 , wherein the stream of audio data for the audible content of the video content captures speech of a person or character associated with the given persona embodied in the video content.

4. The method of claim 3 , wherein the stream of vision data for the visual content of the video content captures movements or gestures of the person or character associated with the given persona embodied in the video content as the one or more spoken utterances captured in the stream of audio data for the audible content of the video content are provided by the person or character.

5. The method of claim 1 , wherein the stream of audio data for the audible content of the video content and the stream of vision data for the visual content of the video content are synchronized via one or more video timestamps.

6. The method of claim 5 , wherein training the instance of the given LLM based on at least the given persona training instance comprises:

annotating, based on the one or more video timestamps, the stream of textual context with one or more visual cue timestamps that indicate when the stream of visual cues is to be utilized in controlling a display of a client device and/or for controlling a visualized representation of the instance of the automated assistant.

7. The method of claim 6 , wherein the one or more visual cue timestamps include at least a start visual cue timestamp that indicates when a given visual cue, included in the stream of visual cues, will start being utilized in controlling the display of the client device and/or for controlling the visualized representation of the instance of the automated assistant, and a stop visual cue timestamp that indicates when the given visual cue, included in the stream of visual cues, will stop being utilized in controlling the display of the client device and/or for controlling the visualized representation of the instance of the automated assistant.

8. The method of claim 1 , wherein the one or more movement tracking machine learning models comprise:

an eye gaze machine learning model trained to detect eye movement that is captured in the stream of vision data for the visual content of the video content;

a mouth movement machine learning model trained to detect mouth movement that is captured in the stream of vision data for the visual content of the video content;

a lip movement machine learning model trained to detect lip movement that is captured in the stream of vision data for the visual content of the video content;

a body movement machine learning model trained to detect body movement that is captured in the stream of vision data for the visual content of the video content; and/or

a body posture machine learning model trained to detect body posture that is captured in the stream of vision data for the visual content of the video content.

9. The method of claim 1 , wherein training the instance of the given LLM based on at least the given persona training instance comprises:

applying training instance input, of the given persona training instance, as input across the given LLM to generate predicted output, the training instance input including the stream of audio data for the audible content of the video content and/or the stream of vision data for the visual content of the video content, and the predicted output including a predicted stream of textual content and/or a predicted stream of visual cues;

comparing the predicted output to training instance output, of the given persona training instance, to generate one or more losses, the training instance output including the stream of textual content corresponding to one or more spoken utterances captured in the stream of audio data for the audible content of the video content and/or the stream of visual cues corresponding to one or more movements captured in the stream of vision data for the visual content of the video content; and

causing the given LLM to be updated based on the one or more losses.

10. The method of claim 9 , wherein comparing the predicted output to the training instance output to generate the one or more losses comprises:

generating a first loss, of the one or more losses, based on comparing the predicted stream of textual content, of the predicted output, to the stream of textual content corresponding to one or more spoken utterances captured in the stream of audio data for the audible content of the video content.

11. The method of claim 10 , wherein comparing the predicted output to the training instance output to generate the one or more losses further comprises:

generating a second loss, of the one or more losses, based on comparing the predicted stream of visual cues, of the predicted output, to the stream of visual cues corresponding to one or more movements captured in the stream of vision data for the visual content of the video content.

12. The method of claim 1 , wherein causing the instance of the given LLM to be utilized in subsequently processing the additional audio data capturing the additional spoken utterances directed to the instance of the automated assistant that is assigned the given persona comprises:

receiving a given additional stream of audio data that captures a given additional spoken utterance of a user of a client device, the given additional stream of audio data being generated by one or more microphones of the client device, and the given additional spoken utterance being directed to the instance of the automated assistant executing at least in part at the client device;

generating, based on processing the given additional stream of audio data and using the instance of the given LLM, a given assistant output that is responsive to the spoken utterance and that is specific to the given persona, wherein the given assistant output includes: (i) a given additional stream of textual content; and (ii) a given additional stream of visual cues for controlling a display of the client device responsive to the given additional spoken utterance and/or for controlling a visualized representation of the instance of the automated assistant that is visually rendered for presentation to the user via the display of the client device; and

in response to receiving the given additional stream of audio data that captures the given additional spoken utterance of the user of the client device:

causing synthesized speech audio data capturing synthesized speech corresponding to the given additional stream of textual context to be audibly rendered for presentation to the user via one or more speakers of the client device; and

causing the given additional stream of visual cues to be utilized in controlling the display of the client device and/or in controlling the visualized representation of the instance of the automated assistant.

13. The method of claim 12 , further comprising:

synchronizing, for presentation to the user, the audible rendering of the synthesized speech corresponding to the given additional stream of textual context and the utilization of the given additional stream of visual cues in controlling the display of the client device and/or in controlling the visualized representation of the instance of the automated assistant.

14. The method of claim 13 , wherein the audible rendering of the synthesized speech corresponding to the given additional stream of textual context and the utilization of the given additional stream of visual cues in controlling the display of the client device and/or in controlling the visualized representation of the instance of the automated assistant are synchronized based on one or more visual cue timestamps.

15. A system comprising:

one or more processors; and

memory storing instructions that, when executed, cause the one or more processors to perform operations, the operations comprising:

obtaining, from an online multimedia repository, video content that includes a stream of audio data for audible content of the video content and a stream of vision data for visual content of the video content;

processing, using an automatic speech recognition model, the stream of audio data for the audible content of the video content to generate a stream of textual content corresponding to one or more spoken utterances captured in the stream of audio data for the audible content of the video content;

processing, using one or more movement tracking machine learning models, the stream of vision data for the visual content of the video content to generate the stream of visual cues;

generating, based on processing the stream of audio data and based on processing the stream of video data, a given persona training data instance to be utilized in further training an instance of a given large language model (LLM) that is specific to a given persona, from among of a plurality of disparate personas, embodied in the video content;

training the instance of the given LLM based on at least the given persona training instance; and

causing the instance of the given LLM to be utilized in subsequently processing additional audio data capturing additional spoken utterances directed to an instance of an automated assistant that is assigned the given persona.

16. The system of claim 15 , wherein the stream of audio data for the audible content of the video content captures speech of a person or character associated with the given persona embodied in the video content.

17. The system of claim 15 , wherein the stream of audio data for the audible content of the video content and the stream of vision data for the visual content of the video content are synchronized via one or more video timestamps.

18. The system of claim 15 , wherein the one or more movement tracking machine learning models comprise:

an eye gaze machine learning model trained to detect eye movement that is captured in the stream of vision data for the visual content of the video content;

a mouth movement machine learning model trained to detect mouth movement that is captured in the stream of vision data for the visual content of the video content;

a lip movement machine learning model trained to detect lip movement that is captured in the stream of vision data for the visual content of the video content;

a body movement machine learning model trained to detect body movement that is captured in the stream of vision data for the visual content of the video content; and/or

a body posture machine learning model trained to detect body posture that is captured in the stream of vision data for the visual content of the video content.

19. A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor to perform operations, the operations comprising:

obtaining, from an online multimedia repository, video content that includes a stream of audio data for audible content of the video content and a stream of vision data for visual content of the video content;

processing, using an automatic speech recognition model, the stream of audio data for the audible content of the video content to generate a stream of textual content corresponding to one or more spoken utterances captured in the stream of audio data for the audible content of the video content;

processing, using one or more movement tracking machine learning models, the stream of vision data for the visual content of the video content to generate a stream of visual cues corresponding to one or more movements captured in the stream of vision data for the visual content of the video content;

generating, based on processing the stream of audio data and based on processing the stream of video data, a given persona training data instance to be utilized in further training an instance of a given large language model (LLM) that is specific to a given persona, from among of a plurality of disparate personas, embodied in the video content;

training the instance of the given LLM based on at least the given persona training instance; and

causing the instance of the given LLM to be utilized in subsequently processing additional streams of audio data capturing additional spoken utterances directed to an instance of an automated assistant that is assigned the given persona.

20. The system of claim 15 , wherein the online multimedia repository is a public online multimedia repository that enables a plurality of users to upload and share corresponding video content with other users.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2022
From: BAEUML, MARTIN; AMARASIRIWARDENA, THUSHAN; PIERACCINI, ROBERTO; MARTINI, GIANLUCA
To: GOOGLE LLC
Reel/Frame 060373/0938 →
Continuity (2)
Continuation 17726244 · Apr 21, 2022
Related Publication 20230343324A1 · Oct 26, 2023
References Cited (17)
US 11605387B1 · Muralitharan · 2023 [cited by examiner]
US 20150172262A1 · Ortiz et al. · 2015 [cited by applicant]
US 20190341050A1 · Diamant · 2019 [cited by examiner]
US 20200137230A1 · Spohrer · 2020 [cited by applicant]
US 20200302927A1 · Andruszkiewicz · 2020 [cited by applicant]
US 20200387677A1 · Kim · 2020 [cited by examiner]
US 20210104236A1 · Doggett · 2021 [cited by applicant]
US 20210124422A1 · Forsland · 2021 [cited by examiner]
US 20230127090A1 · Eirinberg · 2023 [cited by examiner]
US 20230259540A1 · Das · 2023 [cited by examiner]
US 20230260326A1 · Zhou · 2023 [cited by examiner]
US 20230343323A1 · Baeuml · 2023 [cited by applicant]
Mufin: Enriching Semantic Understanding of Sentence Embedding using Dual Tune Framework (Year: 2021). [cited by examiner]
European Patent Office; International Search Report and Written Opinion issued in Application No. PCT/US2022/047027; 24 pages; dated Mar. 21, 2023. [cited by applicant]
Goswami, K. et al.; Mufin: Enriching Semantic Understanding of Sentence Embedding using Dual Tune Framework; IEEE; 6 pages; dated 2021. [cited by applicant]
Boseop, K. et al., “What Changes Can Large-scale Language Models Bring?” Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers; arXiv.org, Cornell University Library; arXiv:2109.04650v2… [cited by applicant]
European Patent Office; Invitation to Pay Additional Fees issued in Application No. PCT/US2022/047027; 18 pages; dated Jan. 25, 2023. [cited by applicant]