IP Library Granted Patent US 12711686
Granted Patent B2
US 12711686 · App. 18/223,195 · Granted Aug 18, 2026

Avatar UI with multiple speaking actions for selected text

Inventors: Erkan Volkan (Lodi, NJ); Siva Penke (San Jose, CA); Siva Boggala (San Ramon, CA); Laszlo Gombos (Winchester, MA); Jisun Park (Palo Alto, CA)
Assignee: Samsung Electronics Co., Ltd.
G06T13/40G06T13/205G06T17/00G10L13/02G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711686
App. No.
18/223,195
Granted
Aug 18, 2026
Kind
B2
Abstract

In one embodiment, a method includes determining, by a client computing device, that a user has selected text displayed on a display of the client computing device. The method further includes presenting, in response to the determination, a UI element on the display of the client computing device. The UI element includes a plurality of selectable portions, each associated with a distinct speaking action for a 3D avatar to perform with respect to the selected text. In response to the user's selection, the method includes presenting on the display of the client computing device an animation of the 3D avatar performing a speaking action corresponding to the selected portion; and providing, by the client computing device, speech audio synchronized with the speaking action of the animated 3D avatar.

Claims (50)

1 . A method comprising:

displaying text on a display of a client computing device;

after displaying the text on the display of the client computing device, then determining, by a processor of the client computing device detecting a text-selection function comprising at least one of (1) a user interaction with the display of the client computing device or (2) an input to the client computing device by a human-interface device, that a user has selected at least a portion of the text already being displayed on the display of the client computing device;

presenting, in response to the determination that the user has selected text already being displayed on the display of the computing device, a UI element on the display of the client computing device, the UI element comprising a plurality of selectable portions, each selectable portion associated with a distinct speaking action for a 3D avatar to perform with respect to the selected text when that respective selectable portion is selected by the user, wherein (1) the UI element is displayed only after the user has selected the text already being displayed on the display of the client computing device and (2) the UI element lets the user select, via the plurality of selectable portions, which distinct speaking action the 3D avatar will perform regarding the text selected by the user,

in response to the user's selection of a particular one of the plurality of selectable UI portions corresponding to a speaking action, presenting on the display of the client computing device an animation of the 3D avatar performing the speaking action corresponding to the selected portion; and

providing, by the client computing device, speech audio synchronized with the speaking action of the animated 3D avatar.

2 . The method of claim 1 , wherein one of the plurality of distinct speaking actions comprises an AI-view action that corresponds to the 3D avatar speaking a comment generated by a trained machine-learning model from the selected text.

3 . The method of claim 2 , wherein the AI-view action further corresponds to the 3D avatar speaking the selected text after speaking the generated comment.

4 . The method of claim 2 , wherein, when the UI element corresponding to the AI-view action is selected, the method further comprises:

providing the selected text to the trained machine learning model;

accessing, from the trained machine learning model, output text paraphrasing the selected text;

generating, from the selected text, one or more feature vectors representing an emotional content of the selected text;

providing the output text and the one or more feature vectors to a transformer model; and

receiving, from the transformer model, the comment.

5 . The method of claim 1 , wherein one of the plurality of distinct speaking actions comprises a speak action that corresponds to the 3D avatar speaking the selected text.

6 . The method of claim 1 , wherein one of the plurality of distinct speaking actions comprises a pronounce action that corresponds to an enhanced view of the 3D avatar's mouth while speaking the selected text.

7 . The method of claim 1 , wherein one of the plurality of distinct speaking actions comprises a read-from-here action that corresponds to the 3D avatar speaking the selected text and at least a portion of the subsequent text following the selected text.

8 . The method of claim 1 , wherein the UI element further comprises a share action that corresponds to recording a video of the 3D avatar performing a speaking action with respect to the selected text.

9 . The method of claim 1 , wherein the UI element further comprises a second avatar displayed in connection with the plurality of selectable portions.

10 . The method of claim 9 , wherein the second avatar comprises a relatively smaller view of the 3D avatar.

11 . The method of claim 1 , wherein the 3D avatar is presented within a threshold distance of the selected text.

12 . The method of claim 1 , wherein the selected text is displayed on a web browser executing on the client computing device.

13 . One or more non-transitory computer readable storage media storing software that is operable when executed by one or more processors to:

display text on a display of a client computing device;

after displaying the text on the display of the client computing device, then determine, by detecting a text-selection function comprising at least one of (1) a user interaction with the display of the client computing device or (2) an input to the client computing device by a human-interface device, that a user has selected at least a portion of the text already being displayed on the display of the client computing device;

present, in response to the determination that the user has selected text already being displayed on the display of the computing device, a UI element on the display of the client computing device, the UI element comprising a plurality of selectable portions, each selectable portion associated with a distinct speaking action for a 3D avatar to perform with respect to the selected text when that respective selectable portion is selected by the user, wherein (1) the UI element is displayed only after the user has selected the text already being displayed on the display of the client computing device and (2) the UI element lets the user select, via the plurality of selectable portions, which distinct speaking action the 3D avatar will perform regarding the text selected by the user;

in response to the user's selection of a particular one of the plurality of selectable UI portions corresponding to a speaking action, present on the display of the client computing device an animation of the 3D avatar performing the speaking action corresponding to the selected portion; and

provide speech audio synchronized with the speaking action of the animated 3D avatar.

14 . The media of claim 13 , wherein one of the plurality of distinct speaking actions comprises an AI-view action that corresponds to the 3D avatar speaking a comment generated by a trained machine-learning model from the selected text.

15 . The media of claim 14 , wherein the AI-view action further corresponds to the 3D avatar speaking the selected text after speaking the generated comment.

16 . The media of claim 14 , wherein, when the UI element corresponding to the AI-view action is selected, the software is further operable when executed by one or more processors to:

provide the selected text to the trained machine learning model;

access, from the trained machine learning model, output text paraphrasing the selected text;

generate, from the selected text, one or more feature vectors representing an emotional content of the selected text;

provide the output text and the one or more feature vectors to a transformer model; and

receive, from the transformer model, the comment.

17 . A system comprising one or more non-transitory computer readable storage media storing instructions; and one or more processors coupled to the non-transitory computer readable storage media, the one or more processors operable to execute the instructions to:

display text on a display of a client computing device;

after displaying the text on the display of the client computing device, then determine, by detecting a text-selection function comprising at least one of (1) a user interaction with the display of the client computing device or (2) an input to the client computing device by a human-interface device, that a user has selected at least a portion of the text already being displayed on the display of the client computing device;

present, in response to the determination that the user has selected text already being displayed on the display of the computing device, a UI element on the display of the client computing device, the UI element comprising a plurality of selectable portions, each selectable portion associated with a distinct speaking action for a 3D avatar to perform with respect to the selected text when that respective selectable portion is selected by the user, wherein (1) the UI element is displayed only after the user has selected the text already being displayed on the display of the client computing device and (2) the UI element lets the user select, via the plurality of selectable portions, which distinct speaking action the 3D avatar will perform regarding the text selected by the user;

in response to the user's selection of a particular one of the plurality of selectable UI portions corresponding to a speaking action, present on the display of the client computing device an animation of the 3D avatar performing the speaking action corresponding to the selected portion; and

provide speech audio synchronized with the speaking action of the animated 3D avatar.

18 . The system of claim 17 , wherein one of the plurality of distinct speaking actions comprises an AI-view action that corresponds to the 3D avatar speaking a comment generated by a trained machine-learning model from the selected text.

19 . The system of claim 18 , wherein the AI-view action further corresponds to the 3D avatar speaking the selected text after speaking the generated comment.

20 . The system of claim 18 , wherein, when the UI element corresponding to the AI-view action is selected, the one or more processors are further operable to execute the instructions to:

provide the selected text to the trained machine learning model;

access, from the trained machine learning model, output text paraphrasing the selected text;

generate, from the selected text, one or more feature vectors representing an emotional content of the selected text;

provide the output text and the one or more feature vectors to a transformer model; and

receive, from the transformer model, the comment.