IP Library Granted Patent US 12,028,302
Granted Patent B2
US 12,028,302 · App. 18/215,221 · Granted Jul 2, 2024

Assistance during audio and video calls

Inventors: Fredrik Bergenlid (Mountain View, CA); Vladyslav Lysychkin (Zurich, CH); Denis Burakov (Zurich, CH); Behshad Behzadi (Mountain View, CA); Andrea Terwisscha Van Scheltinga (Zurich, CH); Quentin Lascombes De Laroussilhe (Zurich, CH); Mikhail Golikov (Mountain View, CA); Koa Metter (Seattle, WA); Ibrahim Badr (Mountain View, CA); Zaheed Sabur (Zurich, CH)
Assignee: Google LLC
H04L51/10G06F16/44G10L15/22H04N7/15H04N21/4394H04N21/4788G10L15/005G10L15/16G10L2015/223G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,028,302
App. No.
18/215,221
Granted
Jul 2, 2024
Kind
B2
Abstract

Implementations relate to providing information items for display during a communication session. In some implementations, a computer-implemented method includes receiving, during a communication session between a first computing device and a second computing device, first media content from the communication session. The method further includes determining a first information item for display in the communication session based at least in part on the first media content. The method further includes sending a first command to at least one of the first computing device and the second computing device to display the first information item.

Claims (59)

1. A computer-implemented method comprising:

receiving session content during a video communication session between a first computing device and a second computing device;

determining, based on the session content, that the session content includes a request for media from the first computing device;

outputting, with a trained machine-learning model, context from the session content;

altering the media based on the context; and

sending a first command to at least one of the first computing device or the second computing device to display the altered media.

2. The method of claim 1 , wherein:

the context includes an identification of emotions in the session content; and

altering the media based on the context includes altering the media based on the emotions in the session content.

3. The method of claim 1 , further comprising:

outputting, with the trained machine-learning model, a speech-to-text conversion of an audio portion of the session content.

4. The method of claim 1 , wherein the trained machine-learning model is personalized for a particular user by:

modifying one or more parameters for the trained machine-learning model based on the particular user;

saving the modified parameters for the trained machine-learning model; and

initializing the trained machine-learning model using the modified parameters for the particular user.

5. The method of claim 1 , wherein the trained machine-learning model is trained using training data and audio and video sources and wherein audio or video conversations that are not from the audio and video sources are excluded from the training data.

6. The method of claim 1 , wherein the trained machine-learning model outputs a determination that the session content includes the request for media based on a user associated with the first computing device making an explicit invocation for assistance in the session content.

7. The method of claim 6 , wherein the trained machine-learning model outputs the determination that the session content includes the request for media based on at least one selected from the group of a gender of a user associated with the first computing device, an estimated age of the user, an identification of the user being indoors or outdoors, a language spoken by the user, and combinations thereof.

8. The method of claim 1 , further comprising:

outputting, with the trained machine-learning model, at least one determination selected from the group of whether a face is present in the session content, a position of the face in the session content, a number of faces in the session content, and combinations thereof.

9. The method of claim 1 , further comprising:

analyzing the session content using a video analysis technique to determine that a user associated with the first computing device explicitly requests the media;

wherein the video analysis technique is at least one selected from the group of face detection, motion detection, gesture detection, and combinations thereof.

10. The method of claim 1 , further comprising:

generating, with the trained machine-learning model, the media in response to the request for the media from the first computing device.

11. A computing device comprising:

a processor; and

a memory coupled to the processor, with instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising:

receiving session content during a video communication session between a first computing device and a second computing device;

determining, based on the session content, that the session content includes a request for media from the first computing device;

outputting, with a trained machine-learning model, context from the session content;

altering the media based on the context; and

sending a first command to at least one of the first computing device or the second computing device to display the altered media.

12. The computing device of claim 11 , wherein:

the context includes an identification of emotions in the session content; and

altering the media based on the context includes altering the media based on the emotions in the session content.

13. The computing device of claim 11 , wherein the operations further comprise:

outputting, with the trained machine-learning model, a speech-to-text conversion of an audio portion of the session content.

14. The computing device of claim 11 , wherein the trained machine-learning model is personalized for a particular user by:

modifying one or more parameters for the trained machine-learning model based on the particular user;

saving the modified parameters for the trained machine-learning model; and

initializing the trained machine-learning model using the modified parameters for the particular user.

15. The computing device of claim 11 , wherein the trained machine-learning model is trained using training data and audio and video sources and wherein audio or video conversations that are not from the audio and video sources are excluded from the training data.

16. A non-transitory computer-readable medium with instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:

receiving session content during a video communication session between a first computing device and a second computing device;

determining, based on the session content, that the session content includes a request for media from the first computing device;

outputting, with a trained machine-learning model, context from the session content;

altering the media based on the context; and

sending a first command to at least one of the first computing device or the second computing device to display the altered media.

17. The non-transitory computer-readable medium of claim 16 , wherein:

the context includes an identification of emotions in the session content; and

altering the media based on the context includes altering the media based on the emotions in the session content.

18. The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise:

outputting, with the trained machine-learning model, a speech-to-text conversion of an audio portion of the session content.

19. The non-transitory computer-readable medium of claim 16 , wherein the trained machine-learning model is personalized for a particular user by:

modifying one or more parameters for the trained machine-learning model based on the particular user;

saving the modified parameters for the trained machine-learning model; and

initializing the trained machine-learning model using the modified parameters for the particular user.

20. The non-transitory computer-readable medium of claim 16 , wherein the trained machine-learning model is trained using training data and audio and video sources and wherein audio or video conversations that are not from the audio and video sources are excluded from the training data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2023
From: BERGENLID, FREDRIK; LYSYCHKIN, VLADYSLAV; BURAKOV, DENIS; BEHZADI, BEHSHAD; VAN SCHELTINGA, ANDREA TERWISSCHA; DE LAROUSSILHE, QUENTIN LASCOMBES; GOLIKOV, MIKHAIL; METTER, KOA; BADR, IBRAHIM; SABUR, ZAHEED
To: GOOGLE LLC
Reel/Frame 064091/0018 →
Continuity (5)
Continuation 17991300 · Nov 21, 2022
Continuation 17031416 · Sep 24, 2020
Continuation 15953266 · Apr 13, 2018
Provisional Application 62538764 · Jul 30, 2017
Related Publication 20230344786A1 · Oct 26, 2023