IP Library Granted Patent US 12,348,469
Granted Patent B2
US 12,348,469 · App. 18/670,389 · Granted Jul 1, 2025

Assistance during audio and video calls

Inventors: Fredrik Bergenlid (Mountain View, CA); Vladyslav Lysychkin (Zurich, CH); Denis Burakov (Zurich, CH); Behshad Behzadi (Mountain View, CA); Andrea Terwisscha Van Scheltinga (Zurich, CH); Quentin Lascombes De Laroussilhe (Zurich, CH); Mikhail Golikov (Mountain View, CA); Koa Metter (Mountain View, CA); Ibrahim Badr (Mountain View, CA); Zaheed Sabur (Zurich, CH)
Assignee: Google LLC
H04L51/10G06F16/44G10L15/22H04N7/15H04N21/4394H04N21/4788G10L15/005G10L15/16G10L2015/223G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,348,469
App. No.
18/670,389
Granted
Jul 1, 2025
Kind
B2
Abstract

Implementations relate to providing information items for display during a communication session. In some implementations, a computer-implemented method includes receiving, during a communication session between a first computing device and a second computing device, first media content from the communication session. The method further includes determining a first information item for display in the communication session based at least in part on the first media content. The method further includes sending a first command to at least one of the first computing device and the second computing device to display the first information item.

Claims (52)

1. A computer-implemented method comprising:

receiving session video content during a video communication session between a first computing device and a second computing device;

detecting, in the session video content, a gesture performed by a user associated with the first computing device;

determining, with a trained machine-learning model and based on the gesture, that the user invoked a request for assistance, wherein the request comprises a request for media and wherein the trained machine-learning model outputs a confidence score associated with the determination;

in response to the confidence score meeting a threshold, outputting, by the trained machine-learning model, the media; and

sending a first command to at least one of the first computing device or the second computing device to display the media.

2. The method of claim 1 , wherein the first command causes the at least one of the first computing device or the second computing device to display the media in a user interface as an illustration that is overlaid atop the session video content.

3. The method of claim 2 , further comprising:

determining, with the trained machine-learning model, an emotion associated with the user during the video communication session;

wherein the illustration is selected based on the gesture and the emotion associated with the user during the video communication session.

4. The method of claim 2 , wherein the illustration is selected from a group of a heart, a birthday-related graphic, a trophy, and combinations thereof.

5. The method of claim 1 , further comprising:

performing speech-to-text translation of audio associated with the user using the video communication session to obtain text;

wherein determining that the user associated with the first computing device invoked the request for media is further based on the text.

6. The method of claim 1 , wherein the trained machine-learning model is personalized to the user by:

modifying one or more parameters of a pre-trained machine-learning model based on the user to obtain the trained machine-learning model.

7. The method of claim 1 , wherein determining that the user associated with the first computing device invoked the request for media is further based on at least one detection selected from a group of face detection, motion detection, and combinations thereof.

8. A computing device comprising:

a processor; and

a memory coupled to the processor, with instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising:

receiving session video content during a video communication session between a first computing device and a second computing device;

detecting, in the session video content, a gesture performed by a user associated with the first computing device;

determining, with a trained machine-learning model and based on the gesture, that the user invoked a request for assistance, wherein the request comprises a request for media and wherein the trained machine-learning model outputs a confidence score associated with the determination;

in response to the confidence score meeting a threshold, outputting, by the trained machine-learning model, the media; and

sending a first command to at least one of the first computing device or the second computing device to display the media.

9. The computing device of claim 8 , wherein the first command causes the at least one of the first computing device or the second computing device to display the requested media in a user interface as an illustration that is overlaid atop the session video content.

10. The computing device of claim 9 , wherein the operations further include:

determining, with the trained machine-learning model, an emotion associated with the user during the video communication session;

wherein the illustration is selected based on the gesture and the emotion associated with the user during the video communication session.

11. The computing device of claim 9 , wherein the illustration is selected from a group of a heart, a birthday-related graphic, a trophy, and combinations thereof.

12. The computing device of claim 8 , wherein the operations further include:

performing speech-to-text translation of audio associated with the user using the video communication session to obtain text;

wherein determining that the user associated with the first computing device invoked the request for media is further based on the text.

13. The computing device of claim 8 , wherein the trained machine-learning model is personalized to the user by:

modifying one or more parameters of a pre-trained machine-learning model based on the user to obtain the trained machine-learning model.

14. The computing device of claim 8 , wherein determining that the user associated with the first computing device invoked the request for media is further based on at least one detection selected from a group of face detection, motion detection, and combinations thereof.

15. A non-transitory computer-readable medium with instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:

receiving session video content during a video communication session between a first computing device and a second computing device;

detecting, in the session video content, a gesture performed by a user associated with the first computing device;

determining, with a trained machine-learning model and based on the gesture, that the user invoked a request for assistance, wherein the request comprises a request for media and wherein the trained machine-learning model outputs a confidence score associated with the determination;

in response to the confidence score meeting a threshold, outputting, by the trained machine-learning model, the media; and

sending a first command to at least one of the first computing device or the second computing device to display the media.

16. The non-transitory computer-readable medium of claim 15 , wherein the first command causes the at least one of the first computing device or the second computing device to display the requested media in a user interface as an illustration that is overlaid atop the session video content.

17. The non-transitory computer-readable medium of claim 16 , wherein the operations further include:

determining, with the trained machine-learning model, an emotion associated with the user during the video communication session;

wherein the illustration is selected based on the gesture and the emotion associated with the user during the video communication session.

18. The non-transitory computer-readable medium of claim 16 , wherein the illustration is selected from a group of a heart, a birthday-related graphic, a trophy, and combinations thereof.

19. The non-transitory computer-readable medium of claim 15 , wherein the operations further include:

performing speech-to-text translation of audio associated with the user using the video communication session to obtain text;

wherein determining that the user associated with the first computing device invoked the request for media is further based on the text.

20. The non-transitory computer-readable medium of claim 15 , wherein the trained machine-learning model is personalized to the user by:

modifying one or more parameters of a pre-trained machine-learning model based on the user to obtain the trained machine-learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2024
From: BERGENLID, FREDRIK; LYSYCHKIN, VLADYSLAV; BURAKOV, DENIS; BEHZADI, BEHSHAD; VAN SCHELTINGA, ANDREA TERWISSCHA; DE LAROUSSILHE, QUENTIN LASCOMBES; GOLIKOV, MIKHAIL; METTER, KOA; BADR, IBRAHIM; SABUR, ZAHEED
To: GOOGLE LLC
Reel/Frame 067508/0847 →
Continuity (6)
Continuation 18215221 · Jun 28, 2023
Continuation 17991300 · Nov 21, 2022
Continuation 17031416 · Sep 24, 2020
Continuation 15953266 · Apr 13, 2018
Provisional Application 62538764 · Jul 30, 2017
Related Publication 20240314094A1 · Sep 19, 2024
References Cited (23)
US 8996635B1 · Byttow et al. · 2015 [cited by applicant]
US 9462112B2 · Woolsey et al. · 2016 [cited by applicant]
US 9576270B1 · Afshar et al. · 2017 [cited by applicant]
US 10075676B2 · Segal · 2018 [cited by applicant]
US 20050078613A1 · Covell · 2005 [cited by examiner]
US 20060172749A1 · Sweeney · 2006 [cited by applicant]
US 20080092158A1 · Bhatnagar et al. · 2008 [cited by applicant]
US 20080297588A1 · Kurtz et al. · 2008 [cited by applicant]
US 20120192087A1 · Lemmey · 2012 [cited by examiner]
US 20130321390A1 · Latta · 2013 [cited by examiner]
US 20140164532A1 · Lynch et al. · 2014 [cited by applicant]
US 20140267543A1 · Kerger et al. · 2014 [cited by applicant]
US 20150088514A1 · Typrin · 2015 [cited by applicant]
US 20180129385A1 · Blumenfeld et al. · 2018 [cited by applicant]
US 20200286130A1 · Betan et al. · 2020 [cited by applicant]
US 20210385299A1 · Liao · 2021 [cited by examiner]
International Bureau of WIPO, International Preliminary Report on Patentability for International Patent Application No. PCT/US2018/027639, Feb. 4, 2020, 6 pages. [cited by applicant]
USPTO, Notice of Allowance for U.S. Appl. No. 18/215,221, Feb. 22, 2024, 10 pages. [cited by applicant]
USPTO, Notice of Allowance for U.S. Appl. No. 17/031,416, Jul. 25, 2022, 10 pages. [cited by applicant]
USPTO, Notice of Allowance for U.S. Appl. No. 15/953,266, Jun. 1, 2020, 5 pages. [cited by applicant]
USPTO, Notice of Allowance for U.S. Appl. No. 17/991,300, Mar. 29, 2023, 10 pages. [cited by applicant]
USPTO, Non-final Office Action for U.S. Appl. No. 15/953,266, Nov. 21, 2019, 11 pages. [cited by applicant]
WIPO, “Written Opinion and International Search report in International Application No. PCT/US2018/027639”, Jun. 4, 2018, 8 pages. [cited by applicant]