IP Library › Granted Patent US 12,621,409
Granted Patent B2
US 12,621,409 · App. 18/255,399 · Granted May 5, 2026

Multi-party optimization for audiovisual enhancement

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: GOOGLE LLC
H04N7/147G06F3/011H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,621,409
App. No.
18/255,399
Filed
Jun 1, 2023
Granted
May 5, 2026
Kind
B2
Art Unit
2693
USPC
348/14.08
Abstract

A computer-implemented method for selectively applying audiovisual enhancement functions to an audiovisual communications stream includes transmitting, by a sender computing system, an audiovisual communication stream to a receiver computing system, obtaining, by the sender computing system, one or more receiver perception feedback signals associated with the audiovisual communication stream, the one or more receiver perception feedback signals obtained as output from one or more receiver perception feedback models at the receiver computing system and descriptive of perception of the audiovisual communication stream by a user operating the receiver computing system, and applying, by the sender computing system, one or more audiovisual enhancement functions to the audiovisual communication stream based at least in part on the one or more receiver perception feedback signals.

Claims (69)

1 . A computer-implemented method for selectively applying audiovisual enhancement functions to an audiovisual communications stream, the method comprising:

transmitting, by a sender computing system, an audiovisual communication stream to a receiver computing system;

obtaining, by the sender computing system, one or more receiver perception feedback signals associated with the audiovisual communication stream, the one or more receiver perception feedback signals obtained as output from one or more receiver perception feedback models at the receiver computing system and descriptive of perception of the audiovisual communication stream by a user operating the receiver computing system, wherein the one or more receiver perception feedback models comprise a gaze detection model; and

applying, by the sender computing system, one or more audiovisual enhancement functions to the audiovisual communication stream based at least in part on the one or more receiver perception feedback signals, wherein the one or more audiovisual enhancement functions comprise an audio fidelity adjustment function, and wherein the method comprises:

responsive to the one or more receiver perception feedback signals indicating that the user's gaze is on a particular spatial region of a visual component of the transmitted audiovisual communication stream, applying, by the sender computing system, an audio fidelity adjustment to a part of an audio component of the audiovisual communication stream corresponding to the particular spatial region.

2 . The computer-implemented method of claim 1 , wherein the one or more receiver perception feedback models comprise an ambient noise level recognition model, and wherein the method comprises:

responsive to the one or more receiver perception feedback signals indicating that the ambient noise level of an environment of the receiver computing system has increased, applying, by the sender computing system, the one or more audiovisual enhancement functions to increase a fidelity of an audio component of the transmitted audiovisual communication stream.

3 . The computer-implemented method of claim 1 , wherein the one or more audiovisual enhancement functions comprise a focus filter configured to provide an improved fidelity at a focus region, and wherein the method comprises:

applying, by the sender computing system, the focus filter to the audiovisual stream to provide an improved fidelity at spatial region of a visual component of the transmitted audiovisual communication stream that is indicated by the receiver perception feedback signals to be a focus of the gaze of the user operating the receiver computing system.

4 . The computer-implemented method of claim 1 wherein the one or more receiver perception feedback models comprise a user reaction recognition model and the one or more audiovisual enhancement functions comprise a video fidelity adjustment and/or an audio fidelity adjustment, and wherein the method comprises:

responsive to the one or more receiver perception feedback signals indicating that the user may be having difficulty comprehending the transmitted audiovisual communication stream, applying, by the sender computer system, the one or more audiovisual enhancement functions to increase a fidelity of an audio component of the transmitted audiovisual communication stream and/or increase a fidelity of at least part of a visual component of the transmitted audiovisual communication stream.

5 . The computer-implemented method of claim 4 , wherein the method comprises:

responsive to the one or more receiver perception feedback signals indicating that the user may be having difficulty comprehending the visual component of the transmitted audiovisual communication stream, applying by the sender computer system the one or more audiovisual enhancement functions to increase a fidelity of a region of a visual component of the transmitted audiovisual communication stream that is indicated to be a focus of the gaze of the user operating the receiver computing system.

6 . The computer-implemented method of claim 1 , wherein the method comprises:

responsive to the one or more receiver perception feedback signals indicating that the user's gaze is on the particular spatial region of the visual component of the transmitted audiovisual communication stream, applying by the sender computer system the one or more audiovisual enhancement functions to increase a fidelity of a part of the audio component of the transmitted audiovisual communication stream that derives from an entity that is depicted in the particular spatial region of the visual component of the transmitted audiovisual communication stream.

7 . The computer-implemented method of claim 1 , wherein the one or more receiver perception feedback models comprise an automatic speech recognition model and the one or more audiovisual enhancement functions comprise an audio fidelity adjustment, and wherein the method comprises:

responsive to the one or more receiver perception feedback signals indicating a decrease in a confidence associated with recognition, using the automatic speech recognition model, of speech that is present in an audio component of the audiovisual communication stream received at the receiver computing system, applying, by the sender computer system, the one or more audiovisual enhancement functions to increase a fidelity of the audio component of the transmitted audiovisual communication stream.

8 . The computer-implemented method of claim 1 , the one or more audiovisual enhancement functions comprise a video fidelity adjustment and/or an audio fidelity adjustment, and wherein the method comprises:

responsive to the one or more receiver perception feedback signals indicating that the gaze of the user operating the receiver computing system is focused outside a display region on which the visual component of the audiovisual communication stream is presented, applying, by the sender computer system, the one or more audiovisual enhancement functions to decrease a fidelity of the visual component of the transmitted audiovisual communication stream and/or to increase a fidelity of an audio component of the transmitted audiovisual communication stream.

9 . The computer-implemented method of claim 1 , wherein the one or more receiver perception feedback signals are continuously received from the receiver computing system to provide real-time feedback.

10 . The computer-implemented method of claim 1 , wherein the one or more audiovisual enhancement functions comprise one or more of:

a choice of compression scheme type;

a video fidelity adjustment;

an audio fidelity adjustment;

a focus filter configured to provide an improved fidelity at a focus region; and

one or more audiovisual enhancement functions comprise an ambient noise filter.

11 . The computer-implemented method of claim 1 , wherein the one or more receiver perception feedback models comprise one or more of:

an ambient noise level recognition model;

a user reaction recognition model; and

an automatic speech recognition model.

12 . A computing system configured for selectively applying audiovisual enhancement functions to an audiovisual communications stream, the computing system comprising:

a sender computing system comprising one or more processors and one or more memory devices storing computer readable instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:

transmitting an audiovisual communication stream to a receiver computing system;

receiving, from the receiver computing system, one or more receiver perception feedback signals;

applying one or more audiovisual enhancement functions to the audiovisual communication stream based at least in part on the one or more receiver perception feedback signals, wherein the one or more receiver perception feedback signals are generated by a gaze detection model and the one or more audiovisual enhancement functions comprise a video fidelity adjustment and/or an audio fidelity adjustment; and

responsive to the one or more receiver perception feedback signals indicating that the gaze of a user operating the receiver computing system is focused outside a display region on which a visual component of the audiovisual communication stream is presented, applying, by the sender computer system, the one or more audiovisual enhancement functions to decrease a fidelity of the visual component of the audiovisual communication stream and/or to increase a fidelity of an audio component of the audiovisual communication stream.

13 . The computing system of claim 12 , wherein the one or more receiver perception feedback signals are continuously received from the receiver computing system to provide real-time feedback.

14 . The computing system of claim 12 , wherein the sender computing system comprises a sender encoder model, the sender encoder model configured to encode the audiovisual stream prior to transmission to the receiver computing system.

15 . The computing system of claim 12 , wherein the one or more audiovisual enhancement functions comprise one or more of:

a choice of compression scheme type;

a focus filter configured to provide an improved fidelity at a focus region; and

one or more audiovisual enhancement functions comprise an ambient noise filter.

16 . The computing system of claim 12 , wherein the one or more receiver perception feedback signals are generated by one or more of:

an ambient noise level recognition model;

a user reaction recognition model; and

an automatic speech recognition model.

17 . A computing system configured for selectively applying audiovisual enhancement functions to an audiovisual communications stream, the computing system comprising:

a receiver computing system comprising one or more processors and one or more memory devices storing computer readable instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising

receiving, from a sender computing system, an audiovisual communication stream;

obtaining one or more receiver perception feedback signals associated with the audiovisual communication stream, the one or more receiver perception feedback signals obtained as output from one or more receiver perception feedback models comprising a gaze detection model; and

transmitting the one or more receiver perception feedback signals to the sender computing system, the one or more receiver perception feedback signals indicating that a gaze of a user is on a particular spatial region of a visual component of the audiovisual communication stream; and

receiving, from the sender computing system, a modified audiovisual communication stream, the modified audiovisual communication stream including an audio fidelity adjustment to a part of an audio component of the audiovisual communication stream corresponding to the particular spatial region.

18 . The computing system of claim 17 , wherein the receiver computing system comprises a receiver decoder model, the receiver decoder model configured to decode the audiovisual stream from the sender computing system.

19 . The computing system of claim 17 , wherein the one or more audiovisual enhancement functions comprise one or more of:

a choice of compression scheme type;

a video fidelity adjustment;

an audio fidelity adjustment;

a focus filter configured to provide an improved fidelity at a focus region; and

one or more audiovisual enhancement functions comprise an ambient noise filter.

20 . The computing system of claim 17 , wherein the one or more receiver perception feedback models comprise one or any combination of:

a gaze detection model;

an ambient noise level recognition model;

a user reaction recognition model; and

an automatic speech recognition model.

21 . A computer-implemented method for selectively applying audiovisual enhancement functions in a multiparty communication, the method comprising:

receiving, by a receiver computing system, at least one audiovisual communication stream comprising a plurality of audiovisual communication channels from a plurality of sender computing systems;

obtaining, by the receiver computing system, a plurality of receiver perception feedback signals respectively associated with each of the plurality of audiovisual communication channels, the one or more receiver perception feedback signals obtained as output from one or more receiver perception feedback models comprising a gaze detection model at the receiver computing system and descriptive of perception of a respective audiovisual communication channel by a user operating the receiver computing system; and

transmitting, by the receiver computing system, the plurality of receiver perception feedback signals to a respective sender computing system of the plurality of sender computing systems, the one or more receiver perception feedback signals indicating that the gaze of the user operating the receiver computing system is focused outside a display region on which a visual component of the audiovisual communication stream is presented; and

receiving, from the sender computing system, at least one modified audiovisual communication stream, the modified audiovisual communication stream including a decreased fidelity of the visual component of the audiovisual communication stream and/or an increased fidelity of an audio component of the audiovisual communication stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2023
From: SHARIFI, MATTHEW
To: GOOGLE LLC
Reel/Frame 063821/0043 →
Continuity (1)
Related Publication 20240040080A1 · Feb 1, 2024
References Cited (15)
US 10382722B1 · Peters et al. · 2019 [cited by applicant]
US 20130028443A1 · Pance · 2013 [cited by examiner]
US 20130132521A1 · Fonseca, Jr. · 2013 [cited by examiner]
US 20160014476A1 · Caliendo, Jr. · 2016 [cited by applicant]
US 20160277244A1 · Reichert, Jr. · 2016 [cited by examiner]
US 20170293356A1 · Khaderi et al. · 2017 [cited by applicant]
US 20190320114A1 · Lee · 2019 [cited by examiner]
EP 3174287 · 2017 [cited by applicant]
Ephrat et al., “Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation.”, arXiv:1804.03619v2, Aug. 9, 2018, 11 pages. [cited by applicant]
Gergshorn, “Google is Using AI to Compress Photos, just like on HBO's Silicon Valley.”, Aug. 23, 2016, https://qz.com/763649/google-is-working-on-a-way-for-ai-to-compress-your-photos-just-like-on-hbos-silicon-valley, re… [cited by applicant]
International Preliminary Report on Patentability for PCT/US2021/020632, mailed on Sep. 14, 2023, 12 pages. [cited by applicant]
NVIDIA.com, “NVIDIA Maxine.”, 2023, https://developer.nvidia.com/maxine, retrieved on Sep. 29, 2023, 5 pages. [cited by applicant]
QZ.com, “Google is Using AI to Compress Photos, Just Like on HBO's Silicon Valley.”, Aug. 23, 2016, https://qz.com/763649/google-is-working-on-a-way-for-ai-to-compress-your-photos-just-like-on-hbos-silicon-valley, retri… [cited by applicant]
Tagliasacchi et al., “SEANet: A Multi-modal Speech Enhancement Network.”, arXiv:2009.02095v2, Oct. 1, 2020, 5 pages. [cited by applicant]
International Search Report for Application No. PCT/US2021/020632, mailed on Jan. 24, 2022, 5 pages. [cited by applicant]