IP Library Granted Patent US 12,561,108
Granted Patent B2
US 12,561,108 · App. 18/765,096 · Granted Feb 24, 2026

Resolving time-delays using generative models

Inventors: Jakob Nicolaus Foerster (San Francisco, CA); Ioannis Alexandros Assael (London, GB)
Assignee: GDM Holding LLC
G06F3/162G06F3/165G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,108
App. No.
18/765,096
Granted
Feb 24, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating audio output samples predicted to be communicated by a user. One example system includes a first user device having a first user. The first user device initiates a communication session between the first user and a second user of a second user device. The first user device obtains a neural network model of the second user. The neural network model is trained to generate, conditioned on audio input samples received up to a current time step, an audio output sample predicted to be communicated by the second user at a next time step. The user device repeatedly provides received audio input samples as input to the neural network model and plays audio output samples generated by the neural network model in place of received audio input samples communicated by the second user.

Claims (38)

1 . A method comprising:

initiating a communication session during which a first user device transmits data to and receives data from a second user device; and

at each of multiple time steps during the communication session:

generating, by using a neural network model, a predicted data sample that is a prediction of a current data sample to be communicated by the second user device at a current time step,

presenting the predicted data sample generated by using the neural network model for display in place of the current data sample communicated by the second user device at the current time step, and

after presenting the predicted data sample for display in place of the current data sample, providing the current data sample communicated by the second user device as part of a model input comprising past data samples received by the first user device to the neural network model.

2 . The method of claim 1 , wherein the data transmitted to and received from the second user device by the first user device comprises video data.

3 . The method of claim 1 , wherein the current data sample comprises a video frame.

4 . The method of claim 1 , wherein the first user device and the second user device are physically apart from each other.

5 . The method of claim 1 , wherein the communication session is a video conferencing session.

6 . The method of claim 1 , wherein the communication session is a gaming session.

7 . The method of claim 1 , wherein the communication session is a collaborative activity session.

8 . The method of claim 1 , wherein the neural network model is configured to process the model input at a future time step to generate a predicted future data sample that is a prediction of a future data sample to be communicated by the second user at the future time step.

9 . A system comprising:

one or more computers; and

one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

initiating a communication session during which a first user device transmits data to and receives data from a second user device; and

at each of multiple time steps during the communication session:

generating, by using a neural network model, a predicted data sample that is a prediction of a current data sample to be communicated by the second user device at a current time step,

presenting the predicted data sample generated by using the neural network model for display in place of the current data sample communicated by the second user device at the current time step, and

after presenting the predicted data sample for display in place of the current data sample, providing the current data sample communicated by the second user device as part of a model input comprising past data samples received by the first user device to the neural network model.

10 . The system of claim 9 , wherein the data transmitted to and received from the second user device by the first user device comprises video data.

11 . The system of claim 9 , wherein the current data sample comprises a video frame.

12 . The system of claim 9 , wherein the first user device and the second user device are physically apart from each other.

13 . The system of claim 9 , wherein the communication session is a video conferencing session.

14 . The system of claim 9 , wherein the communication session is a gaming session.

15 . The system of claim 9 , wherein the communication session is a collaborative activity session.

16 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

initiating a communication session during which a first user device transmits data to and receives data from a second user device; and

at each of multiple time steps during the communication session:

generating, by using a neural network model, a predicted data sample that is a prediction of a current data sample to be communicated by the second user device at a current time step,

presenting the predicted data sample generated by using the neural network model for display in place of the current data sample communicated by the second user device at the current time step, and

after presenting the predicted data sample for display in place of the current data sample, providing the current data sample communicated by the second user device as part of a model input comprising past data samples received by the first user device to the neural network model.

17 . The non-transitory computer-readable storage media of claim 16 , wherein the data transmitted to and received from the second user device by the first user device comprises video data.

18 . The non-transitory computer-readable storage media of claim 16 , wherein the current data sample comprises a video frame.

19 . The non-transitory computer-readable storage media of claim 16 , wherein the first user device and the second user device are physically apart from each other.

20 . The non-transitory computer-readable storage media of claim 16 , wherein the communication session is a video conferencing session.

21 . The non-transitory computer-readable storage media of claim 16 , wherein the communication session is a collaborative activity session.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2025
From: FOERSTER, JAKOB NICOLAUS; ASSAEL, IOANNIS ALEXANDROS
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 070754/0333 →
Priority Claims (1)
GR 20180100238 · Jun 1, 2018 · national
Continuity (3)
Continuation 17886951 · Aug 12, 2022
Continuation 16428370 · May 31, 2019
Related Publication 20240361974A1 · Oct 31, 2024
References Cited (8)
US 10127918B1 · Kamath Koteshwara · 2018 [cited by examiner]
US 10594757B1 · Shevchenko et al. · 2020 [cited by applicant]
US 20170140753A1 · Jaitly et al. · 2017 [cited by applicant]
US 20180210872A1 · Podnnajersky et al. · 2018 [cited by applicant]
US 20180350395A1 · Simko et al. · 2018 [cited by applicant]
US 20190051272A1 · Lewis · 2019 [cited by examiner]
US 20190052471A1 · Panattoni · 2019 [cited by examiner]
US 20200068434A1 · Pedersen · 2020 [cited by examiner]