IP Library Granted Patent US 12,475,366
Granted Patent B2
US 12,475,366 · App. 17/316,516 · Granted Nov 18, 2025

Updating a neural network model on a computation device

Inventors: Peter Exner (Malmö, SE); Hannes Bergkvist (Rydebäck, SE)
Assignee: Sony Group Corporation
G06N3/08G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,366
App. No.
17/316,516
Granted
Nov 18, 2025
Kind
B2
Abstract

A method for updating a neural network model on a computation device. The method includes estimating a bandwidth for data download from a server device to the computation device; estimating a time point of available computation capacity of the computation device; computing a maximum partition size as a function of the bandwidth and the time point; and causing download of a selected partition of the neural network model from the server device to the computation device, the selected partition being determined based on the maximum partition size. The method enables the selected partition to be executed by the computation device upon downloading and thereby provides for seamless updating of the neural network model while it is being executed on the computation device.

Claims (25)

1 . A method of updating a neural network model on a computation device, said method comprising:

estimating a bandwidth for data download from a server device to the computation device;

estimating a time point of available computation capacity of the computation device;

computing a maximum partition size as a function of the bandwidth and the time point; and

causing download of a selected partition of the neural network model from the server device to the computation device, said selected partition being determined based on the maximum partition size; and

updating and executing the selected partition of the neural network model, downloaded from the server, on the computation device,

wherein updating the selected partition includes replacing a corresponding partition of the neural network model on the computation device with the selected partition downloaded from the server device, and

wherein the method is performed while the neural network model is being executed on the computation device.

2 . The method of claim 1 , wherein the selected partition is determined to have a size substantially equal to or less than the maximum partition size.

3 . The method of claim 1 , wherein execution of the selected partition is initiated at the time point of available computation capacity.

4 . The method of claim 1 , wherein the maximum partition size is computed as a function of a product of the bandwidth and a time interval from a selected time to the time point.

5 . The method of claim 1 , wherein one or more existing partitions of the neural network model have been executed on the computation device at said time point.

6 . The method of claim 5 , wherein the selected partition is determined as a function of the one or more existing partitions.

7 . The method of claim 5 , wherein the selected partition is determined as a function of a dependence within the neural network model on output generated by the one or more existing partitions.

8 . The method of claim 5 , wherein the selected partition is determined so as to operate on the output generated by the one or more existing partitions.

9 . The method of claim 5 , further comprising: evaluating the output generated by the one or more existing partitions to identify one or more partitions to be excluded from execution, wherein the selected partition is determined while excluding the one or more partitions.

10 . The method of claim 5 , wherein said causing the download comprises: transmitting, by the computation device to the server device, size data indicative of the maximum partition size and status data indicative of the one or more existing partitions that have been executed at the time point.

11 . The method of claim 1 , further comprising: determining the selected partition by the computation device, wherein said causing the download comprises: transmitting, by the computation device to the server device, data indicative of the selected partition.

12 . The method of claim 1 , wherein the selected partition is determined as a function of the available computation capacity of the computation device at the time point.

13 . The method of claim 1 , wherein the selected partition is determined among a plurality of predefined partitions of the neural network model and based on a predefined dependence between the predefined partitions.

14 . The method of claim 1 , wherein the selected partition is determined by dynamically partitioning the neural network model on demand.

15 . The method of claim 1 , which is performed by the computation device.

16 . The method of claim 1 , which is repeatedly performed at consecutive current time points to update the neural network model on the computation device, and wherein said causing, at a respective current time point, results in download of partition data of a size substantially equal to the maximum partition size estimated at the respective current time point, said partition data comprising the selected partition, and optionally one or more further selected partitions.

17 . A computation device comprising a communication circuit for communicating with a server device, and logic to control the computation device to perform the method in accordance with claim 1 .

18 . A non-transitory computer-readable medium comprising computer instructions which, when executed by a processing system, cause the processing system to perform the method in accordance with claim 1 .

Assignments (3)
CHANGE OF NAME Recorded May 11, 2021
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 056207/0641 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2021
From: SONY EUROPE BV
To: SONY CORPORATION
Reel/Frame 056214/0167 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2021
From: BERGKVIST, HANNES; EXNER, PETER
To: SONY EUROPE BV
Reel/Frame 056215/0026 →
Priority Claims (1)
SE 2050693-7 · Jun 11, 2020 · national
Continuity (1)
Related Publication 20210390402A1 · Dec 16, 2021
References Cited (27)
US 8793205B1 · Fisher · 2014 [cited by examiner]
US 9613310B2 · Buibas · 2017 [cited by applicant]
US 10387770B2 · Brothers · 2019 [cited by applicant]
US 11755884B2 · Eilert · 2023 [cited by examiner]
US 20040162638A1 · Solomon · 2004 [cited by examiner]
US 20190205744A1 · Mondello · 2019 [cited by applicant]
US 20190205765A1 · Mondello · 2019 [cited by applicant]
US 20190258924A1 · Hamidouche · 2019 [cited by applicant]
US 20190392305A1 · Gu · 2019 [cited by applicant]
US 20210027166A1 · Gorokhov · 2021 [cited by examiner]
US 20230259744A1 · Moradi · 2023 [cited by examiner]
EP 3447645A1 · 2019 [cited by applicant]
WO 2019185981A1 · 2019 [cited by applicant]
WO 2020046859A1 · 2020 [cited by applicant]
Kang, Yiping, et al. “Neurosurgeon: Collaborative intelligence between the cloud and mobile edge.” (Year: 2017). [cited by examiner]
Luo, Liang, et al. “Parameter hub: a rack-scale parameter server for distributed deep neural network training.” (Year: 2018). [cited by examiner]
Shi, Wenqi, et al. “Improving device-edge cooperative inference of deep learning via 2-step pruning.” (Year: 2019). [cited by examiner]
Jeong, Hyuk-Jin, et al. “IONN: Incremental offloading of neural network computations from mobile devices to edge servers.” (Year: 2018). [cited by examiner]
Li, En, et al. “Edge AI: On-demand accelerating deep neural network inference via edge computing.” (Year: 2019). [cited by examiner]
Gao, et al. “Tetris: Scalable and efficient neural network acceleration with 3d memory.” (Year: 2017). [cited by examiner]
Chen, Yu-Hsin, et al. “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks.” (Year: 2016). [cited by examiner]
Hinton et al., “Distilling the knowledge in a neural network.” (Year: 2015). [cited by examiner]
Kang et al. “Neurosurgeon: Collaborative intelligence between the cloud and mobile edge”, (Year: 2017). [cited by examiner]
Shi, et al. “Improving device-edge cooperative inference of deep learning via 2-step pruning.” (Year: 2019). [cited by examiner]
Luo, et al. “Parameter hub: a rack-scale parameter server for distributed deep neural network training.” (Year: 2018). [cited by examiner]
Wang et al., “Convergence of Edge Computing and Deep Learning: A Comprehensive Survey,” IEEE Communication Surveys & Tutorials; 1907.08349v3; dated Jan. 28, 2020, 36 pages. [cited by applicant]
Office Action and Swedish Search Report from corresponding Swedish Application No. 2050693-7, mailed on Apr. 16, 2021, 5 pages. [cited by applicant]