IP Library › Granted Patent US 12,537,747
Granted Patent B2
US 12,537,747 · App. 18/290,201 · Granted Jan 27, 2026

Method for model inference

Inventors: Qin Mu (Beijing, CN); Wei Hong (Beijing, CN); Zhongyuan Zhao (Beijing, CN); Kexin Xiong (Beijing, CN)
Assignees: BEIJING XIAOMI MOBILE SOFTWARE CO., LTD.; BEIJING UNIVERSITY OF POSTS AND TELECOMMUNICATIONS
H04L41/16H04L41/044H04L41/34H04L47/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,537,747
App. No.
18/290,201
Granted
Jan 27, 2026
Kind
B2
Abstract

A method for model inference, applicable for an operation administration and maintenance (OAM) entity, includes determining a first model corresponding to model subscription request information in response to receiving the model subscription request information sent by a control radio access network (RAN) device; and obtaining a first number of model segmentation blocks by segmenting the first model, and distributing the first number of model segmentation blocks to a first number of control RAN devices.

Claims (65)

1 . A method for model inference, applicable for an operation administration and maintenance (OAM) entity, comprising:

determining a first model corresponding to model subscription request information in response to receiving the model subscription request information sent by a control radio access network (RAN) device; and

obtaining a first number of model segmentation blocks by segmenting the first model, and distributing the first number of model segmentation blocks to a first number of control RAN devices.

2 . The method according to claim 1 , wherein each of the first number of model segmentation blocks correspondingly has allocation information;

wherein the allocation information comprises an inference sequence of the first number of model segmentation blocks and a control RAN device corresponding to each of the first number of model segmentation blocks.

3 . The method according to claim 1 , wherein the first number of control RAN devices comprises a first control RAN device, and the first control RAN device is a control RAN device accessed by a terminal;

wherein distributing the first number of model segmentation blocks to the first number of control RAN devices comprises:

determining a plurality of auxiliary control RAN devices from control RAN devices adjacent to the first control RAN device;

determining a second number of control RAN devices from the plurality of auxiliary control RAN devices based on a computing power occupation state and a load of each of the plurality of auxiliary control RAN devices; wherein the second number of control RAN devices are control RAN devices within the first number of control RAN devices except for the first control RAN device; and

sending a first model segmentation block to the first control RAN device and distributing a remaining number of model segmentation blocks to the second number of control RAN devices, based on the inference sequence of the first number of model segmentation blocks.

4 . The method according to claim 3 , further comprising:

receiving model performance update data sent by the first control RAN device; and

updating the first model based on the model performance update data, determining an updated model parameter of the first model, and sending the updated model parameter of the first model to the first control RAN device.

5 . The method according to claim 3 , further comprising:

updating a distributed RAN device accessed by the terminal in response to receiving a first model analysis subscription update request; wherein the first model analysis subscription update request indicates the terminal to make a switch on the distributed RAN device without making a switch on the control RAN device;

or

updating a distributed RAN device accessed by the terminal and re-segmenting the first model, in response to receiving a second model analysis subscription update request; wherein the second model analysis subscription update request indicates the terminal to make a switch on both the distributed RAN device and the control RAN device.

6 . A device for model inference, comprising:

a processor; and

a memory for storing instructions executable by the processor;

wherein the processor is configured to perform the method for model inference of claim 1 .

7 . A method for model inference, applicable for a control radio access network (RAN) device, comprising:

in response to receiving a model analysis subscription request sent by a distributed RAN device, obtaining model subscription request information by processing the model analysis subscription request, and sending the model subscription request information to an operation administration and maintenance (OAM) entity; and

receiving model segmentation blocks sent by the OAM entity; wherein the model segmentation blocks are model segmentation blocks determined by segmenting a first model, and the first model is determined by the OAM entity based on the model subscription request information.

8 . The method according to claim 7 , further comprising:

after sending the model subscription information to the OAM entity, sending a model inference data request to the distributed RAN device, wherein the model inference data request is configured to acquire model inference data; and

obtaining inference intermediate information of the model segmentation blocks by inferring the model segmentation blocks based on the model inference data.

9 . The method according to claim 8 , wherein the model segmentation blocks correspondingly have allocation information; and the allocation information comprises an inference sequence of a first number of model segmentation blocks and a control RAN device corresponding to each of the first number of model segmentation blocks; and

the method further comprises:

in response to the control RAN device being not a last control RAN device, sending the inference intermediate information to a next control RAN device based on the inference sequence; and

in response to the control RAN device being the last control RAN device, determining a first inference result corresponding to the first model after the model inference is completed, and sending the first inference result to a first control RAN device, wherein the first control RAN device is a control RAN device accessed by a terminal.

10 . The method according to claim 9 , further comprising:

receiving the first inference result in response to the control RAN device being the first control RAN device; and

sending the first inference result to a first distributed RAN device, wherein the first distributed RAN device is a distributed RAN device accessed by the terminal.

11 . The method according to claim 10 , wherein after sending the first inference result to the first distributed RAN device, the method further comprises:

receiving performance data sent by the first distributed RAN device, wherein the performance data is real performance data after the terminal adjusts an execution strategy based on the first model; and

obtaining model performance update data by processing the performance data, and sending the model performance update data to the OAM entity.

12 . The method according to claim 7 , wherein sending the model subscription request information to the OAM entity comprises:

sending the model subscription request information to the OAM entity in response to the control RAN device being a first control RAN device;

wherein the first control RAN device is a control RAN device corresponding to a first distributed RAN device accessed by a terminal.

13 . The method according to claim 12 , further comprising:

in response to the control RAN device being the first control RAN device, determining a second distributed RAN device that resends the model analysis subscription request if the model analysis subscription request is received again, wherein the second distributed RAN device is a distributed RAN device re-accessed after the terminal makes a switch on the distributed RAN device; and

sending a first inference result to the second distributed RAN device, and sending a model subscription update request to the OAM entity.

14 . The method according to claim 13 , further comprising:

in response to the control RAN device being a second control RAN device, if the model analysis subscription request is received again, determining the second distributed RAN device that resends the model analysis subscription request and a second control RAN device that receives the model analysis subscription request again, wherein the second control RAN device is a control RAN device corresponding to the second distributed RAN device, and the second distributed RAN device is a distributed RAN device re-accessed after the terminal makes a switch on the distributed RAN device; and

sending the first inference result to the second control RAN device, and sending the model subscription update request to the OAM entity.

15 . A device for model inference, comprising:

a processor; and

a memory for storing instructions executable by the processor;

wherein the processor is configured to perform the method for model inference of claim 7 .

16 . A method for model inference, applicable for a distributed radio access network (RAN) device, comprising:

sending a model analysis subscription request to a control RAN device in response to receiving the model analysis subscription request sent by a terminal;

wherein the model analysis subscription request is configured to acquire a first model from an operation administration and maintenance (OAM) entity; and the first model comprises a first number of model segmentation blocks.

17 . The method according to claim 16 , further comprising:

receiving a model inference data request sent by the control RAN device, wherein the model inference data request is configured to acquire model inference data; and

acquiring the model inference data from the terminal, and sending the model inference data to the control RAN device.

18 . The method according to claim 16 , further comprising:

receiving a first inference result sent by a first control RAN device, in response to the distributed RAN device being a first distributed RAN device; and

sending the first inference result to the terminal.

19 . The method according to claim 18 , wherein after sending the first inference result to the terminal, the method further comprises:

receiving performance data sent by the terminal in response to the distributed RAN device being the first distributed RAN device; wherein the performance data is real performance data after the terminal adjusts an execution strategy based on the first model; and

sending the performance data to the first control RAN device.

20 . The method according to claim 16 , further comprising:

in response to the distributed RAN device being a second distributed RAN device, if the model analysis subscription request resent by the terminal is received, determining to send the model analysis subscription request to a control RAN device corresponding to the second distributed RAN device;

wherein the second distributed RAN device is a distributed RAN device re-accessed after the terminal makes a switch on the distributed RAN device.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2023
From: MU, QIN
To: BEIJING XIAOMI MOBILE SOFTWARE CO., LTD.
Reel/Frame 065525/0537 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2023
From: HONG, WEI
To: BEIJING XIAOMI MOBILE SOFTWARE CO., LTD.
Reel/Frame 065525/0604 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2023
From: ZHAO, ZHONGYUAN; XIONG, KEXIN
To: BEIJING UNIVERSITY OF POSTS AND TELECOMMUNICATIONS
Reel/Frame 065525/0694 →
Continuity (1)
Related Publication 20240323099A1 · Sep 26, 2024
References Cited (12)
US 20240048988A1 · Pravinchandra Bhatt · 2024 [cited by examiner]
US 20240064066A1 · Chong · 2024 [cited by examiner]
US 20240112087A1 · Balasubramaniam · 2024 [cited by examiner]
US 20240265307A1 · Mu · 2024 [cited by examiner]
US 20240276247A1 · Mu · 2024 [cited by examiner]
PCT/CN2021/092900, International Search Report dated Dec. 27, 2021, 3 pages. [cited by applicant]
CMCC “Solutions for AI-based load balancing” 3GPP TSG-RAN WG3 #112e, R3-212505, May 2021, 7 pages. [cited by applicant]
Lenovo et al. “Discussion on standard impact to support AI functionality” 3GPP TSG-RAN WG3 #112e R3-212180, May 2021, 3 pages. [cited by applicant]
Deutsche Telekom “High-level principles and definitions for the AI/ML-based functional framework for RAN intelligence” 3GPP TSG-RAN3 Meeting #112-e, R3-211632, May 2021, 8 pages. [cited by applicant]
Fraunhofer HHI “Use case and requirements for orchestration of AI/ML based closed loops to enable autonomous networks” International Telecommunication Union, Apr. 2021, 10 pages. [cited by applicant]
3GPP TR 22.874 V0.1.0 Technical Specification Group Services and System Aspects; Study on traffic characteristics and performance requirements for AI/ML model transfer in 5GS (Release 18) Sep. 2020, 55 pages. [cited by applicant]
European Patent Application No. 21941214.5, Search and Opinion dated Jan. 27, 2025, 16 pages. [cited by applicant]