IP Library Granted Patent US 12694701
Granted Patent B1
US 12694701 · App. 18/205,204 · Granted Jul 28, 2026

Partial neural network activation

Inventors: Jorge Albericio Latorre (Brooklyn, NY); Michael Ranzinger (Park City, UT)
Assignee: NVIDIA Corporation
G06V30/19113
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694701
App. No.
18/205,204
Granted
Jul 28, 2026
Kind
B1
Abstract

Apparatuses, systems, and techniques to perform inferencing with neural networks. In at least one embodiment, portions of one or more neural networks are selected for use in inferencing based, at least in part, on information to be inferenced by the one or more neural networks.

Claims (31)

1 . A processor, comprising:

one or more circuits to:

cause one or more layers of one or more neural networks to generate output values from performing a portion of inferencing based, at least in part, on input data;

cause one or more sub-networks to be selected from among a plurality of sub-networks of the one or more neural networks to complete the inferencing based, at least in part, on the input data or the output values from the one or more layers; and

provide the output values from the one or more layers to the selected one or more sub-networks, wherein the output values are duplicated to each selected sub-network such that the portion of inferencing is common to the plurality of sub-networks to complete the inferencing.

2 . The processor of claim 1 , wherein the one or more neural networks comprise two or more sub-models to be selected based, at least in part, on one or more router layers of the one or more neural networks.

3 . The processor of claim 2 , wherein the one or more router layers select one of the two or more sub-models based, at least in part, on one or more features detected in one or more inputs to the one or more neural networks.

4 . The processor of claim 2 , wherein a sub-model selected from the two or more sub-models is to be loaded into a memory by the processor.

5 . The processor of claim 2 , wherein the one or more neural networks comprise one or more stem layers to perform inferencing common to two or more of the two or more sub-models.

6 . The processor of claim 5 , wherein the one or more router layers select a sub-model of the two or more sub-models based, at least in part, on one or more outputs of the one or more stem layers.

7 . The processor of claim 2 , wherein the two or more sub-models are respectively trained to recognize text in different environments.

8 . A method, comprising:

performing a portion of inferencing by one or more layers of one or more neural networks to generate output values based, at least in part, on input data;

selecting one or more sub-networks from among a plurality of sub-networks of the one or more neural networks to complete the inferencing based, at least in part, on the input data or the output values from the one or more layers; and

providing the output values from the one or more layers to the selected one or more sub-networks, wherein the output values are duplicated to each selected sub-network such that the portion of inferencing is common to the plurality of sub-networks to complete the inferencing.

9 . The method of claim 8 , wherein the one or more neural networks comprise two or more sub-models to be selected based, at least in part, on one or more router layers of the one or more neural networks.

10 . The method of claim 9 , wherein the one or more router layers select one of the two or more sub-models based, at least in part, on one or more features detected in one or more inputs to the one or more neural networks.

11 . The method of claim 9 , further comprising using one or more processors to load a selected sub-model selected from the two or more sub-models to a memory.

12 . The method of claim 9 , wherein the one or more neural networks comprise one or more stem layers to perform inferencing common to two or more of the two or more sub-models.

13 . The method of claim 12 , wherein the one or more router layers select a sub-model of the two or more sub-models based, at least in part, on one or more outputs of the one or more stem layers.

14 . The method of claim 9 , wherein the two or more sub-models are respectively trained to recognize text in different environments.

15 . A system, comprising:

one or more processors to:

cause one or more layers of one or more neural networks to generate output values from performing a portion of inferencing based, at least in part, on input data;

cause one or more sub-networks to be selected from among a plurality of sub-networks of the one or more neural networks to complete the inferencing based, at least in part, on the input data or the output values from the one or more layers; and

provide the output values from the one or more layers to the selected one or more sub-networks, wherein the output values are duplicated to each selected sub-network such that the portion of inferencing is common to the plurality of sub-networks to complete the inferencing.

16 . The system of claim 15 , wherein the one or more neural networks comprise two or more sub-models to be selected based, at least in part, on one or more router layers of the one or more neural networks.

17 . The system of claim 16 , wherein the one or more router layers select one of the two or more sub-models based, at least in part, on one or more features detected in one or more inputs to the one or more neural networks.

18 . The system of claim 16 , wherein the one or more neural networks comprise one or more stem layers to perform inferencing common to two or more of the two or more sub-models.

19 . The system of claim 18 , wherein the one or more router layers select a sub-model of the two or more sub-models based, at least in part, on one or more outputs of the one or more stem layers.

20 . The system of claim 16 , wherein the two or more sub-models are respectively trained to recognize text in different environments.