IP Library › Granted Patent US 12,417,406
Granted Patent B2
US 12,417,406 · App. 17/507,188 · Granted Sep 16, 2025

Virtualizing external memory as local to a machine learning accelerator

Inventors: Lawrence J. Madar, III (San Francisco, CA); Temitayo Fadelu (San Francisco, CA); Harshit Khaitan (San Jose, CA); Ravi Narayanaswami (San Jose, CA)
Assignee: Google LLC
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,406
App. No.
17/507,188
Granted
Sep 16, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for virtualizing external memory as local to a machine learning accelerator. One ambient computing system comprises: an ambient machine learning engine; a low-power CPU; and an SRAM that is shared among at least the ambient machine learning engine and the low-power CPU; wherein the ambient machine learning engine comprises virtual address logic to translate from virtual addresses generated by the ambient machine learning engine to physical addresses within the SRAM.

Claims (37)

1. A device comprising:

a main memory shared by multiple client devices, wherein the multiple client devices include an ambient computing device, wherein the ambient computing device comprises: multiple ambient processing devices including an ambient machine learning (ML) engine; and a shared local memory that is shared by the multiple ambient processing devices,

wherein the ambient computing device includes virtual address logic configured to translate virtual addresses used by the ambient ML engine into physical addresses of the shared local memory, the ambient computing device being configured to:

load first instructions for a first ambient processing device of the multiple ambient processing devices into the shared local memory;

execute the first instructions using the first ambient processing device;

load second instructions for the ambient ML engine into the shared local memory, the second instructions overwriting at least part of the first instructions;

execute the second instructions using the ambient ML engine, to generate a request to read or write data to a virtual address;

translate the virtual address to a physical address of the shared local memory using the virtual address logic; and

read or write the data to the physical address of the shared local memory.

2. The device of claim 1 , wherein the ambient computing device is configured to process sensor signals before other client devices sharing the main memory are activated from a low-power state.

3. The device of claim 1 , wherein the multiple client devices include a main ML engine that generates physical addresses in the main memory.

4. The device of claim 1 , wherein upon receiving an interrupt, the device is configured to stream model parameters from the main memory into the shared local memory.

5. The device of claim 4 , wherein streaming the model parameters into the shared local memory overwrites space in the shared local memory used by one or more other ambient processing devices.

6. The device of claim 4 , wherein streaming the model parameters into the shared local memory comprises streaming the model parameters into available space in the shared local memory.

7. The device of claim 4 , wherein the interrupt represents receipt of one or more sensor signals to be processed.

8. The device of claim 4 , wherein the ambient ML engine is configured to perform an inference pass of a machine learning model by using the virtual address logic to access the model parameters that were streamed into the shared local memory.

9. The device of claim 1 , wherein the shared local memory comprises multiple memory banks configured to be individually powered down when entering a low-power state.

10. The device of claim 1 , wherein the multiple ambient processing devices comprise at least one of a direct memory access controller, one or more other ML engines, or one or more processors.

11. The device of claim 1 , wherein the ambient ML engine comprises a single ML compute tile.

12. A system comprising:

multiple client devices including an ambient computing device; and

a main memory shared by the multiple client devices, wherein the ambient computing device comprises multiple ambient processing devices and a shared local memory that is shared by the multiple ambient processing devices including an ambient machine learning (ML) engine,

wherein the ambient computing device includes virtual address logic configured to translate virtual addresses used by the ambient ML engine into physical addresses of the shared local memory, the ambient computing device being configured to:

load first instructions for a first ambient processing device of the multiple ambient processing devices into the shared local memory;

execute the first instructions using the first ambient processing device;

load second instructions for the ambient ML engine into the shared local memory, the second instructions overwriting at least part of the first instructions;

execute the second instructions using the ambient ML engine, to generate a request to read or write data to a virtual address;

translate the virtual address to a physical address of the shared local memory using the virtual address logic; and

read or write the data to the physical address of the shared local memory.

13. The system of claim 12 , wherein the ambient computing device is configured to process sensor signals before other client devices sharing the main memory are activated from a low-power state.

14. The system of claim 12 , wherein the multiple client devices include a main ML engine that generates physical addresses in the main memory.

15. The system of claim 12 , wherein upon receiving an interrupt, the device is configured to stream model parameters from the main memory into the shared local memory.

16. The system of claim 15 , wherein streaming the model parameters into the shared local memory overwrites space in the shared local memory used by one or more other ambient processing devices.

17. The system of claim 15 , wherein streaming the model parameters into the shared local memory comprises streaming the model parameters into available space in the shared local memory.

18. The system of claim 15 , wherein the interrupt represents receipt of one or more sensor signals to be processed.

19. The system of claim 15 , wherein the ambient ML engine is configured to perform an inference pass of a machine learning model by using the virtual address logic to access the model parameters that were streamed into the shared local memory.

20. The system of claim 12 , wherein the shared local memory comprises multiple memory banks configured to be individually powered down when entering a low-power state.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE 1ST ASSIGNOR NAME PREVIOUSLY RECORDED AT REEL: 57868 FRAME: 162. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jul 10, 2025
From: MADAR, LAWRENCE J., III; FADELU, TEMITAYO; KHAITAN, HARSHIT; NARAYANASWAMI, RAVI
To: GOOGLE LLC
Reel/Frame 071912/0943 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2021
From: MADAR, LAWRENCE J.; FADELU, TEMITAYO; KHAITAN, HARSHIT; NARAYANASWAMI, RAVI
To: GOOGLE LLC
Reel/Frame 057868/0162 →
Continuity (2)
Continuation 16397481 · Apr 29, 2019
Related Publication 20220044153A1 · Feb 10, 2022
References Cited (48)
US 9710265B1 · Temam · 2017 [cited by applicant]
US 10746792B1 · Diamant et al. · 2020 [cited by applicant]
US 11176493B2 · Madar, III et al. · 2021 [cited by applicant]
US 20140047251A1 · Kottilingal et al. · 2014 [cited by applicant]
US 20140176572A1 · Vembu et al. · 2014 [cited by applicant]
US 20170206464A1 · Clayton et al. · 2017 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180322390A1 · Das et al. · 2018 [cited by applicant]
US 20190243756A1 · Ray et al. · 2019 [cited by applicant]
US 20200183482A1 · Sebot · 2020 [cited by examiner]
US 20200211151A1 · Vaidyanathan · 2020 [cited by examiner]
US 20200342350A1 · Madar, III et al. · 2020 [cited by applicant]
US 20210019631A1 · Das et al. · 2021 [cited by applicant]
CN 104335180 · 2015 [cited by applicant]
CN 106708753 · 2017 [cited by applicant]
CN 106951926 · 2017 [cited by applicant]
CN 109508782 · 2019 [cited by applicant]
EP 1988467 · 2008 [cited by applicant]
EP 3385850A1 · 2018 [cited by applicant]
EP 3396533 · 2018 [cited by applicant]
JP 2008282396 · 2008 [cited by applicant]
JP 2008310700 · 2008 [cited by applicant]
JP 2015195031A · 2015 [cited by applicant]
JP 2016500186A · 2016 [cited by applicant]
JP 2017084370A · 2017 [cited by applicant]
JP 2018133016A · 2018 [cited by applicant]
TW 201911039 · 2019 [cited by applicant]
WO WO2015183404 · 2015 [cited by applicant]
WO WO2017218009 · 2017 [cited by applicant]
WO WO2019023046A1 · 2019 [cited by applicant]
WO WO2019032808A1 · 2019 [cited by applicant]
WO WO2019104228 · 2019 [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2021-557065, mailed on Dec. 5, 2023, 7 pages (with machine translation). [cited by applicant]
Office Action in European Appln. No. 19824100.2, mailed on May 28, 2024, 4 pages. [cited by applicant]
Office Action in Japanese Appln. No. 2023-186440, mailed on Jan. 7, 2024, 16 pages (with machine translation). [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2019/063424, dated Nov. 2, 2021, 7 pages. [cited by applicant]
Office Action in Chinese Appln. No. 201980094598.6, mailed on Sep. 6, 2023, 14 pages (with English translation). [cited by applicant]
Office Action in Japanese Appln. No. 2021-557065, mailed on Jul. 11, 2023, 6 pages (with English translation). [cited by applicant]
Hu et al., “Design and implementation of the ARM virtual machine for multithreading” Informationization, No. 11, Jun. 10, 2009, 5 pages (with English abstract). [cited by applicant]
devblogs.nvidia.com [online], “Inside pascal: Nvidia's newest computing platform,” Nvidia Developer, Apr. 2016, retrieved on Nov. 25, 2019, retrieved from URL <https://devblogs.nvidia.com/inside-pascal/>, 19 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2019/063424, dated Apr. 1, 2020, 13 pages. [cited by applicant]
Office Action in Taiwanese Appln. No. 108143271, dated Nov. 18, 2020, 13 pages (with English translation). [cited by applicant]
Office Action in European Appln. No. 19824100.2, dated Mar. 22, 2023, 6 pages. [cited by applicant]
Office Action in Taiwanese Appln. No. 111136960, dated Feb. 3, 2023, 11 pages (with English translation). [cited by applicant]
Wikipedia.org [online], “Ubiquitous Computing” available on or before Apr. 2019, via Internet Archive: Wayback Machine URL <https://web.archive.org/web/20190427161726/https://en.wikipedia.org/wiki/Ubiquitous_computing>,… [cited by applicant]
Office Action in Japanese Appln. No. 2021-557065, dated Jan. 24, 2023, 10 pages (with English translation). [cited by applicant]
Extended European Search Report in European Appln. No. 25154610.7, mailed on Jun. 30, 2025, 11 pages. [cited by applicant]
Wikipedia.org [online], “Virtual memory,” Oct. 20, 2018, retrieved on Jul. 9, 2025, retrieved from URL<https://en.wikipedia.org/w/index.php?title=Virtual_memory&oldid=864864520 >, 9 pages. [cited by applicant]