IP Library Granted Patent US 12,417,394
Granted Patent B2
US 12,417,394 · App. 17/204,935 · Granted Sep 16, 2025

System and method for AI model watermarking

Inventors: Laurent Charette (Vancouver, CA); Lingyang Chu (Burnaby, CA); Lanjun Wang (Toronto, CA); Yong Zhang (Richmond, CA)
Assignee: Huawei Cloud Computing Technologies Co., Ltd.
G06N7/01G06N20/00G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,394
App. No.
17/204,935
Granted
Sep 16, 2025
Kind
B2
Abstract

Method and system for watermarking prediction outputs generated by a first AI model to enable detection of a target AI model that has been distilled from the prediction outputs. Includes receiving, at the first AI model, a set of input data samples from a requesting device; storing at least a subset of the input data samples to maintain a record of the input data samples; predicting, using the first AI model, a respective set of prediction outputs that each include a probability value, the AI model using a watermark function to insert a periodic watermark signal in the probability values of the prediction outputs; and outputting, from the first AI model, the prediction outputs including the periodic watermark signal.

Claims (30)

1. A method for watermarking prediction outputs generated by a first AI model to enable detection of a target AI model that has been distilled from the watermarked prediction outputs, comprising:

receiving, at the first AI model, a set of input data samples from a requesting device;

storing at least a subset of the input data samples to maintain a record of the input data samples;

predicting, using the first AI model, a respective set of watermarked prediction outputs that each include a probability value, the first AI model using a watermark embedding function to insert a periodic watermark signal in the probability values to generate the watermarked prediction outputs; and

outputting, from the first AI model, the watermarked prediction outputs including the periodic watermark signal.

2. The method of claim 1 wherein the periodic watermark signal is configured to cause an AI model that is distilled from the respective set of prediction outputs to insert a periodic signal into prediction outputs of the AI model that can be detected as matching the periodic watermark signal.

3. The method of claim 1 , further comprising, prior to the predicting, including the watermark embedding function in a preliminary AI model to generate the first AI model.

4. The method of claim 3 , wherein the preliminary AI model is an untrained model, the method comprising training the preliminary AI model using a loss that is based on the outputs of the preliminary AI model with the included watermark embedding function.

5. The method of claim 1 , comprising defining a key corresponding to the watermark embedding function, the key including a random projection vector, wherein the watermark embedding function inserts the periodic watermark signal based on the random projection vector.

6. The method of claim 5 , further comprising determining if the target AI model has been distilled from the first AI model by:

submitting a query to the target AI model that includes at least some of the stored subset of the input data samples;

receiving prediction outputs from the target AI model corresponding to the input data samples;

determining, based on the key, if a periodic signal that matches the periodic watermark signal can be detected in the prediction outputs from the target AI model.

7. The method of claim 6 wherein the key further includes information that identifies a frequency of the periodic watermark signal and a target prediction output to monitor for the periodic watermark signal.

8. The method of claim 1 , wherein the watermark embedding function is configured to modify softmax outputs of a softmax layer of the first AI model by inserting the periodic watermark signal into the probability values of the prediction outputs.

9. A method of determining if a target AI model has been distilled from a first AI model by:

submitting a query to the target AI model that includes input data samples that were previously provided to the first AI model;

receiving prediction outputs from the target AI model corresponding to the input data samples; and

determining, based on a predetermined key, if a periodic signal that matches a known periodic watermark signal can be detected in the prediction outputs from the target AI model.

10. The method of claim 9 wherein the determining comprises:

determining, based on a Fourier power spectrum of the prediction outputs and a projection vector included in the predetermined key, if a signal power that corresponds to the frequency of the known periodic watermark signal can be detected in the prediction outputs from the target AI model.

11. A computer system comprising one or more processing units and one or more non-transient memories storing computer implementable instructions for execution by the one or more processing devices, wherein execution of the computer implementable instructions configures the computer system to perform a method for watermarking prediction outputs generated by a first AI model to enable detection of a target AI model that has been distilled from the watermarked prediction outputs, comprising:

receiving, at the first AI model, a set of input data samples from a requesting device;

storing at least a subset of the input data samples to maintain a record of the input data samples;

predicting, using the first AI model, a respective set of watermarked prediction outputs that each include a probability value, the first AI model using a watermark embedding function to insert a periodic watermark signal in the probability values to generate the watermarked prediction outputs; and

outputting, from the first AI model, the watermarked prediction outputs including the periodic watermark signal.

12. The computer system of claim 11 wherein the periodic watermark signal is configured to cause an AI model that is distilled from the respective set of prediction outputs to insert a periodic signal into prediction outputs of the AI model that can be detected as matching the periodic watermark signal.

13. The computer system of claim 12 wherein the method includes, prior to the predicting, including the watermark embedding function in a preliminary AI model to generate the first AI model.

14. The computer system of claim 13 wherein the preliminary AI model is an untrained model, the method comprising training the preliminary AI model using a loss that is based on the outputs of the preliminary AI model with the included watermark embedding function.

15. The computer system of claim 13 , wherein the method includes defining a key corresponding to the watermark embedding function, the key including a random projection vector, wherein the watermark embedding function inserts the periodic watermark signal based on the random projection vector.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2022
From: HUAWEI TECHNOLOGIES CO., LTD.
To: HUAWEI CLOUD COMPUTING TECHNOLOGIES CO., LTD.
Reel/Frame 059267/0088 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2021
From: CHARETTE, LAURENT; CHU, LINGYANG; WANG, LANJUN; ZHANG, YONG
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 056024/0675 →
Continuity (1)
Related Publication 20220300842A1 · Sep 22, 2022
References Cited (16)
US 20190370440A1 · Gu et al. · 2019 [cited by applicant]
US 20200184044A1 · Zatloukal · 2020 [cited by applicant]
US 20200210764A1 · Hamedi · 2020 [cited by examiner]
US 20210264161A1 · Saraee · 2021 [cited by examiner]
US 20210272011A1 · Yonetani · 2021 [cited by examiner]
US 20210342490A1 · Briancon · 2021 [cited by examiner]
US 20210344997A1 · Anderson · 2021 [cited by examiner]
US 20220308895A1 · Ben-Elazar · 2022 [cited by examiner]
CN 112334917A · 2021 [cited by examiner]
WO WO2022018736A1 · 2022 [cited by examiner]
Tang Yong et al.,“An algorithm of digital watermark based on auto-correlation function in the discrete wavelet transform domain”,Yanshan University, vol. 31, No. 3, with an English abstract, total 6 pages. May 2015. [cited by applicant]
Jie Zhang et al., “Deep Model Intellectual Property Protection via Deep Watermarking”,Journal of Latex Class Files, vol. 14, No. 8, total 14 pages. Aug. 2015. [cited by applicant]
Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al, 9 pages, Mar. 9, 2015. [cited by applicant]
Entangled Watermarks as a Defence Against Model Extraction, Hengrui Jia et al., 18 pages, Feb. 19, 2021. [cited by applicant]
Deep Nerual Network Fingerprints by Conferrable Adversarial Examples, Nils Lukas et al., 18 paes, Jan. 20, 2021. [cited by applicant]
DAWN: Dynamic Adversarial Watermarking of Neural Networks, Sebastian Szyller et al., 16 pages, Jun. 18, 2020. [cited by applicant]