Speculative textual watermarking in machine learning
Aspects of the present disclosure involve watermark-preserving speculative text generation. Aspects include generating, by a draft machine learning model based on a prompt, a lookahead comprising a plurality of candidate next tokens with draft probabilities, wherein the draft machine learning model has been trained on watermarked outputs generated, based on a watermarking key, using a target machine learning model. Aspects include evaluating, in a single inference pass of the target machine learning model, the plurality of candidate next tokens, based on the watermarking key, to generate target probabilities. Aspects include accepting a first token subset of the plurality of candidate next tokens based on comparing the target probabilities to the draft probabilities. Aspects include generating, by the target machine learning model, one or more replacement tokens for a second token subset comprising each candidate next token of the plurality of candidate next tokens not included in the first token subset.
1 . A system for watermark-preserving speculative text generation, comprising:
one or more processors; and
a non-transitory computer readable medium storing instructions that, when executed by the one or more processors, cause the system to:
receive a generation request comprising a prompt;
generate, by a draft machine learning model based on the prompt, a lookahead comprising a plurality of candidate next tokens with respective draft probabilities, wherein the draft machine learning model has been trained on watermarked outputs generated, based on a watermarking key, using a target machine learning model;
evaluate, in a single inference pass of the target machine learning model configured to perform a generative watermarking process, the plurality of candidate next tokens, based on the watermarking key, to generate respective target probabilities;
accept a first token subset of the plurality of candidate next tokens based on comparing the respective target probabilities to the respective draft probabilities, wherein the accepted first token subset comprises each candidate next token of the plurality of candidate next tokens for which a respective target probability of the respective target probabilities is greater than or equal to a respective draft probability of the respective draft probabilities;
generate, by the target machine learning model, one or more replacement tokens for a second token subset comprising each candidate next token of the plurality of candidate next tokens not included in the first token subset, wherein the second token subset comprises each candidate next token of the plurality of candidate next tokens for which a respective target probability of the respective target probabilities is less than a respective draft probability of the respective draft probabilities; and
output, in response to the generation request, text that includes the accepted first token subset and the one or more replacement tokens.
2 . The system of claim 1 , wherein the generative watermarking process performed by the target machine learning model comprises tournament sampling across multiple layers using the watermarking key and repeated context masking.
3 . The system of claim 1 , wherein the draft machine learning model is trained on watermarked outputs generated by the target machine learning model using the watermarking key so as to approximate a distribution of the target machine learning model in connection with the generative watermarking process.
4 . The system of claim 1 , wherein the evaluating in a single inference pass comprises computing, in one forward pass of the target machine learning model, based on the prompt, a corresponding target probability of the respective target probabilities for each candidate next token of the lookahead.
5 . The system of claim 1 , wherein the generating of the one or more replacement tokens is performed by the target machine learning model using a same generative watermarking process as used to compute the respective target probabilities, thereby preserving watermark guarantees.
6 . The system of claim 1 , wherein the lookahead comprising the plurality of candidate next tokens has a configured length that is greater than one.
7 . The system of claim 1 , wherein the outputting of the text further comprises storing metadata in connection with the text to enable subsequent watermark detection without utilizing the target machine learning model.
8 . The system of claim 1 , wherein the draft machine learning model generates the plurality of candidate next tokens using a same set of decoding settings as the target machine learning model so that the respective draft probabilities are comparable to the respective target probabilities, and wherein the same set of decoding settings comprises one or more of:
top-k sampling;
top-p sampling; or
temperature.
9 . The system of claim 1 , wherein the accepting the first token subset and the generating the one or more replacement tokens are performed positionally with respect to the lookahead such that accepted candidate next tokens are appended, and rejected positions are filled by the one or more replacement tokens, in an original order of the lookahead.
10 . The system of claim 1 wherein the accepting of the first token subset is based on determining, for each position of the lookahead, whether a corresponding respective target probability of the respective target probabilities is greater than or equal to a corresponding respective draft probability of the respective draft probabilities.
11 . The system of claim 1 , wherein the draft machine learning model has fewer tunable parameters than the target machine learning model.
12 . A system for training a watermark-aware draft machine learning model, comprising:
one or more processors; and
a non-transitory computer readable medium storing instructions that, when executed by the one or more processors, cause the system to:
receive a training configuration comprising a watermarking key and a set of training prompts;
generate, by a target machine learning model configured to perform a generative watermarking process based on the watermarking key, a plurality of watermarked outputs for the set of training prompts;
assemble a training corpus comprising, for each training prompt of the set of training prompts, the training prompt, a corresponding watermarked output of the plurality of watermarked outputs, the watermarking key, and one or more target distribution signals derived from the target machine learning model under the generative watermarking process; and
train, using the training corpus, a draft machine learning model to predict, for lookahead positions, candidate next tokens with respective draft probabilities aligned to the distribution signals from the target machine learning model, wherein:
the trained draft machine learning model is used to generate a lookahead comprising candidate next tokens with associated draft probabilities;
the target machine learning model evaluates, in a single inference pass, the candidate next tokens to generate respective target probabilities;
a first token subset of the candidate next tokens is accepted, wherein the accepted first token subset comprises each candidate next token of the candidate next tokens for which a respective target probability of the respective target probabilities is greater than or equal to an associated draft probability of the associated draft probabilities; and
one or more replacement tokens are generated for a second token subset of the candidate next tokens, wherein the second token subset comprises each candidate next token of the candidate next tokens for which a respective target probability of the respective target probabilities is less than an associated draft probability of the associated draft probabilities.
13 . The system of claim 12 , wherein the instructions, when executed by the one or more processors, further cause the system to:
validate the draft machine learning model by computing, on a hold-out evaluation set, an acceptance metric that measures a proportion of candidate next tokens whose target probabilities from the target machine learning model according to the generative watermarking process are greater than or equal to corresponding draft probabilities from the draft machine learning model for a configured lookahead length;
store trained parameters of the draft machine learning model together with metadata comprising a key identifier for the watermarking key, and the acceptance metric; and
deploy the draft machine learning model for use in watermark-preserving speculative text generation with the target machine learning model.
14 . The system of claim 12 , wherein the target distribution signals comprise one or more of logits, probabilities, acceptance masks for lookahead positions, or tournament-layer assignments produced by the target machine learning model according to the generative watermarking process.
15 . The system of claim 12 , wherein the training the draft machine learning model comprises optimizing a loss function that penalizes divergence from the target distribution signals for the lookahead positions.
16 . The system of claim 12 , wherein the generative watermarking process performed by the target machine learning model comprises tournament sampling across multiple layers using the watermarking key and repeated context masking.
17 . The system of claim 12 , wherein the draft machine learning model is configured with a same set of decoding settings as the target machine learning model.
18 . The system of claim 12 , wherein the draft machine learning model has fewer tunable parameters than the target machine learning model.
19 . A method for watermark-preserving speculative text generation, comprising:
receiving a generation request comprising a prompt;
generating, by a draft machine learning model based on the prompt, a lookahead comprising a plurality of candidate next tokens with respective draft probabilities, wherein the draft machine learning model has been trained on watermarked outputs generated, based on a watermarking key, using a target machine learning model;
evaluating, in a single inference pass of the target machine learning model configured to perform a generative watermarking process, the plurality of candidate next tokens, based on the watermarking key, to generate respective target probabilities;
accepting a first token subset of the plurality of candidate next tokens based on comparing the respective target probabilities to the respective draft probabilities, wherein the accepted first token subset comprises each candidate next token of the plurality of candidate next tokens for which a respective target probability of the respective target probabilities is greater than or equal to a respective draft probability of the respective draft probabilities;
generating, by the target machine learning model, one or more replacement tokens for a second token subset comprising each candidate next token of the plurality of candidate next tokens not included in the first token subset, wherein the second token subset comprises each candidate next token of the plurality of candidate next tokens for which a respective target probability of the respective target probabilities is less than a respective draft probability of the respective draft probabilities; and
outputting, in response to the generation request, text that includes the accepted first token subset and the one or more replacement tokens.