IP Library Patent Application 19417113
Patent Application
App. No. 19/417,113

GENERATIVE NEURAL NETWORKS WITH INVISIBLE TOKENS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/417,113
Abstract

Methods, systems, and apparatuses, including computer programs encoded on computer storage media, for processing a network input using a generative neural network to generate an output sequence of output tokens. The system selects each output token from a vocabulary of tokens that includes a plurality of visible tokens and one or more pairs of invisible tokens. The system processes the output sequence of output tokens to generate a final output sequence by removing, from the output sequence, the beginning invisible token, the end invisible token, and each visible token that is between the beginning invisible token and the end invisible token. The system then provides the final output sequence in response to the network input.

Claims (50)

1 . A method performed by a set of one or more computers, the method comprising:

receiving a network input;

processing the network input using a generative neural network to generate an output sequence of output tokens, wherein each output token is selected from a vocabulary of tokens that includes a plurality of visible tokens and one or more pairs of invisible tokens, each pair of invisible tokens comprising a respective beginning invisible token and a respective end invisible token;

processing the output sequence of output tokens to generate a final output sequence, comprising:

determining that the output sequence includes a beginning invisible token from one of the pairs followed by an end invisible token from the same pair; and

in response, removing, from the output sequence, the beginning invisible token, the end invisible token, and each visible token that is between the beginning invisible token and the end invisible token in the output sequence; and

providing the final output sequence in response to the network input.

2 . The method of claim 1 , wherein the generative neural network comprises an auto-regressive neural network that auto-regressively generates tokens from the vocabulary, and wherein the output sequence comprises visible tokens that are after the end invisible token in the output sequence and that are generated conditioned on the visible tokens that are between the beginning invisible token and the end invisible token in the output sequence.

3 . The method of claim 1 , wherein the visible tokens comprise text tokens that represent text data.

4 . The method of claim 1 , wherein the visible tokens comprise image tokens that represent image data.

5 . The method of claim 1 , wherein the visible tokens comprise audio tokens that represent audio data.

6 . The method of claim 1 , wherein:

the network input is received from a user device,

the set of one or more computers are remote from the user device, and

providing the final output sequence in response to the network input comprises providing the final output sequence to the user device.

7 . The method of claim 6 , wherein removing, from the output sequence, the beginning invisible token, the end invisible token, and each visible token that is between the beginning invisible token and the end invisible token in the output sequence is performed by the set of one or more computers prior to providing the final output sequence to the user device, such that the tokens between the beginning invisible token and the end invisible token are not transmitted to the user device.

8 . The method of claim 1 , wherein:

the network input is received as input from a user device,

the set of one or more computers includes only the user device, and

providing the final output sequence in response to the network input comprises providing the final output sequence for presentation on the user device.

9 . The method of claim 1 , wherein:

the set of one or more computers includes a server remote from a user device and the user device,

the server performs the processing of the network input using the generative neural network to generate the output sequence and transmits the output sequence to the user device, and

the user device performs the processing of the output sequence to generate the final output sequence by removing the beginning invisible token, the end invisible token, and each visible token that is between the beginning invisible token and the end invisible token.

10 . The method of claim 1 , wherein the network input comprises an initial prompt that characterizes a media item to be generated by the generative neural network, and wherein the tokens between the beginning invisible token and end invisible token represent an expanded prompt for generating the media item.

11 . The method of claim 10 , wherein the media item is an image, a video, or an audio sample.

12 . The method of claim 11 , wherein the generative neural network is configured to generate the media item conditioned on the tokens between the beginning invisible token and end invisible token, and wherein the method further comprises:

providing the media item in response to the network input.

13 . The method of claim 1 , wherein the network input comprises a query, wherein the tokens between the beginning invisible token and end invisible token represent intermediate data for generating a response to the query, and wherein the output sequence further comprises visible tokens that follow the end invisible token that represent the response to the query generated conditioned on the intermediate data.

14 . The method of claim 13 , wherein the network input further comprises an image or a video and the query is a query about the image or video.

15 . The method of claim 14 , wherein the intermediate data is grounding data specifying locations in the image or video of one or more objects.

16 . The method of claim 13 , wherein the intermediate data is a reasoning output.

17 . The method of claim 1 , wherein the vocabulary of tokens includes a plurality of distinct pairs of invisible tokens, and wherein removing, from the output sequence, the beginning invisible token, the end invisible token, and each visible token that is between the beginning invisible token and the end invisible token in the output sequence is performed according to a specific handling action determined based on an identity of the pair of invisible tokens included in the output sequence.

18 . The method of claim 17 , wherein the plurality of distinct pairs of invisible tokens includes a first pair associated with a first handling action and a second pair associated with a second handling action that is different from the first handling action.

19 . A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations, the operations comprising:

receiving a network input;

processing the network input using a generative neural network to generate an output sequence of output tokens, wherein each output token is selected from a vocabulary of tokens that includes a plurality of visible tokens and one or more pairs of invisible tokens, each pair of invisible tokens comprising a respective beginning invisible token and a respective end invisible token;

processing the output sequence of output tokens to generate a final output sequence, comprising:

determining that the output sequence includes a beginning invisible token from one of the pairs followed by an end invisible token from the same pair; and

in response, removing, from the output sequence, the beginning invisible token, the end invisible token, and each visible token that is between the beginning invisible token and the end invisible token in the output sequence; and

providing the final output sequence in response to the network input.

20 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations, the operations comprising:

receiving a network input;

processing the network input using a generative neural network to generate an output sequence of output tokens, wherein each output token is selected from a vocabulary of tokens that includes a plurality of visible tokens and one or more pairs of invisible tokens, each pair of invisible tokens comprising a respective beginning invisible token and a respective end invisible token;

processing the output sequence of output tokens to generate a final output sequence, comprising:

determining that the output sequence includes a beginning invisible token from one of the pairs followed by an end invisible token from the same pair; and

in response, removing, from the output sequence, the beginning invisible token, the end invisible token, and each visible token that is between the beginning invisible token and the end invisible token in the output sequence; and

providing the final output sequence in response to the network input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 29, 2025
From: SORICUT, RADU; BISHOP, COLTON M
To: GDM HOLDING LLC
Reel/Frame 073324/0861 →