MLchartDataset catalogue

Patent · US2026105650A1 · A1 · US

Systems and methods for generating multimodal data using a single-tower architecture with a data generation subsystem

(11) Publication number
US2026105650A1
(21) Application number
19/359,570
(22) Filing date
2025-10-15
(30) Priority date
2024-04-29
(43) Publication date
2026-04-16
(51) IPC
G06F 40/284; G06F 40/40; G06T 11/00; G06V 10/774; G06V 10/82; G06V 30/19; G10L 25/18; G10L 25/30
(52) CPC
  • G06T Image data processing or generation, in general: 11/00
  • G06F Electric digital data processing: 40/284, 40/40
  • G06V Image or video recognition or understanding: 10/774, 10/82, 30/19147
  • G10L Speech analysis techniques or speech synthesis; speech recognition; speech or voice processing techniques; speech or audio coding or decoding: 25/18, 25/30
(73) Assignee
GDM Holding LLC
(72) Inventors
Mostafa Dehghani; Phillip LIPPE; Emiel Hoogeboom; Jonathan Heek
(54) Title
Systems and methods for generating multimodal data using a single-tower architecture with a data generation subsystem
(57) Abstract

A computer-implemented method of generating multimodal data. The method comprises using a token generation neural network to generate, autoregressively, an output sequence of multimodal tokens, and in response to a next multimodal token being a start-of-image token, generating an image using an image generation subsystem conditioned on features representing the current sequence of multimodal tokens obtained from the token generation neural network. The method further comprises processing the image to convert pixels of the image into a sequence of image tokens, each image token comprising a block encoding of values of the pixels in a different region of the image that maps a set of values of the pixels to a respective image token, and appending the sequence of image tokens to the current output sequence of multimodal tokens as the next multimodal tokens in the output sequence of multimodal tokens.

Full text
View on Google Patents

Claims (1)

  1. (canceled) 2. A computer-implemented method of generating multimodal data using a generative system comprising a token generation neural network, the method comprising: receiving a prompt sequence that defines an input sequence of multimodal tokens, and processing the input sequence of multimodal tokens using the token generation neural network to generate an output sequence of multimodal tokens, wherein a multimodal token represents a data element of one of a plurality of modalities; wherein generating the output sequence of multimodal tokens comprises autoregressively, for each successive position in the output sequence of multimodal tokens: processing a combined sequence comprising the input sequence of multimodal tokens and a current output sequence of multimodal tokens, using the token generation neural network, to generate a next multimodal token for the output sequence of multimodal tokens, and appending the next multimodal token to the current output sequence of multimodal tokens; the method further comprising, in response to the next multimodal token being a start-of-image token: generating an image by performing a reverse diffusion process conditioned on features representing the current output sequence of multimodal tokens obtained from the token generation neural network; and appending a sequence of image tokens representing the image to the current output sequence of multimodal tokens as subsequent multimodal tokens in the output sequence of multimodal tokens. 3. The method of claim 2, further comprising continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens, to generate further multimodal tokens for the output sequence of multimodal tokens. 4. The method of claim 3, wherein the token generation neural network comprises one or more self-attention neural network layers, and wherein continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens comprises using bi-directional attention whilst the self-attention neural network layers are processing the image tokens when continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens. 5. The method of claim 4, wherein continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens comprises using self-attention with a causal mask whilst the self-attention neural network layers are processing other multimodal tokens of a different one of the plurality of modalities when continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens 6. The method of claim 2, wherein the input sequence of multimodal tokens comprises multimodal tokens representing text, audio, or image data elements, and wherein the output sequence of multimodal tokens comprises multimodal tokens representing text or audio data elements. 7. The method of claim 2, wherein performing the reverse diffusion process comprises: initializing a latent vector representation of the image, by sampling values for the latent vector representation from a noise distribution; and at each of a series of time steps: determining an updated version of the latent vector representation, by processing the latent vector representation conditioned on the features representing the current output sequence of multimodal tokens, to determine a reduced noise version of the latent vector representation. 8. The method of claim 7, wherein processing the latent vector representation conditioned on the features representing the current output sequence of multimodal tokens comprises attending over the features representing the current output sequence. 9. The method of claim 2, wherein the reverse diffusion process is performed in a latent variable space. 10. The method of claim 2, comprising: processing the combined sequence including the start-of-image token using the token generation neural network to generate features of a summary multimodal token; and wherein performing the reverse diffusion process comprises performing the reverse diffusion process conditioned on the features of the summary multimodal token. 11. The method of claim 2, wherein the image comprises an audio spectrogram; the method further comprising converting the audio spectrogram to time series audio data for an audio waveform. 12. The method of claim 11, wherein the prompt sequence comprises text or audio that defines an audio generation task; and wherein the time series audio data for the audio waveform defines audio that is specified by the prompt sequence. 13. The method of claim 2, wherein the prompt sequence comprises text or audio that defines an image generation task or image processing task; and wherein the image defines a result of the task; in particular wherein: i) the task comprises generating an image specified by the prompt; ii) the prompt includes an image and the task comprises generating a modified version of the image, where a modification to be performed is described by the prompt; iii) the prompt includes an image and the task is an optical character recognition task that involves generating an output sequence of multimodal tokens that represents words or characters in the image; iv) the prompt includes an image and the task comprises generating an output sequence of multimodal tokens that represents an answer to a question about the image; v) the prompt includes an image and identifies one or more objects in the image and the task comprises generating an output sequence of multimodal tokens that defines a presence, location, orientation, or count of one or more of the objects in the image; vi) the prompt includes an image and the task comprises generating an output sequence of multimodal tokens that describes a content of the image or that classifies a content of the image into one or more of a plurality of categories; or vii) the prompt includes an image and defines a goal for a mechanical agent acting in a real world environment and the task comprises generating an output sequence of multimodal tokens that defines one or more actions to be performed by the mechanical agent to achieve the goal. 14. The method of claim 2, wherein the generative system has been trained on interleaved sequences of text-and-image sequences. 15. The method of claim 2, wherein the generative system has been trained on a token prediction objective and on a diffusion model training objective. 16. The method of claim 2, wherein performing the reverse diffusion process comprises: initializing a latent representation of the image by sampling noise from a noise distribution, and at each of one or more updating iterations, updating the latent representation to reduce noise in the latent representation conditioned on the features. 17. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or computers cause the one or more computers to perform operations for generating multimodal data using a generative system comprising a token generation neural network, the operations comprising: receiving a prompt sequence that defines an input sequence of multimodal tokens, and processing the input sequence of multimodal tokens using the token generation neural network to generate an output sequence of multimodal tokens, wherein a multimodal token represents a data element of one of a plurality of modalities; wherein generating the output sequence of multimodal tokens comprises autoregressively, for each successive position in the output sequence of multimodal tokens: processing a combined sequence comprising the input sequence of multimodal tokens and a current output sequence of multimodal tokens, using the token generation neural network, to generate a next multimodal token for the output sequence of multimodal tokens, and appending the next multimodal token to the current output sequence of multimodal tokens; the method further comprising, in response to the next multimodal token being a start-of-image token: generating an image by performing a reverse diffusion process conditioned on features representing the current output sequence of multimodal tokens obtained from the token generation neural network; and appending a sequence of image tokens representing the image to the current output sequence of multimodal tokens as subsequent multimodal tokens in the output sequence of multimodal tokens. 18. The system of claim 17, wherein performing the reverse diffusion process comprises: initializing a latent representation of the image by sampling noise from a noise distribution, and at each of one or more updating iterations, updating the latent representation to reduce noise in the latent representation conditioned on the features. 19. The system of claim 17, further comprising continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens, to generate further multimodal tokens for the output sequence of multimodal tokens. 20. The system of claim 19, wherein the token generation neural network comprises one or more self-attention neural network layers, and wherein continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens comprises using bi-directional attention whilst the self-attention neural network layers are processing the image tokens when continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens. 21. One or more non-transitory computer storage media storing instructions that when executed by one or computers cause the one or more computers to perform operations for generating multimodal data using a generative system comprising a token generation neural network, the operations comprising: receiving a prompt sequence that defines an input sequence of multimodal tokens, and processing the input sequence of multimodal tokens using the token generation neural network to generate an output sequence of multimodal tokens, wherein a multimodal token represents a data element of one of a plurality of modalities; wherein generating the output sequence of multimodal tokens comprises autoregressively, for each successive position in the output sequence of multimodal tokens: processing a combined sequence comprising the input sequence of multimodal tokens and a current output sequence of multimodal tokens, using the token generation neural network, to generate a next multimodal token for the output sequence of multimodal tokens, and appending the next multimodal token to the current output sequence of multimodal tokens; the method further comprising, in response to the next multimodal token being a start-of-image token: generating an image by performing a reverse diffusion process conditioned on features representing the current output sequence of multimodal tokens obtained from the token generation neural network; and appending a sequence of image tokens representing the image to the current output sequence of multimodal tokens as subsequent multimodal tokens in the output sequence of multimodal tokens.

Description

This specification relates to processing data using machine learning models.

Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input.

Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.

This specification describes a method, implemented as a computer program on one or more computers in one or more locations, for generating multimodal data. A method of training a system for generating multimodal data items is also described. Corresponding systems are also described. Some implementations of the described techniques address issues of “negative transfer”, where training on multiple modalities adversely affects performance on individual modalities.

In a first aspect there is described a computer-implemented method of generating multimodal data using a system. The system includes a token generation neural network and a data, e.g., image, generation subsystem. The data, e.g., image, generation subsystem can comprise an image generation neural network; it can implement a diffusion model.

Record as JSON
{
  "publication_number": "US2026105650A1",
  "country": "US",
  "kind": "A1",
  "title": "Systems and methods for generating multimodal data using a single-tower architecture with a data generation subsystem",
  "abstract": "A computer-implemented method of generating multimodal data. The method comprises using a token generation neural network to generate, autoregressively, an output sequence of multimodal tokens, and in response to a next multimodal token being a start-of-image token, generating an image using an image generation subsystem conditioned on features representing the current sequence of multimodal tokens obtained from the token generation neural network. The method further comprises processing the image to convert pixels of the image into a sequence of image tokens, each image token comprising a block encoding of values of the pixels in a different region of the image that maps a set of values of the pixels to a respective image token, and appending the sequence of image tokens to the current output sequence of multimodal tokens as the next multimodal tokens in the output sequence of multimodal tokens.",
  "claims": [
    "1. (canceled) 2. A computer-implemented method of generating multimodal data using a generative system comprising a token generation neural network, the method comprising: receiving a prompt sequence that defines an input sequence of multimodal tokens, and processing the input sequence of multimodal tokens using the token generation neural network to generate an output sequence of multimodal tokens, wherein a multimodal token represents a data element of one of a plurality of modalities; wherein generating the output sequence of multimodal tokens comprises autoregressively, for each successive position in the output sequence of multimodal tokens: processing a combined sequence comprising the input sequence of multimodal tokens and a current output sequence of multimodal tokens, using the token generation neural network, to generate a next multimodal token for the output sequence of multimodal tokens, and appending the next multimodal token to the current output sequence of multimodal tokens; the method further comprising, in response to the next multimodal token being a start-of-image token: generating an image by performing a reverse diffusion process conditioned on features representing the current output sequence of multimodal tokens obtained from the token generation neural network; and appending a sequence of image tokens representing the image to the current output sequence of multimodal tokens as subsequent multimodal tokens in the output sequence of multimodal tokens. 3. The method of claim 2, further comprising continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens, to generate further multimodal tokens for the output sequence of multimodal tokens. 4. The method of claim 3, wherein the token generation neural network comprises one or more self-attention neural network layers, and wherein continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens comprises using bi-directional attention whilst the self-attention neural network layers are processing the image tokens when continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens. 5. The method of claim 4, wherein continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens comprises using self-attention with a causal mask whilst the self-attention neural network layers are processing other multimodal tokens of a different one of the plurality of modalities when continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens 6. The method of claim 2, wherein the input sequence of multimodal tokens comprises multimodal tokens representing text, audio, or image data elements, and wherein the output sequence of multimodal tokens comprises multimodal tokens representing text or audio data elements. 7. The method of claim 2, wherein performing the reverse diffusion process comprises: initializing a latent vector representation of the image, by sampling values for the latent vector representation from a noise distribution; and at each of a series of time steps: determining an updated version of the latent vector representation, by processing the latent vector representation conditioned on the features representing the current output sequence of multimodal tokens, to determine a reduced noise version of the latent vector representation. 8. The method of claim 7, wherein processing the latent vector representation conditioned on the features representing the current output sequence of multimodal tokens comprises attending over the features representing the current output sequence. 9. The method of claim 2, wherein the reverse diffusion process is performed in a latent variable space. 10. The method of claim 2, comprising: processing the combined sequence including the start-of-image token using the token generation neural network to generate features of a summary multimodal token; and wherein performing the reverse diffusion process comprises performing the reverse diffusion process conditioned on the features of the summary multimodal token. 11. The method of claim 2, wherein the image comprises an audio spectrogram; the method further comprising converting the audio spectrogram to time series audio data for an audio waveform. 12. The method of claim 11, wherein the prompt sequence comprises text or audio that defines an audio generation task; and wherein the time series audio data for the audio waveform defines audio that is specified by the prompt sequence. 13. The method of claim 2, wherein the prompt sequence comprises text or audio that defines an image generation task or image processing task; and wherein the image defines a result of the task; in particular wherein: i) the task comprises generating an image specified by the prompt; ii) the prompt includes an image and the task comprises generating a modified version of the image, where a modification to be performed is described by the prompt; iii) the prompt includes an image and the task is an optical character recognition task that involves generating an output sequence of multimodal tokens that represents words or characters in the image; iv) the prompt includes an image and the task comprises generating an output sequence of multimodal tokens that represents an answer to a question about the image; v) the prompt includes an image and identifies one or more objects in the image and the task comprises generating an output sequence of multimodal tokens that defines a presence, location, orientation, or count of one or more of the objects in the image; vi) the prompt includes an image and the task comprises generating an output sequence of multimodal tokens that describes a content of the image or that classifies a content of the image into one or more of a plurality of categories; or vii) the prompt includes an image and defines a goal for a mechanical agent acting in a real world environment and the task comprises generating an output sequence of multimodal tokens that defines one or more actions to be performed by the mechanical agent to achieve the goal. 14. The method of claim 2, wherein the generative system has been trained on interleaved sequences of text-and-image sequences. 15. The method of claim 2, wherein the generative system has been trained on a token prediction objective and on a diffusion model training objective. 16. The method of claim 2, wherein performing the reverse diffusion process comprises: initializing a latent representation of the image by sampling noise from a noise distribution, and at each of one or more updating iterations, updating the latent representation to reduce noise in the latent representation conditioned on the features. 17. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or computers cause the one or more computers to perform operations for generating multimodal data using a generative system comprising a token generation neural network, the operations comprising: receiving a prompt sequence that defines an input sequence of multimodal tokens, and processing the input sequence of multimodal tokens using the token generation neural network to generate an output sequence of multimodal tokens, wherein a multimodal token represents a data element of one of a plurality of modalities; wherein generating the output sequence of multimodal tokens comprises autoregressively, for each successive position in the output sequence of multimodal tokens: processing a combined sequence comprising the input sequence of multimodal tokens and a current output sequence of multimodal tokens, using the token generation neural network, to generate a next multimodal token for the output sequence of multimodal tokens, and appending the next multimodal token to the current output sequence of multimodal tokens; the method further comprising, in response to the next multimodal token being a start-of-image token: generating an image by performing a reverse diffusion process conditioned on features representing the current output sequence of multimodal tokens obtained from the token generation neural network; and appending a sequence of image tokens representing the image to the current output sequence of multimodal tokens as subsequent multimodal tokens in the output sequence of multimodal tokens. 18. The system of claim 17, wherein performing the reverse diffusion process comprises: initializing a latent representation of the image by sampling noise from a noise distribution, and at each of one or more updating iterations, updating the latent representation to reduce noise in the latent representation conditioned on the features. 19. The system of claim 17, further comprising continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens, to generate further multimodal tokens for the output sequence of multimodal tokens. 20. The system of claim 19, wherein the token generation neural network comprises one or more self-attention neural network layers, and wherein continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens comprises using bi-directional attention whilst the self-attention neural network layers are processing the image tokens when continuing to process the combined sequence after appending the sequence of image tokens to the current output sequence of multimodal tokens. 21. One or more non-transitory computer storage media storing instructions that when executed by one or computers cause the one or more computers to perform operations for generating multimodal data using a generative system comprising a token generation neural network, the operations comprising: receiving a prompt sequence that defines an input sequence of multimodal tokens, and processing the input sequence of multimodal tokens using the token generation neural network to generate an output sequence of multimodal tokens, wherein a multimodal token represents a data element of one of a plurality of modalities; wherein generating the output sequence of multimodal tokens comprises autoregressively, for each successive position in the output sequence of multimodal tokens: processing a combined sequence comprising the input sequence of multimodal tokens and a current output sequence of multimodal tokens, using the token generation neural network, to generate a next multimodal token for the output sequence of multimodal tokens, and appending the next multimodal token to the current output sequence of multimodal tokens; the method further comprising, in response to the next multimodal token being a start-of-image token: generating an image by performing a reverse diffusion process conditioned on features representing the current output sequence of multimodal tokens obtained from the token generation neural network; and appending a sequence of image tokens representing the image to the current output sequence of multimodal tokens as subsequent multimodal tokens in the output sequence of multimodal tokens."
  ],
  "description_excerpt": "This specification relates to processing data using machine learning models.\n\nMachine learning models receive an input and generate an output, e.g., a predicted output, based on the received input.\n\nNeural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.\n\nThis specification describes a method, implemented as a computer program on one or more computers in one or more locations, for generating multimodal data. A method of training a system for generating multimodal data items is also described. Corresponding systems are also described. Some implementations of the described techniques address issues of “negative transfer”, where training on multiple modalities adversely affects performance on individual modalities.\n\nIn a first aspect there is described a computer-implemented method of generating multimodal data using a system. The system includes a token generation neural network and a data, e.g., image, generation subsystem. The data, e.g., image, generation subsystem can comprise an image generation neural network; it can implement a diffusion model.",
  "cpc": [
    "G06T 11/00",
    "G06F 40/284",
    "G06F 40/40",
    "G06V 10/774",
    "G06V 10/82",
    "G06V 30/19147",
    "G10L 25/18",
    "G10L 25/30"
  ],
  "ipc": [
    "G06F 40/284",
    "G06F 40/40",
    "G06T 11/00",
    "G06V 10/774",
    "G06V 10/82",
    "G06V 30/19",
    "G10L 25/18",
    "G10L 25/30"
  ],
  "assignees": [
    "GDM Holding LLC"
  ],
  "inventors": [
    "Mostafa Dehghani",
    "Phillip LIPPE",
    "Emiel Hoogeboom",
    "Jonathan Heek"
  ],
  "filing_date": "2025-10-15",
  "publication_date": "2026-04-16",
  "priority_date": "2024-04-29",
  "application_number": "US-202519359570-A",
  "family_id": "95895637",
  "cited_by_count": 0
}

Record 31 of 8,000 in Patents full text (MLC-0201). Request the full dataset.