Detailed Action
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 5, 9-10, 13 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Maharana et al (“StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation”), Chang et al (“MaskGIT: Masked Generative Image Transformer”), Zhou et al (“Conditional Prompt Learning for Vision-Language Models”), and Navarrete et al (US 20210049447 A1), hereinafter Maharana, Chang, Zhou, and Navarrete respectively.
Regarding claim 9, Maharana teaches a processing system comprising: a memory storing a pretrained generative image transformer
“We convert the pretrained DALL-E into StoryDALL-E model for the story continuation task.” – Pg 8, Par 2, Lines 1-2
NOTE: Maharana teaches the StoryDALL-E model designed for generating a sequence of images to form a coherent visual, see Fig. 1. This model uses a pretrained DALL- model as a part of the model’s function. “The DALL-E model introduced in [38] is a text-to-image synthesis pipeline which comprises of a discrete variational autoencoder (dVAE) in the first stage and an autoregressive transformer in the second stage” - Pg 21, Section A.2, Par 1. This functionally corresponds to the pretrained generative image transformer. A processor system that has a memory storing this model would naturally be required.
and a prompt token generator;
“We initialize a parameterization network MLP(.), which takes a matrix of trainable parameters P′θ of dimensions Pidx and dim(hi) as input and generates the prompt Pθ. These trainable matrices are randomly initialized and trained from scratch on the downstream task and dataset. Pθ is appended to the word embeddings of input caption, along with the global story embeddings. Together, these additional embedding vectors act as ‘virtual tokens’ of a task-specific prompt, and are attended to by each of the caption as image tokens." - Pg 7, Par 5
NOTE: The parameterization network multi later perceptron (MLP) as disclosed by Maharana functionally corresponds to the prompt token generator. This component receives trainable parameters P′θ and generates a prompt Pθ. Generated word embeddings are appended to the caption embeddings and act as “virtual tokens” of a task specific prompt, see Pg 7, Par 5.
and one or more processors coupled to the memory and configured to train the prompt token generator according to a training method comprising: for each given training example of a plurality of training examples,
“We evaluate our approach StoryDALL-E on two existing datasets, PororoSV and FlintstonesSV, and introduce a new dataset DiDeMoSV collected from a video-captioning dataset.” – Abstract
NOTE: Maharana discloses several datasets for training StoryDALL-E which would naturally provide a plurality of training examples for the training process.
the given training example including a target token sequence representing a first vector-quantized image
"A pretrained VQVAE encoder [35] is used to transform RGB images into small 2D grids of image tokens, which are flattened and concatenated with the modified inputs in StoryDALL-E" - Pg 8, Par 2 Lines 2-4
NOTE: Maharana discloses using a Vector Quantized Variational Autoencoder (VQVAE) to transform an image to grids of image token. VQVAE is the vector-quantized image representation used by the DALL-E model therefore the training image is converted into the vector quantized image representation.
generating, using the prompt token generator, a first sequence of prompt tokens
“We initialize a parameterization network MLP(.), which takes a matrix of trainable parameters P′θ of dimensions Pidx and dim(hi) as input and generates the prompt Pθ…Together, these additional embedding vectors act as ‘virtual tokens’ of a task-specific prompt, and are attended to by each of the caption as image tokens." - Pg 7, Par 5
NOTE: The parametrization network MLP which functionally corresponds to the prompt token generator generates virtual tokens of a task-specific prompt which correspond to generating the first sequence of prompt tokens
generating, using the pretrained generative image transformer, a first output token sequence based at least in part on the first sequence of prompt tokens,
“an autoregressive transformer is used to model the joint distribution over the text and image tokens. For a given text input s and target image x, these models learn the distribution of image tokens p(x)…. The prediction of image tokens at each time step is influenced by the text tokens and previously predicted image tokens via the self-attention layer.” – Pg 21-22, Section A.2, Par 3, Lines 3-10
NOTE: Maharana discloses that the autoregressive transformer receives input virtual tokens (prompt tokens) from the MLP component (which functionally corresponds to the prompt token generator. The transformer then predicts image tokens sequentially which can be understood as the output token sequence which are naturally based off the prompt tokens used as input.
the first output token sequence representing a second vector-quantized image;
“The frames are encoded using pretrained VQVAE and sent as inputs to the pretrained DALL-E. The inputs are prepended with input-agnostic prompt (in prompt-tuning setting only) and global story embeddings corresponding to each sample in the story continuation dataset. The output of StoryDALL-E is decoded using VQ-VAE to generate the predicted image." - Fig. 2 Caption
NOTE: Maharana discloses that the DALL-E output is the discrete image-token representation decoded by the VQVAE into the predicted image. The output token sequence would then naturally represent the newly generated predicted vector quantized image (second vector-quantized image)
Maharana does not teach the given training example includingthe given training example including a target token sequence representing a first vector-quantized image and a first set of one or more identifiers, at least one identifier of the first set of one or more identifiers relating to a subject of the first vector-quantized image:
“We evaluate the performance of our model on class conditional image synthesis on ImageNet 256ˆ256 and 512ˆ512.” – Pg 5, Section 4.2, Lines 1-3
NOTE: Chang discloses MaskGIT which is a model configured to tokenize an image and learns to generate those tokens through a training process comprising of replacing some of the tokens with a mask and using a transformer to predict the missing tokens, see abstract. The model is trained using ImageNet class labels, see Fig 14 caption. The class labels functionally correspond to the training examples comprising a first set of one or more identifiers relating to the subject of the image. After the combination, the class labels which functionally correspond to the identifiers related to the image as taught by Chang can modify the training examples including a target token sequence representing a first vector quantized image as taught by Maharana so that the training examples can comprise both the target token sequence and identifiers related to the first vector-quantized image.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to have the given training example including a target token sequence representing a first vector-quantized image and a first set of one or more identifiers, at least one identifier of the first set of one or more identifiers relating to a subject of the first vector-quantized image. One would be motivated to make this combination because having training examples comprise a target token sequence and at least one identifier related to the subject of the first vector-quantized image would lead to a predicted result of a more accurate training system. The target token sequence provides a ground truth which can be compared to output predicted token sequences which can guide the training of the prompt token generator. The identifiers further adds assistance to the prompt token generator by ensuring the predicted output does not stray too far from the intended subject.
Maharana in view of Chang still does not teach generating, using the prompt token generator, a first sequence of prompt tokens based at least in part on the first set of one or more identifiers. However, Zhou teaches generating, using the prompt token generator, a first sequence of prompt tokens based at least in part on the first set of one or more identifiers,
“Specifically, on top of the M context vectors, we further learn a lightweight neural network, called Meta-Net, to generate for each input a conditional token (vector)… The prompt for the i-th class is thus conditioned on the input, i.e., ti(x) = {v1(x),v2(x),...,vM(x),ci}.” – Pg 4. Section 3.2, Par 2-3
NOTE: Maharana as modified shows a prompt token generator and obtaining identifiers, but does not teach that the generated tokens are based in part of the identifiers. However, Zhou discloses the CoCoOp method which comprises a Meta-Net neural network configured to output conditional tokens for each image. CoCoOp’s input is an image feature that is derived from the image whose subject is being represented. The image feature related to the subject of the image corresponds to the identifier related to the subject of the image since the image feature would consist of information identifying the subject of the image. The prompt token formula represented as “ti” includes image feature “x”, see Pg 4, Section 3.1-3.2. After the combination, the concept of using image features which corresponds to the identifiers to generate token prompts as used in Zhou’s method can modify the identifiers as taught by Chang. This modification is then added to Maharana’s StoryDALL-E architecture so that the parameterization network MLP component (prompt token generator) which generates virtual token sequences can take input the identifiers obtained from Chang in order for the generated prompt tokens to be condition on information that identifies the subject of the image as shown by Zhou’s method.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Maharana by incorporating the teachings of Zhou to generating, using the prompt token generator, a first sequence of prompt tokens based at least in part on the first set of one or more identifiers. One would be motivated to make this combination because it would lead to the predicted results of creating more dynamic and accurate sequences of token prompts that adapt to the subject of the image from the training example.
Maharana in view of Chang and Zhou still does not teach comparing the first output token sequence to the target token sequence to generate a loss value for the given training example; and modifying one or more parameters of the prompt token generator based at least in part on the loss values generated for the plurality of training examples. However, Navarrete teaches comparing the first output token sequence to the target token sequence to generate a loss value for the given training example
“acquiring training input data and training target data; processing the training input data by using the neural network to obtain training output data; calculating a loss value of a loss function of the neural network according to the training output data and the training target data” – Par 16, Lines 5-10
NOTE: Navarrete teaches a standard training process of a neural network which comprises of acquiring training input data and training target data and calculating the loss value by comparing the output of the neural network with the target training data. After the combination, the comparison of the training output data and target data to compute a loss value taught by Navarrete can be added to Maharana as modified by Chang and Zhou. This will allow the output sequence tokens generated by the pretrained generative image transformer to be compared to the training examples which comprises the target token sequence to compute a loss value using the training process of Navarrete.
and modifying one or more parameters of the prompt token generator based at least in part on the loss values generated for the plurality of training examples.
“modifying parameters of the neural network based on the training weight of the at least one nonlinear layer and the loss value, obtaining a trained neural network in a case where the loss function of the neural network satisfies a predetermined condition, and continuously inputting the training input data and the training target data to repeatedly execute the above training process in a case where the loss function of the neural network does not satisfy the predetermined condition.” – Par 16, Lines 10-18
NOTE: Navarrete discloses modifying the parameters of the neural network based on the loss value obtained from comparing the target input data and output data. This training process is continuously performed until a predetermined conditioned is satisfied. After the combination, the modification of parameters of a neural network based on obtained loss values as taught by Navarrete can be added to Maharana as modified. This will then allow Maharana’s system to compare the training images comprising target token sequences and the output token sequence from the pretrained generative image transformer to obtain a loss value. The loss value will then be used to modify parameters of the prompt token generator as described in the training process taught by Navarrete.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Maharana by incorporating the teachings of Navarrete to compare the first output token sequence to the target token sequence to generate a loss value for the given training example. One would be motivated to make this combination since its standard practice for training models by obtaining a loss value between the predicted result and target ground truth. By comparing the target token sequence with the output token sequence, the loss value can help determine how close the predicted output token sequence was to the known target token sequence. This training process will lead to the predicted result of an improved prompt token generator with a small loss value that can accurately generate the appropriate tokens.
Regarding claim 1, the claim recites similar limitations to claim 9. Therefore, method claim 1 corresponds to the system disclosed in claim 9 and is rejected for the same reasons of obviousness as used above.
Regarding claim 20, the claim recites similar limitations to claim 1, 5, 9-10, 13 or 18. Therefore, non-transitory computer program product claim 20 corresponds to the system disclosed in claim 1, 5, 9 , 10, 13, 18 and is rejected for the same reasons of obviousness as used above.
Regarding claim 10, Maharana in view of Chang, Zhou, and Navarrete teaches the system of claim 9. Maharana as combined further teaches wherein the prompt token generator comprises a multi-layer perceptron.
“We initialize a parameterization network MLP(.), which takes a matrix of trainable parameters P′θ of dimensions Pidx and dim(hi) as input and generates the prompt Pθ. These trainable matrices are randomly initialized and trained from scratch on the downstream task and dataset. Pθ is appended to the word embeddings of input caption, along with the global story embeddings. Together, these additional embedding vectors act as ‘virtual tokens’ of a task-specific prompt, and are attended to by each of the caption as image tokens." - Pg 7, Par 5
NOTE Maharana discloses a parameterization network MLP component that generates prompts to be appended to word embeddings of an input caption to function as “virtual tokens”. This functionally corresponds to the prompt token generator comprising a multi-layer perceptron (MLP)
Regarding claim 13, Maharana in view of Chang, Zhou, and Navarrete teaches the system of claim 9. Maharana as combined further teaches wherein the one or more processors are further configured to: generate, using the prompt token generator, a second sequence of prompt tokens based at least in part on a second set of one or more identifiers;
“We initialize a parameterization network MLP(.), which takes a matrix of trainable parameters P′θ of dimensions Pidx and dim(hi) as input and generates the prompt Pθ…Together, these additional embedding vectors act as ‘virtual tokens’ of a task-specific prompt, and are attended to by each of the caption as image tokens." – Maharana Pg 7, Par 5
NOTE: Claim 13 describes a similar limitation regarding a prompt token generator generating the prompt tokens based in part on a set of identifiers as claim 9. However, the claim’s limitations merely describe a second loop of the training process. One of ordinary skill would recognize that after the first instance of the training, a second iteration can occur to produce the second sequence of prompt tokens based on a second set of identifiers. See rejection of claim 9.
and generate, using the pretrained generative image transformer, a second output token sequence based at least in part on the second sequence of prompt tokens, the second output token sequence representing a third vector-quantized image.
“an autoregressive transformer is used to model the joint distribution over the text and image tokens. For a given text input s and target image x, these models learn the distribution of image tokens p(x)…. The prediction of image tokens at each time step is influenced by the text tokens and previously predicted image tokens via the self-attention layer.” – Pg 21-22, Section A.2, Par 3, Lines 3-10
NOTE: Claim 13 describes a similar limitation regarding a generative image transformer generating the output token sequence representing a vector-quantized image based in part on a sequence of prompt tokens as claim 9. One of ordinary skill in the art would recognize that the second sequence of prompt tokens as generated during the second loop of the training process from the prompt token generator is then fed to the generative image transformer to produce a second output sequence representing a third vector-quantized image. See rejection of claim 9.
Regarding claim 5, the claim recites similar limitations to claim 13. Therefore, method claim 5 corresponds to the system disclosed in claim 13 and is rejected for the same reasons of obviousness as used above.
Regarding claim 18, Maharana in view of Chang, Zhou, and Navarrete teaches the system of claim 9. Maharana as combined further teaches wherein the pretrained generative image transformer is an autoregressive image transformer.
“an autoregressive transformer is used to model the joint distribution over the text and image tokens.” – Maharana Pg 21, Section A.2, Par 3, Lines 3-4
Regarding claim 19, Maharana in view of Chang, Zhou, and Navarrete teaches the system of claim 9. Maharana as combined further teaches wherein the pretrained generative image transformer is a non-autoregressive image transformer.
“MaskGIT adopts a novel non-autoregressive decoding method” – Chang Pg 2, Par 2, Lines 5-6 and “We introduce a novel decoding method where all tokens in the image are generated simultaneously in parallel. This is feasible due to the bi-directional self-attention of MTVM.”
NOTE: Chang discloses a decoding method where all tokens in the image are generated simultaneously in parallel. This is functionally the same as a non-autoregressive model. After the combination the non-autoregressive decoding method used for the generative image transformer as taught by Chang can be used as substitute to Maharana’s autoregressive image transformer.
Claim(s) 2, 11-12, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Maharana, Chang, Zhou, Navarrete, and He et al (“HyperPrompt: Prompt-based Task-Conditioning of Transformers”), hereinafter He respectively.
Regarding claim 2, Maharana in view of Chang, Zhou, and Navarrete teaches the method of claim 1. Maharana as combined does not teach wherein the prompt token generator comprises two or more multi-layer perceptrons, and wherein modifying the one or more parameters of the prompt token generator comprises modifying one or more parameters of each of the two or more multi-layer perceptrons. However, He teaches wherein the prompt token generator comprises two or more multi-layer perceptrons,
we apply two local HyperNetworks hm k and hm v to transform the global prompt Pτ into layer-specific and task-specific prompts as shown in Figure 2(b): Pm τ,k = hm k(Pτ) = Um k (Relu(Dm k(Pτ))), (3) HyperPrompt-Sep. In the opposite extreme of HyperPrompt-Share, each task can have its own local HyperNetworks hm τ,k(Pτ) and hm τ,v(Pτ) as following: Pm τ,k = hm τ,k(Pτ) = Um τ,k(Relu(Dm τ,k(Pτ))), Pm τ,v = hm τ,v(Pτ) = Um τ,v(Relu(Dm τ,v(Pτ))), where Dm τ,k/v and Um (5) (6) τ,k/v are down-projection and up-projection matrices for the τ task, respectively. Pm τ,v = hm v(Pτ) = Um v (Relu(Dm v (Pτ))), (4)” – Pg 4, Par 1
NOTE: He teaches a prompt generator component that uses two “hypernetworks” which are two projection networks: “where Im τ is the input to the shared global HyperNetworks as shown in Figure 1(c). ht is a MLP consisting of two feed-forward layers and a ReLU non-linearity, which takes the concatenation of kτ and zm as input.”, see Pg 5, Par 1. This corresponds to two multi-layer perceptron architecture since each of the two hypernetworks contains the MLP architecture. After the combination, the concept of using two multi-layer perceptron architectures for prompt generation as taught by He can modify Maharana’s prompt token generator that already uses one multi-layer perceptron for generating prompt tokens so that the prompt token generator can use two multi-layer perceptrons for generating prompt tokens.
and wherein modifying the one or more parameters of the prompt token generator comprises modifying one or more parameters of each of the two or more multi-layer perceptrons.
“Prompt-tuning is an alternative [26] to full model fine-tuning where the pretrained model weights are frozen and instead, a small sequence of task-specific vectors is optimized for the downstream task. We initialize a parameterization network MLP(.), which takes a matrix of trainable parameters P′ θ of dimensions Pidx and dim(hi) as input and generates the prompt Pθ. These trainable matrices are randomly initialized and trained from scratch on the downstream task and dataset….New parameters as well as pretrained weights are optimized in full-model finetuning whereas only the parameters of the prompt, story encoder and cross-attention layers are optimized during prompt-tuning” – Pg 7 Par 4 and Pg 8 Par 1
NOTE: Maharana discloses modifying the parameters of a prompt token generator using a multi-layer perceptron architecture as discussed in the rejection of claim 1. The steps for modifying the parameters of one multi-layer perceptron would apply to the second multi-layer perceptron. After the combination, the concept of using two multi-layer perception based architectures as shown in He’s hypernetworks can modify Maharana’s prompt token generator to use two multi-layer perceptrons and modify the parameters of the MLPs during the training process.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Maharana by incorporating the teachings of He to have the prompt token generator comprise two or more multi-layer perceptrons and wherein modifying the one or more parameters of the prompt token generator comprises modifying one or more parameters of each of the two or more multi-layer perceptrons. One would be motivated to make this combination since using two or more multi-layer perceptrons would lead to the predicted result of a more efficient prompt token generator. The input to the prompt token generator would be able to be processed faster with two or more multi-layer perceptrons and would generate more accurate prompt tokens during the training process.
Regarding claim 11, Maharana in view of Chang, Zhou, and Navarrete teaches the system of claim 9. The claim recites similar limitations to claim 2. Therefore, system claim 11 corresponds to the system method in claim 2 and is rejected for the same reasons of obviousness as used above.
Regarding claim 12, Maharana in view of Chang, Zhou, Navarrete, and He teaches the system of claim 12. The claim recites similar limitations to claim 2. Therefore, system claim 12 corresponds to the system method in claim 2 and is rejected for the same reasons of obviousness as used above.
Regarding claim 20, the claim recites similar limitations to claim 2 or 11 or 12. Therefore, non-transitory computer program product claim 20 corresponds to the system disclosed in claim 2, or 11, or 12 and is rejected for the same reasons of obviousness as used above.
Claim(s) 3-4, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Maharana, Chang, Zhou, Navarrete, and Abhiram et al (US 11429820 B2), hereinafter Abhiram respectively.
Regarding claim 3, Maharana in view of Chang, Zhou, and Navarrete teaches the method of claim 1. Maharana as combined further teaches wherein the first set of one or more identifiers of the given training example comprises a class identifier relating to the subject of the first vector-quantized image.
“Process 300 includes a step in which the system receives a first training bundle comprising one or more first training images, a true class category vector for an object instance depicted in the one or more first training images, and a first true object instance identifier for the object instance depicted in the one or more first training images (302).” – Col 4, Lines 59-64
NOTE: Abhiram discloses a training bundle (plurality of training examples) which comprise a true class category vector for an object depicted in the image. Abhiram provides an example: “The true class category vector may be, for example, a vector or list of object labels 216 (e.g., if the object instance depicted is a German shepherd dog, the vector may indicate a probability of 1 or a Boolean ‘true’ corresponding to the class category German shepherd dog,” – Col 4, Lines 65-68 and Col 5 Lines 1-2. This functionally corresponds to a class identifier related to the subject of the image. After the combination, obtaining the class identifiers of training examples as taught by Abhiram can modify the training examples consisting of vector-quantized images taught by Maharana as combined so that the training examples can comprise class identifiers related to the subject of the first vector-quantized image.
It would have been obvious to one of ordinary skill before the effective filing date of the present invention to modify Maharana by incorporating the teachings of Abhiram to have the first set of one or more identifiers of the given training example comprises a class identifier relating to the subject of the first vector-quantized image because it would lead to the predicted result of providing additional information that distinguishes the subject in the image which would allow the prompt token generator to produce accurate prompt tokens during the training process.
Regarding claim 4, Maharana in view of Chang, Zhou, and Navarrete teaches the method of claim 1. Maharana as combined further teaches wherein the first set of one or more identifiers of the given training example comprises an instance identifier relating to the first vector-quantized image.
“Process 300 includes a step in which the system receives a first training bundle comprising one or more first training images, a true class category vector for an object instance depicted in the one or more first training images, and a first true object instance identifier for the object instance depicted in the one or more first training images (302).” – Col 4, Lines 59-64
NOTE: Abhiram further discloses the training image examples comprising a true object instance identifier. The instance identifier is defined as unique ID to differentiate between other images within the same class: “In certain embodiments, instead of a true class category vector, the process may use another data structure to label the object instance(s) in the training bundle such as a dictionary, e.g., {‘category’: ‘German shepherd dog’}. The first true object instance identifier may be a unique ID for the individual depicted, such as a social security number or a unique integer that has been previously assigned to the object instance”, see col 5, Lines 4-11. This functionally corresponds to the training examples comprising instance identifiers relating to the first image. After the combination obtaining the instance identifiers of training examples as taught by Abhiram can modify the training examples consisting of vector-quantized images taught by Maharana as combined so that the training examples can comprise instance identifiers related to the subject of the first vector-quantized image.
Regarding claim 20, the claim recites similar limitations to claim 3 or 4. Therefore, non-transitory computer program product claim 20 corresponds to the system disclosed in claim 3 or 4 and is rejected for the same reasons of obviousness as used above.
Allowable Subject Matter
Claims 6-8, and 14-17 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claim 14, Maharana in view of Chang, Zhou, and Navarrete teaches the system of claim 9. Maharana as combined does not teach wherein the one or more processors are further configured to: generate, using the prompt token generator, a second sequence of prompt tokens based at least in part on a second set of one or more identifiers; generate, using the prompt token generator, a third sequence of prompt tokens based at least in part on a third set of one or more identifiers, the third set of one or more identifiers differing from the second set of one or more identifiers by at least one identifier; generate one or more intermediate sequences of prompt tokens based on the second sequence of prompt tokens and the third sequence of prompt tokens; generate, using the pretrained generative image transformer, a second output token sequence based at least in part on the second sequence of prompt tokens, the second output token sequence representing a third vector-quantized image; and generate, using the pretrained generative image transformer, a third output token sequence based at least in part on the second output token sequence and one of the one or more intermediate sequences of prompt tokens, the third output token sequence representing a fourth vector-quantized image.
Although Maharana as combined teaches wherein the one or more processors are further configured to: generate, using the prompt token generator, a second sequence of prompt tokens based at least in part on a second set of one or more identifiers, see rejection of claim 13, the combination fails to establish generating a third sequence of prompt tokens based at least in part of a third set of one or more identifiers. The claim’s limitation expresses the need for both the second and third sequences of prompt tokens generated by the prompt token generator to exist in order to produce an intermediate sequence of prompt token sequences based on the second and third token sequences. The combination does not show generating this intermediate token sequence between two generated token sequences and therefore could not be used as input to the pretrained generative transformer to generate a third output sequence based on an intermediate sequence and the second output token sequence. None of the prior art searched, alone or in combination, renders obvious to the limitations of claim 14 and is therefore allowable if rewritten in independent form. Dependent claim 15-17 depend on claim 14 and therefore would also be allowable.
Regarding claim 6, the claim recites similar limitations to claim 14. Therefore, method claim 6 corresponds to the system disclosed in claim 14 and would be allowable for the same reasons as descried above. Dependent claims 7-8 depend on claim 6 and therefore would also be allowable.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID V. NGUYEN whose telephone number is (571)272-6111. The examiner can normally be reached M-F 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Y Poon can be reached at 571-270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DAVID VAN NGUYEN/
Examiner, Art Unit 2617
/KING Y POON/ Supervisory Patent Examiner, Art Unit 2617