Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Claims 1 – 15 and 17 – 21 are pending in this application. Claims 1, 17 and 18 are independent.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1 – 13, 15 and 17 – 21 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by NPL (Khan et al. | Transformers in Vision: A Survey | (2021-01-04) Last Update Date: 2022-01-19).
Regarding independent claims 1, 17, and 18, Khan et al. teaches:
A system (e.g., FIGS. 7, 8 of Khan et al.) comprising: one or more computers (e.g., GPU of Khan et al.); and one or more storage devices (e.g., memory of Khan et al.) storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising: obtaining an input image for a computer vision task (See at least Khan et al., page 4, #2.3 Transformer Model; FIGS. 2, 3, 7, 8; "…Being an auto-regressive model, the decoder of the Transformer uses previous predictions…takes inputs from the encoder as well as the previous outputs to predict the next word of the sentence in the translated language…" Also, see pages 2, #2.1 Self-Attention in Transformers & 3, #2.2 (Self) Supervised Pre-training); processing the input image through a sequence transduction neural network (e.g., encoder + auto-regressive decoder neural network of Khan et al.) that is configured to process the input image to generate a guiding code sequence (e.g., tokens of Khan et al.) that includes a fixed number of vectors (See at least Khan et al., page 4, #2.3 Transformer Model; FIGS. 2, 3, 7, 8; "…Being an auto-regressive model, the decoder of the Transformer uses previous predictions…takes inputs from the encoder as well as the previous outputs to predict the next word of the sentence in the translated language…" Also, see at least pages 2, #2.1 Self-Attention in Transformers & 3, #2.2 (Self) Supervised Pre-training); and providing the guiding code sequence as input to a base computer vision neural network (e.g., feed-forward layers/Vision Transformer (FIG. 6) of Khan et al.) that is configured to process the guiding code sequence and the input image to generate a network output for the computer vision task (e.g., the task of predicting the next word of the sentence of Khan et al.) (See at least Khan et al., page 3, FIG. 3; "…The Transformer encoder (middle row) operates on the input language sequence and converts it to an embedding before passing it on to the encoder blocks…The blocks consisting of multi-head attention (top row) and feed-forward layers are repeated N times in both the encoder and decoder…" Also, see at least pages 4, #2.3 Transformer Model; 5, #3.1.1 Self-Attention in CNNs; 6, #3.2.1 Uniform-scale Vision Transformers and FIGS. 2, 6 – 8).
Regarding dependent claims 2 and 19, Khan et al. teaches:
wherein each vector in the guiding code sequence is selected from a discrete vocabulary of vectors (See at least Khan et al., page 7, FIG. 6; "…In standard ViTs, the number of the tokens and token feature dimension are kept fixed throughout…" Also, see at least pages 4, #2.3 Transformer Model; 5, #3.1.1 Self-Attention in CNNs; 6, #3.2.1 Uniform-scale Vision Transformers and FIGS. 2, 6 – 8).
Regarding dependent claims 3 and 20, Khan et al. teaches:
wherein the sequence transduction neural network (e.g., encoder + auto-regressive decoder neural network of Khan et al.) comprises: an encoder neural network configured to process the input image to generate an encoded representation of the input image, and an auto-regressive decoder neural network configured to auto-regressively generate an output sequence that specifies the guiding code sequence conditioned on the encoded representation of the input image (See at least Khan et al., page 4, #2.3 Transformer Model; FIGS. 2, 3, 7, 8; "…Being an auto-regressive model, the decoder of the Transformer uses previous predictions…takes inputs from the encoder as well as the previous outputs to predict the next word of the sentence in the translated language…" Also, see at least pages 2, #2.1 Self-Attention in Transformers & 3, #2.2 (Self) Supervised Pre-training).
Regarding dependent claims 4 and 21, Khan et al. teaches:
wherein: the encoder neural network is a Vision Transformer backbone neural network (e.g., Vision Transformer (FIG. 6) of Khan et al.); and the auto-regressive decoder neural network is an auto-regressive Transformer decoder (e.g., auto-regressive model of Khan et al.) (See at least Khan et al., page 4, #2.3 Transformer Model; FIGS. 2, 3, 7, 8; "…Being an auto-regressive model, the decoder of the Transformer uses previous predictions…takes inputs from the encoder as well as the previous outputs to predict the next word of the sentence in the translated language…" Also, see at least pages 2, #2.1 Self-Attention in Transformers & 3, #2.2 (Self) Supervised Pre-training).
Regarding dependent claim 5, Khan et al. teaches:
wherein the base computer vision neural network is a feedforward neural network (e.g., feed-forward layered network of Khan et al.).
Regarding dependent claim 6, Khan et al. teaches:
wherein the base computer vision neural network is a Vision Transformer (e.g., Vision Transformer (FIG. 6) of Khan et al.).
Regarding dependent claim 7, Khan et al. teaches:
wherein the guiding code sequence is a prediction of a sequence that would be generated by a restricted oracle neural network by processing a ground truth label for the computer vision task for the input image (See at least Khan et al., page 9, FIG. 7; "…Detection Transformer (DETR) [13] treats the object detection task as a set prediction problem and uses the Transformer network to encode relationships between set elements. A bipartite set loss is used to uniquely match the box predictions with the ground-truth boxes (shown on the right two columns)…" Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model).
Regarding dependent claim 8, Khan et al. teaches:
wherein generating the guiding code sequence requires generating more than one hundred times fewer values than generating the network output for the computer vision task (See at least Khan et al., page 9; Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model and FIGS. 2, 3, 6 – 8).
Regarding dependent claim 9, Khan et al. teaches:
wherein the network output for the computer vision task is structured output that includes one or more predicted values for each of a plurality of pixels in the output image (See at least Khan et al., page 9; Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model and FIGS. 2, 3, 6 – 8).
Regarding dependent claim 10, Khan et al. teaches:
wherein the computer vision task is object detection (See at least Khan et al., page 8, #3.3 Transformers for Object Detection; FIG. 7; Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model).
Regarding dependent claim 11, Khan et al. teaches:
wherein the base computer vision neural network and the sequence transduction neural network have been trained by performing training operations comprising: obtaining a set of first training data for the computer vision task that comprises a plurality of training images and, for each training image, a ground truth output for the computer vision task (See at least Khan et al., page 9, FIG. 7; "…Detection Transformer (DETR) [13] treats the object detection task as a set prediction problem and uses the Transformer network to encode relationships between set elements. A bipartite set loss is used to uniquely match the box predictions with the ground-truth boxes (shown on the right two columns)…" Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model); training the base computer vision neural network (e.g., encoder + auto-regressive decoder neural network of Khan et al.) jointly with a restricted oracle neural network (e.g., Vision Transformer (FIG. 6) of Khan et al.) on the first training data (See at least Khan et al., page 8, #3.2.4 Self-Supervised Vision Transformers & #3.3 Transformers for Object Detection; FIG. 7; Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model), wherein the restricted oracle neural network is configured to process a ground truth output for the computer vision task to generate a training guiding code sequence for the corresponding training image (See at least Khan et al., page 9, FIG. 7; "…Detection Transformer (DETR) [13] treats the object detection task as a set prediction problem and uses the Transformer network to encode relationships between set elements. A bipartite set loss is used to uniquely match the box predictions with the ground-truth boxes (shown on the right two columns)…" Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model), and wherein, during the training, the computer vision neural network receives as input (i) a training image and (ii) a training guiding code sequence (e.g., sequence of tokens of Khan et al.) for the training image generated by the restricted oracle neural network (See at least Khan et al., page 9, FIG. 7; Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model); and after training the base computer vision neural network jointly with the restricted oracle neural network, training the sequence transduction neural network (e.g., encoder + auto-regressive decoder neural network of Khan et al.) on second training data that includes a plurality of training examples, each training example including: (i) a training image, and (ii) a ground truth guiding code sequence generated by processing a ground truth output for the computer vision task (e.g., sequence-to-sequence prediction of Khan et al.) for the training image using the trained restricted oracle neural network (See at least Khan et al., page 9, FIG. 7; Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model and FIGS. 2, 3, 8, 9;).
Regarding dependent claim 12, Khan et al. teaches:
wherein each vector in the guiding code sequence is selected from a discrete vocabulary of vectors (See at least Khan et al., page 2, FIG. 7; Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model and FIGS. 2, 3, 8, 9;), and wherein training the base computer vision neural network jointly with the restricted oracle neural network comprises: learning the discrete vocabulary of vectors (See at least Khan et al., page 2, FIG. 7; Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model and FIGS. 2, 3, 8, 9;).
Regarding dependent claim 13, Khan et al. teaches:
wherein the restricted oracle neural network is configured to map a ground truth output for the computer vision task to a sequence of encoded vectors (See at least Khan et al., page 13, "…Transformers at full potential train large-scale data, [19] take the clean (ground-truth) images from the ImageNet benchmark and synthesize their degraded versions for different tasks…" Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training & page 4, #2.3 Transformer Model), and generate the training guiding code sequence by mapping each encoded vector to a nearest vector in the discrete vocabulary of vectors (See at least Khan et al., page 13, "…Transformers at full potential train large-scale data, [19] take the clean (ground-truth) images from the ImageNet benchmark and synthesize their degraded versions for different tasks…" Also, see at least pages 2, #2.1 Self-Attention in Transformers; 3, #2.2 (Self) Supervised Pre-training, page 4, #2.3 Transformer Model, and page 7 Token to Token ViT).
Allowable Subject Matter
Dependent claim 14 is objected to as being allowable – including all of the limitations of its base claim(s) and any intervening and/or dependent claims, if re-written in independent form.
Conclusion
The prior art made of record and not relied upon is considered pertinent to Applicant's disclosure: See the Notice of References Cited (PTO–892)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IDOWU O OSIFADE whose telephone number is (571)272-0864. The Examiner can normally be reached on Monday-Friday 8:00am-5:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the Examiner’s Supervisor, JOHN VILLECCO can be reached on (571) 272 – 7319. The fax phone number for the organization where this application or proceeding is assigned is (571) 273 – 8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov.
Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at (866) 217 – 9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call (800) 786 – 9199 (IN USA OR CANADA) or (571) 272 – 1000.
/IDOWU O OSIFADE/Primary Examiner, Art Unit 2675