CTFR 17/893,038 CTFR 98066 Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Claim Status 07-15-aia AIA Claim(s) 1, 4-7, 10, 11, 14-17, 20, 22 is/are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Huang (”Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning”, 2021), hereinafter referred to as Huang . Response to Amendment The amendment filed on 05/19/2026 has been entered. Claims 1, 11, 20 were amended Claims 2, 3, 8, 9, 12, 13, 18, 19, 21 was/were cancelled. Claims 1, 4-7, 10, 11, 14-17, 20, 22 remain pending in the application. Response to Arguments Applicant’s arguments (remarks filed 05/19/2026) have been considered but are not fully persuasive. Applicant argues on pg 6 (last two paragraphs) of the remarks filed 05/19/2026, reproduced below: PNG media_image1.png 448 860 media_image1.png Greyscale Upon further review of the reference and in light of applicant's argument, the examiner respectfully disagrees as follows: first of all, the amended claims state “textual representation”, which Huang discloses. See Figure 3, reproduced below: PNG media_image2.png 510 746 media_image2.png Greyscale . The VD, or Visual Dictionary, has the sample indices of 191, which contains the concept of “head”, as a textual representation. Additionally, Figure 1 shows the relationship between objects, reproduced below: PNG media_image3.png 498 734 media_image3.png Greyscale . See the “Ours” examples of the Huang framework. Applicant further argues on pg 7, reproduced below: PNG media_image4.png 776 864 media_image4.png Greyscale . Upon further review of the reference and in light of applicant's argument, the examiner respectfully disagrees as follows: first of all, see Examiner’s counter arguments above. Examiner notes that Applicant appears to inject specification limitations into the claims, but does not appear to specifically state the limitations. Applicant is reminded that although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns , 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). Additionally, Huang uses a “Transformer for vision” (Abstract). Accordingly, the claims, as they are currently written, do not place the instant application into an allowable state. Claim Rejections - 35 USC § 102 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 07-12-aia AIA (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 07-15-aia AIA Claim(s) 1, 4-7, 10, 11, 14-17, 20, 22 is/are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Huang (”Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning”, 2021), hereinafter referred to as Huang . Regarding claims 1, 11, and 20, Huang teaches A method comprising (Huang, abstract: “In this paper, we propose SOHO to “See Out of tHe bOx” that takes a whole image as input, and learns vision-language representation in an end-to-end manner”), at a device: processing an image (Huang, pg 3, column 1, Section 3, ¶1, “The visual encoder takes an image as input”, which is being interpreted as “processing an image”), by a vision transformer (Huang, see image below: “Transformer” and “visual” is being interpreted as including a “vision transformer”) pretrained (Huang, see image below: “we propose a novel Masked Visual Modeling pre-training” is being interpreted as involving “pretrained”) on a predefined (Huang, see image below, “based on the virtual visual semantic labels produced by the visual dictionary” is being interpreted as involving “predefined”, which is being interpreted as existing before; just as classification labels are predefined in classical machine learning algorithms) concept-feature dictionary (Huang, pg 4, column 2, Section 3.3, ¶1, reproduced below: PNG media_image5.png 340 550 media_image5.png Greyscale . “Visual dictionary” is being interpreted as involving a concept-feature dictionary as seen from the “image-text matching pre-training tasks”) that correlates image features with image concepts (Huang, pg 8, Section 4.4, reproduced below: PNG media_image6.png 338 552 media_image6.png Greyscale . “VD index” is being interpreted as image concept [as can be seen in Figure 3 where index 191 is the image concept “head” and index 1074 is the image concept “building”], “visual feature” is being interpreted as “image features”) that indicate, via a textual representation (Huang, pg 8, Figure 3, reproduced below: PNG media_image2.png 510 746 media_image2.png Greyscale . The index 191 and “head” are being interpreted to involve a textual representation), relationships between objects represented in the image features to infer an associated concept for the image (Huang, Figure 1 and Figure 1 text shows “Ours: A couple sit on the shore next to a boat on the sea”, which is being interpreted as involving inferring “an associated concept for the image” that shows the relationship between objects. For example, “A couple sit on the shore” shows the relationship between “A couple” and “the shore”) that indicates a relationship (Huang, Figure 1, “Ours: A couple sit on the shore next to a boat on the sea”. “sit on the shore”, “next to a boat” are being interpreted as indicating “a relationship”) between two or more objects as depicted in the image (Huang, Figure 1, “Ours: A couple sit on the shore next to a boat on the sea”. “Couple, shore, boat, sea” are being interpreted as “two or more objects depicted in the image” ), wherein the vision transformer is comprised of a tokenizer, at least one layer for generating patch embeddings (Huang, see Section 4.4 image above, “image patch”; when combined with pg 3, Figure 2 of “Visual Dictionary-based embedding features”, shows “at least one layer for generating patch embeddings”), and at least one multi-head self-attention layer (Huang, pg 9, Section A.4. Discussion, ¶1, reproduced below: PNG media_image7.png 572 558 media_image7.png Greyscale “Multi-layer Transformer” is being interpreted as involving at least one multi-head. “Self-attention mechanism” is being interpreted as being part of one of the layer, resulting in a “multi-head self-attention layer”. Further, a ResNet-101 backbone and 12-layer Transformer is used in this prior art, seen in pg 5, Section 4.1. As with ordinary skill in the art would know, this contains at least one multi-head self attention layer”, see note * below); and outputting the concept inferred for the image (Huang, Figure 1 and Figure 1 text shows “Ours: A couple sit on the shore next to a boat on the sea”. Which is being interpreted as outputting the concept inferred for the image). * Note : Knowledge reference on the 12-layer transformer (pg 3, column 2 mentions “self-attention heads”, which is being interpreted as multi-head self-attention): https://arxiv.org/abs/1810.04805. Regarding claim 4, Huang teaches The method of claim 1, wherein the image is not labeled with the concept when received as input (Huang, pg 5, column 1, last paragraph: “Detailed comparisons of pre-training dataset usage of most VLPT works, including our train/test image and text numbers, are included in our supplementary material”. “Test image” is being interpreted as splitting the dataset into training and testing sets, as one with ordinary skill in the art would know. The testing sets are being interpreted as not labeled with the concept when received as input). Regarding claim 5, Huang teaches The method of claim 1, wherein the vision transformer performs action prediction and object prediction utilizing the image (Huang, Figure 1 and Figure 1 text shows “Ours: A couple sit on the shore next to a boat on the sea”. “Sit” is being interpreted as “action prediction”. “Couple”, “shore”, “boat”, “sea” are being interpreted as object prediction. Figure 1 shows an image which these predictions are done. SOHO, the method, is being interpreted as involving a “vision transformer”). Regarding claim 6, Huang teaches The method of claim 1, wherein the concept (Huang, see Figure 1 citation below, the caption is being interpreted as involving “the concept”) indicates the relationship (Huang, see Figure 1 citation below, “A couple sit on the shore next to” is being interpreted as involving a relationship) between the two or more objects (Huang, see Figure 1 citation below, “couple” and “boat” are being interpreted as “two or more objects”) as a tuple (Huang, see Figure 1 citation below. A definition of a tuple, as one with ordinary skill in the art would know, if an ordered, finite sequence of elements. The order of the caption matters and it is a finite sequence of elements.) of two objects and an associated action (Huang, Figure 1 and Figure 1 text shows “Ours: A couple sit on the shore next to a boat on the sea”. “Couple” and “boat” are being interpreted as examples of two objects with the association action of “sit”). Regarding claim 7, Huang teaches The method of claim 1, further comprising, at the device: performing one or more classification operations utilizing the concept (Huang, pg 7, column 2, Section 4.2.4, ¶1, reproduced below: PNG media_image8.png 456 548 media_image8.png Greyscale . “Three-classification” is being interpreted as one or more classification operations. “Output of the transformer” is being interpreted as “utilizing the concept”). Regarding claim 10, Huang teaches The method of claim 1, wherein each of a plurality of concepts within the dictionary is represented by a key (Huang, see image below, the “VD index”, or Visual Dictionary Index, is being interpreted as a key), and wherein each of the plurality of concepts within the dictionary is linked to a predefined (Huang, see section 3.3 image from claim 1, “based on the virtual visual semantic labels produced by the visual dictionary” is being interpreted as involving “predefined”, which is being interpreted as existing before; just as classification labels are predefined in classical machine learning algorithms) set of image features (Huang, pg 8, Section 4.4, reproduced below: PNG media_image6.png 338 552 media_image6.png Greyscale . “VD index” is being interpreted as image concept [as can be seen in Figure 3 where index 191 is the image concept “head” and index 1074 is the image concept “building”], “visual feature” is being interpreted as “image features”). Regarding claim 22, Huang teaches The method of claim 1, wherein the vision transformer is configured to include: a global task that, during training, clusters images with the same concept together to produce semantically consistent relational representations (Huang, Figure 1, reproduced below: PNG media_image9.png 744 1102 media_image9.png Greyscale . “Global context” is being interpreted as “global task”. For example, “chatting” or “next to a boat” are examples of relational representations”), and a local task that, during training, guides the vision transformer to discover object-centric semantic correspondence across images (Huang, pg 8, Section 4.4, reproduced below: PNG media_image6.png 338 552 media_image6.png Greyscale . “VD index” is being interpreted as object-centric semantic correspondence [as can be seen in Figure 3 where index 191 is the object-centric semantic correspondence “head” and index 1074 is the object-centric semantic correspondence “building” across images]. This object-centric semantic correspondence is being interpreted as “a local task”; pg 2, column 1, second to last paragraph: “VD can be dynamically updated through our trainable CNN backbone directly from visual-language data during pretraining” The Visual Dictionary is being interpreted as during training guides the vision transformer). Claim 14 is rejected using the same rationale as applied to claim 4 discussed above. Claim 15 is rejected using the same rationale as applied to claim 5 discussed above. Claim 16 is rejected using the same rationale as applied to claim 6 discussed above. Claim 17 is rejected using the same rationale as applied to claim 7 discussed above . Conclusion 07-39 AIA THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure : Kim et al (“Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)”, 2018) discloses a dictionary of concepts (Appendix A) and trained using concept vectors. Liu et al (“SIFT Flow: Dense Correspondence across Scenes and its Applications”, 2010, as cited in IDS filed 09/16/2022) discloses a concept-feature dictionary (interpreted from visual words and SIFT descriptors, pg 6, column 1, first full paragraph). Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHNNY B DUONG whose telephone number is (571)272-1358. The examiner can normally be reached Monday - Thursday 10a-9p (ET). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Bella can be reached at (571)272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.B.D./Examiner, Art Unit 2667 /MATTHEW C BELLA/Supervisory Patent Examiner, Art Unit 2667 Application/Control Number: 17/893,038 Page 2 Art Unit: 2667 Application/Control Number: 17/893,038 Page 3 Art Unit: 2667 Application/Control Number: 17/893,038 Page 4 Art Unit: 2667 Application/Control Number: 17/893,038 Page 5 Art Unit: 2667 Application/Control Number: 17/893,038 Page 6 Art Unit: 2667 Application/Control Number: 17/893,038 Page 7 Art Unit: 2667 Application/Control Number: 17/893,038 Page 8 Art Unit: 2667 Application/Control Number: 17/893,038 Page 9 Art Unit: 2667 Application/Control Number: 17/893,038 Page 10 Art Unit: 2667