Prosecution Insights
Last updated: August 16, 2026
Application No. 18/497,079

VISUAL QUESTION ANSWERING WITH UNLABELED IMAGE AUGMENTATION

Non-Final OA §101§103
Filed
Oct 30, 2023
Priority
Nov 04, 2022 — provisional 63/422,629 +1 more
Examiner
BALDWIN, RANDALL KERN
Art Unit
Tech Center
Assignee
NEC Laboratories America Inc.
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
192 granted / 242 resolved
+19.3% vs TC avg
Strong +28% interview lift
Without
With
+28.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
10 currently pending
Career history
259
Total Applications
across all art units

Statute-Specific Performance

§101
16.3%
-23.7% vs TC avg
§103
41.4%
+1.4% vs TC avg
§102
13.3%
-26.7% vs TC avg
§112
24.3%
-15.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 242 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to the application and claims filed 10/30/2023. Claims 1-20 are pending and have been examined. Claims 1-20 are rejected. Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. The present application claims priority to U.S. Provisional Application No. 63/422,629 filed on 11/04/2022 and to U.S. Provisional Application No. 63/423,945 filed on 11/09/2022. Applicant has not complied with one or more conditions for receiving the benefit of an earlier filing date under 35 U.S.C. 119(e) as follows: The later-filed application must be an application for a patent for an invention which is also disclosed in the prior application (the parent or original nonprovisional application or provisional application). The disclosure of the invention in the parent application and in the later-filed application must be sufficient to comply with the requirements of 35 U.S.C. 112(a) or the first paragraph of pre-AIA 35 U.S.C. 112, except for the best mode requirement. See Transco Products, Inc. v. Performance Contracting, Inc., 38 F.3d 551, 32 USPQ2d 1077 (Fed. Cir. 1994). The disclosure of the prior-filed application, U.S. Provisional Application No. 63/422,629 (hereinafter “the ‘629 provisional application”) fails to provide adequate support or enablement in the manner provided by 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph for one or more claims of this application. Dependent claims 2, 9 and 17 recite, using respective similar language “wherein training the student model includes training the student model to approximate P(T | I), where T = (Q, A), Q is a question, A is an answer and P(T | I) is conditional probability of T in image I.” The as-filed specification of the ‘629 provisional application fails to provide adequate support or enablement for at least this element of claims 2, 9 and 17. That is, the as-filed specification of the ‘629 provisional application fails to provide adequate support or enablement for at least the above-noted “wherein training the student model includes training the student model to approximate P(T | I), where T = (Q, A), Q is a question, A is an answer and P(T | I) is conditional probability of T in image I.” elements of claims 2, 9 and 17. For example, the ‘629 provisional is silent regarding any training of the student model that includes or comprises training the model to approximate P(T | I) where P(T | I) is conditional probability of T in image I, let alone the claimed “wherein training the student model includes training the student model to approximate P(T | I), where T = (Q, A), Q is a question, A is an answer and P(T | I) is conditional probability of T in image I.” as recited, using respective similar language, in claims 2, 9 and 17. Based on their respective dependencies from claims 2 and 9, the specification of the ‘629 provisional application also fails to provide adequate support or enablement for dependent claims 3 and 10. Dependent claims 4, 11 and 18 recite, using respective similar language “wherein training the teacher model includes learning through deep learning to maximize a conditional likelihood of a question-answer pair, given an image.” The as-filed specification of the ‘629 provisional application fails to provide adequate support or enablement for at least this element of claims 4, 11 and 18. That is, the as-filed specification of the ‘629 provisional application fails to provide adequate support or enablement for at least the above-noted “wherein training the teacher model includes learning through deep learning to maximize a conditional likelihood of a question-answer pair, given an image.” elements of claims 4, 11 and 18. For example, the ‘629 provisional is silent regarding any training of a teacher model that includes or comprises any deep learning or training to maximize a conditional likelihood of a question-answer pair, much less the claimed “wherein training the teacher model includes learning through deep learning to maximize a conditional likelihood of a question-answer pair, given an image.” as recited, using respective similar language, in claims 4, 11 and 18. Dependent claims 5 and 12 recite, using respective similar language “wherein pseudolabeling unlabeled images includes producing pseudolabels for unlabeled images Iu by obtaining logits of a decoder, and the logits defining a distribution over tokens of a natural language vocabulary of the teacher model.” The as-filed specification of the ‘629 provisional application fails to provide adequate support or enablement for at least this element of claims 5 and 12. That is, the as-filed specification of the ‘629 provisional application fails to provide adequate support or enablement for at least the above-noted “wherein pseudolabeling unlabeled images includes producing pseudolabels for unlabeled images Iu by obtaining logits of a decoder, and the logits defining a distribution over tokens of a natural language vocabulary of the teacher model.” elements of claims 5 and 12. For example, the ‘629 provisional is silent regarding any decoder, obtaining any logits of a decoder, tokens, or producing pseudolabels for unlabeled images Iu by obtaining logits of a decoder, let alone the claimed “wherein pseudolabeling unlabeled images includes producing pseudolabels for unlabeled images Iu by obtaining logits of a decoder, and the logits defining a distribution over tokens of a natural language vocabulary of the teacher model.” as recited, using respective similar language, in claims 5 and 12. Dependent claims 6, 13 and 19 each recite “wherein, given an image I, a question Q and answer A, the student model approximates P(A| Q, I), while the teacher model approximates P(Q, A | I), where P(A| Q, I) is conditional probability of A in Q, I and P(A, Q| I) is conditional probability of A and Q in I to enable using unlabeled data in training.” The as-filed specification of the ‘629 provisional application fails to provide adequate support or enablement for at least this element of claims 6, 13 and 19. That is, the as-filed specification of the ‘629 provisional application fails to provide adequate support or enablement for at least the above-noted “wherein, given an image I, a question Q and answer A, the student model approximates P(A| Q, I), while the teacher model approximates P(Q, A | I), where P(A| Q, I) is conditional probability of A in Q, I and P(A, Q| I) is conditional probability of A and Q in I to enable using unlabeled data in training.” elements of claims 6, 13 and 19. For example, the ‘629 provisional is silent regarding any student model that approximates P(A| Q, I) and teacher model that approximates P(Q, A | I) where P(A| Q, I) is conditional probability of A in Q, I and P(A, Q| I) is conditional probability of A and Q in I, let alone the claimed “wherein, given an image I, a question Q and answer A, the student model approximates P(A| Q, I), while the teacher model approximates P(Q, A | I), where P(A| Q, I) is conditional probability of A in Q, I and P(A, Q| I) is conditional probability of A and Q in I to enable using unlabeled data in training.” as recited in claims 6, 13 and 19. Thus, the as-filed specification of the ‘629 provisional application fails to provide adequate support or enablement for at least the above-noted elements of claims 2-6, 9-13 and 17-19. However, the disclosure of the prior-filed application, U.S. Provisional Application No. 63/423,945 filed 11/09/2022 (hereinafter “the ‘945 provisional application”) appears to provide adequate support for claims 2-6, 9-13 and 17-19. Therefore, the effective filing date for claims 2-6, 9-13 and 17-19 of the instant application is the filing date of the ‘945 provisional application, 11/09/2022. Examiner will consider if the ‘629 provisional application supports each of the other claims if a rejection would need to rely upon an intervening reference between the actual filing date of the ‘945 provisional application, 11/09/2022, and the 11/04/2022 filing of the ‘629 provisional application. Each claim will receive benefit of the earliest filing date above for which a continuous chain of support can be established for the entirety of the claim. As discussed above, the effective filing date for at least claims 2-6, 9-13 and 17-19 of the instant application is the filing date of the ‘945 provisional application, 11/09/2022. Information Disclosure Statement Acknowledgment is made of the information disclosure statement filed 10/30/2023, which complies with 37 CFR 1.97. As such, the information disclosure statement has been placed in the application file and the information referred to therein has been considered by the examiner. Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference characters not mentioned in the description: Reference characters 204 and 208 in FIG. 2 are not found in the detailed description (see, paragraph 32 describing FIG. 2); and Reference characters 420 and 422 in FIG. 4 are not found in the detailed description (see, paragraphs 46-52 describing FIG. 4); The drawings are also objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference signs mentioned in the description: 700, 720 and 722 (see, paragraphs 48-49 describing FIG. 4 and reciting “A VGA model 720 can be stored in memory 703 along with program code 722 for generating a user interface and responding to queries with visual and textual information.” And “The processing system 700 may also include other elements (not shown), for example, various other input devices and/or output devices can be included in processing system 700, depending upon the particular implementation.”). The drawings are further objected to as failing to comply with 37 CFR 1.84(p)(3) because Figures 1-3 and 5 include letters which do not measure at least .32 cm. (1/8 inch) in height (i.e., most lowercase and subscript characters in FIGs. 1, 3 and 5, and some of the lowercase characters in element 230 in FIG. 2). Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference characters in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Specification The disclosure is objected to because of the following informalities: In the “RELATED APPLICATION INFORMATION” section in paragraph 1 of the specification the recitation of “This application claims priority to U.S. Provisional Application No. 63/422,629, filed on November 4, 2022, and U.S. Provisional Application No. 63/423,945, filed on November 9, 2022, both incorporated herein by reference in their entirety.” is grammatically incorrect. In particular, “both incorporated herein by reference in their entirety” should recite “both incorporated herein by reference in their entireties Appropriate correction is required. Reference characters 204 and 208 shown in Figure 2 are not found in the detailed description (see, paragraph 32 describing FIG. 2), and reference characters 420 and 422 in FIG. 4 are not found in the detailed description (see, paragraphs 46-52 describing FIG. 4). Appropriate correction is required. It appears that references to reference characters 700, 720 and 722 in the specification are typographical errors and should recite 400, 420 and 422 (see, FIG. 4 depicting elements 400, 420 and 422, and paragraphs 48-49 describing FIG. 4 and reciting “A VGA model 720 can be stored in memory 703 along with program code 722 for generating a user interface and responding to queries with visual and textual information.” and “The processing system 700 may also include other elements (not shown), for example, various other input devices and/or output devices can be included in processing system 700, depending upon the particular implementation.”). Appropriate correction is required. The last sentence of paragraph 37, the recitation of “For a sample (Q, A, I) ∈ DQA, the sample is transformed into a target sequence of tokens T (yl , y2 , … yn) by entering (Q, A) into a structured template of having the following form:” is grammatically incorrect and includes an extraneous word, “of”, between “template” and “having”. Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. because the claimed invention is directed to an abstract idea without significantly more. The analysis below of the claims’ subject matter eligibility follows the 2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50-57 (January 7, 2019) (“2019 PEG”) and the 2024 Guidance Update on Patent Subject Matter Eligibility, Including on Artificial Intelligence, 89 Fed. Reg. 58128-58138 (July 17, 2024) (“2024 AI SME Update”). When considering subject matter eligibility under 35 U.S.C. 101, it must be determined whether the claim is directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter (Step 1). If the claim does fall within one of the statutory categories, the second step in the analysis is to determine whether the claim is directed to a judicial exception (Step 2A). The Step 2A analysis is broken into two prongs. In the first prong (Step 2A, Prong 1), it is determined whether or not the claims recite a judicial exception (e.g., mathematical concepts, mental processes, certain methods of organizing human activity). If it is determined in Step 2A, Prong 1 that the claims recite a judicial exception, the analysis proceeds to the second prong (Step 2A, Prong 2), where it is determined whether or not the claims integrate the judicial exception into a practical application. If it is determined at step 2A, Prong 2 that the claims do not integrate the judicial exception into a practical application, the analysis proceeds to determining whether the claim is a patent-eligible application of the exception (Step 2B). If an abstract idea is present in the claim, any element or combination of elements in the claim must be sufficient to ensure that the claim integrates the judicial exception into a practical application, or else amounts to significantly more than the abstract idea itself. Regarding claims 1, 8 and 16, these claims are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 1 is directed to a method, corresponding to a process, claim 8 is directed to a system, corresponding to a machine and claim 16 is directed to a computer program product embodied in a non-transitory computer readable medium, corresponding to an article of manufacture, which are each one of the statutory categories. The claims are directed to an abstract idea. In particular, the claims recite mental processes that are concepts performed in the human mind (including an observation, evaluation, judgment, opinion). The limitations recited in claims 1, 8 and 16, using respective similar language: performing image conditional visual question generation on a visual language model (VLM) and a targeted visual question answer dataset using images to generate question and answer pairs; pseudolabeling unlabeled images using the teacher model to decode synthetic question and answer pairs for the unlabeled images; merging the synthetic question and answer pairs for the unlabeled images with real data from the targeted visual question answer dataset to generate a self-augmented training set as drafted, under their broadest reasonable interpretation (BRI), in view of the specification, cover concepts performed in the human mind (evaluation, judgement, or opinion to generate question-answer pairs based on received/observed data from a targeted visual question answer dataset using a generically-recited visual language model (VLM), labeling observed/received unlabeled images using a generically-recited teacher model, decoding synthetic/artificial question and answer pairs for the unlabeled images and then merging/combining the synthetic question-answer pairs for the observed/received unlabeled images with observed, real data from the targeted visual question-answer dataset in order to generate/create a self-augmented training data set. Regarding the “visual language model (VLM)” and the “teacher model”, no details of the models or their training are recited, and the models are recited at a high level of generality and the models can be constructed by hand with pen and paper. Thus, the claimed VLM and “teacher model”, under the BRI, in light of the specification, could be any models useable to generate question-answer pairs based on received/observed data from a targeted visual question answer dataset, which could be constructed by hand with pen and paper based on a reasonable amount of observed data (i.e., the “visual question answer dataset” and images). Given a sufficiently small set of images and questions and answers in the “visual question answer dataset”, nothing in the claims prohibit the generating, pseudolabeling and merging steps/operations from being performed mentally or with pen and paper. The VLM and teacher model are recited at a high level of generality and therefore are being interpreted as performing a mental process on a generic computer. See MPEP 2106.04(a)(2) § III.C which states that “a concept that is performed in the human mind and applicant is merely claiming that concept performed 1) on a generic computer, or 2) in a computer environment, or 3) is merely using a computer as a tool to perform the concept” still recite a mental process. The claim limitations, under their broadest reasonable interpretations (BRIs), cover performance of the limitations in the mind but for the recitation of generic computer components (e.g., “A computer-implemented method for training a visual question answer model, comprising:” <the above-noted steps> (claim 1), “A system for training a visual question answer model, comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:” <perform the above-noted operations> (claim 8) and “A computer program product for training a visual question answer model, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:” <the above-noted steps> (claim 16). Therefore, the claims are directed to an abstract idea - mental processes. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claims recite the additional elements: “A computer-implemented method for training a visual question answer model, comprising:” <the above-noted steps> (claim 1), “A system for training a visual question answer model, comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:” <perform the above-noted operations> (claim 8) and “A computer program product for training a visual question answer model, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:” <the above-noted steps> (claim 16). These additional elements are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that they amount to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). The claims also recite, using respective similar language, the additional elements: “training a teacher model by performing image conditional visual question generation on a visual language model (VLM) and a targeted visual question answer dataset using images to generate question and answer pairs; … and training a student model using the VLM and the self-augmented training set to return visual answers to text queries.” The “visual language model (VLM)”, “teacher model” and “student model” are recited at a high level of generality as mere instructions to implement an abstract idea on a computer and amounts to the recitation of the words “apply it” (or an equivalent) or amount to no more than mere instructions to implement an abstract idea or other exception on a computer or merely uses a computer as a tool to perform an abstract idea (i.e., as generic computer components performing generic computer functions). See MPEP 2106.05(f). The above-noted training steps/operations are simply generic training to perform the abstract idea of processing data and amount to mere instructions to apply the exception (MPEP 2106.05(f)). That is, the training is invoked as a tool to perform the underlying abstract idea of generating question data, labeling image data, and merging question-answer data (i.e., data processing). Machine learning/training so recited cannot provide an improvement to the functioning of a computer or other technological field. MPEP § 2106.05(a)(I). Step 2B Analysis: The claims do not recite additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of “A computer-implemented method for training a visual question answer model, comprising:” <the above-noted steps> (claim 1), “A system for training a visual question answer model, comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:” <perform the above-noted operations> (claim 8) and “A computer program product for training a visual question answer model, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:” <the above-noted steps> (claim 16) amount to no more than using generic computer components to implement the exception such that they amount to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). As also discussed above, applying the generically-recited models and generic training to perform the abstract idea amounts to no more than mere instructions to apply the exception (MPEP 2106.05(f). Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The additional elements do not provide an inventive concept, and, therefore, the claims are not patent eligible. Regarding claims 2, 9 and 17, these claims are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 2 is directed to a method as depending from claim 1, claim 9 is directed to a system as depending from claim 8 and claim 17 is directed to a computer program product embodied in a non-transitory computer readable medium as depending from claim 16, thus the analysis for patent eligibilities of claims 1, 8 and 16 are incorporated herein. Step 2A Prong 1: The claims each recite “approximate P(T | I), where T = (Q, A), Q is a question, A is an answer and P(T | I) is conditional probability of T in image I.” The additional limitation added by these claims covers mathematical concepts - mathematical relationships, mathematical formulas or equations, or mathematical calculations (e.g., approximating a conditional probability P(T | I) based on Q, A – question-answer data). Therefore, the claims are directed to an abstract idea - mental processes combined with mathematical concepts. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claims recite, using respective similar language, the additional element: “wherein training the student model includes training the student model to approximate P(T | I).” The above-noted training step/operation is simply generic training to perform the abstract idea of calculating/approximating the conditional probability P(T | I) and amounts to mere instructions to apply the exception (MPEP 2106.05(f)). That is, the training is invoked as a tool to perform the underlying abstract idea of approximating a conditional probability P(T | I) (i.e., a mathematical concept). Machine learning/training so recited cannot provide an improvement to the functioning of a computer or other technological field. MPEP § 2106.05(a)(I). Claim 9 also recites the additional element “wherein the computer program causes the hardware processor to train the student model to approximate P(T | I)”. This additional element is recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computer) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Step 2B Analysis: The claims do not recite additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, applying the generically-recited student model and generic training to perform the abstract idea (mathematical concept) amounts to no more than mere instructions to apply the exception (MPEP 2106.05(f). As also discussed above, the additional element “wherein the computer program causes the hardware processor to train the student model to approximate P(T | I)” in claim 9 amounts to no more than using generic computer components to implement the exception such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The additional elements do not provide an inventive concept, and, therefore, the claims are not patent eligible. Regarding claims 3 and 10, these claims are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 3 is directed to a method as depending from claim 2 and claim 10 is directed to a system as depending from claim 9, thus the analysis for patent eligibilities of claims 2 and 9, and of base claims 1 and 8 are incorporated herein. Step 2A Prong 1: The claims recite, using respective similar language, “wherein the targeted visual question answer dataset is generated by transforming data samples into a target sequence of tokens T (yl , y2 , … yn) by entering (Q, A) into a structured template, where Q is a question, A is an answer; and optimizing a loss over all question-image-answer pairs.” The additional limitations added by these claims cover mathematical concepts - mathematical relationships, mathematical formulas or equations, or mathematical calculations (e.g., transforming data based on observed data samples into a target sequence of tokens T and optimizing a loss/loss function over all question-image-answer pairs). In light of the specification, (see, paragraphs 33, 37 and 39-40, “The teacher model 308 is image conditioned to associate question answer pairs with images (I) by optimizing a loss function (LVQG).”, “For a sample (Q, A, I) ∈ DQA, the sample is transformed into a target sequence of tokens T (yl , y2 , … yn) by entering (Q, A) into a structured template”, and “Once T (yl , y2 , … yn) is obtained, the teacher model (VQG) 308 is trained by optimizing: L V Q G =   - ∑ n = 1 N log ⁡ P θ y n y < n ,   x ) (2) over all of the question-image-answer pairs in DQA, where x represents the latent encoded features in the standard encoder-decoder architecture.”), these limitations are directed to a mathematical concepts. Therefore, the claims are directed to an abstract idea - mental processes combined with mathematical concepts. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. Claim 3 does not recite any additional elements and claim 10 only recites the additional element “wherein the computer program causes the hardware processor to generate the targeted visual question answer dataset”. This additional element is recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computer) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Step 2B Analysis: The claims do not recite additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element “wherein the computer program causes the hardware processor to generate the targeted visual question answer dataset” in claim 10 amounts to no more than using generic computer components to implement the exception such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The additional element does not provide an inventive concept, and, therefore, the claims are not patent eligible. Regarding claims 4, 11 and 18, these claims are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 4 is directed to a method as depending from claim 1, claim 11 is directed to a system as depending from claim 8 and claim 18 is directed to a computer program product embodied in a non-transitory computer readable medium as depending from claim 16, thus the analysis for patent eligibilities of claims 1, 8 and 16 are incorporated herein. Step 2A Prong 1: The claims recite, using respective similar language, “maximize a conditional likelihood of a question-answer pair, given an image.” The additional limitation added by these claims covers mathematical concepts - mathematical relationships, mathematical formulas or equations, or mathematical calculations (e.g., maximize a conditional likelihood of a question-answer pair, based on observed/received image data). Therefore, the claims are directed to an abstract idea - mental processes combined with mathematical concepts. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claims recite, using respective similar language, the additional element: “wherein training the teacher model includes learning through deep learning to maximize a conditional likelihood”. The above-noted training step/operation is simply generic deep learning/training to perform the abstract idea of maximize a conditional likelihood of a question-answer pair, given a received/observed image and amounts to mere instructions to apply the exception (MPEP 2106.05(f)). That is, the training is invoked as a tool to perform the underlying abstract idea of maximize a conditional likelihood of a question-answer pair, given an observed image (i.e., a mathematical concept). Machine learning/training so recited cannot provide an improvement to the functioning of a computer or other technological field. MPEP § 2106.05(a)(I). Claim 11 also recites the additional element “wherein the computer program causes the hardware processor to train the teacher model”. This additional element is recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computer) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Step 2B Analysis: The claims do not recite additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, applying the generically-recited teacher model and generic training to perform the abstract idea (mathematical concept) amounts to no more than mere instructions to apply the exception (MPEP 2106.05(f). As also discussed above, the additional element “wherein the computer program causes the hardware processor to train the teacher model” in claim 11 amounts to no more than using generic computer components to implement the exception such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The additional elements do not provide an inventive concept, and, therefore, the claims are not patent eligible. Regarding claims 5 and 12, these claims are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 5 is directed to a method as depending from claim 1 and claim 12 is directed to a system as depending from claim 8, thus the analysis for patent eligibilities of base claims 1 and 8 are incorporated herein. Step 2A Prong 1: The claims recite, using respective similar language, “wherein pseudolabeling unlabeled images includes producing pseudolabels for unlabeled images Iu by obtaining logits of a decoder, and the logits defining a distribution over tokens of a natural language vocabulary of the teacher model.” The additional limitations added by these claims cover mathematical concepts - mathematical relationships, mathematical formulas or equations, or mathematical calculations (e.g., generating/producing pseudolabels for observed/received unlabeled images Iu by obtaining logits of a decoder, the logits defining a distribution over tokens of a natural language vocabulary). In light of the specification, (see, paragraph 41, “To produce a pseudolabel (Q’, A’) for an unlabeled image Iu, L1:N = VQGIC(Iu) is obtained, where LI:N are the logits of the decoder. The logits LI:N define a distribution P(LI:N | LI:N-1) over tokens of the model's natural language vocabulary. Nucleus sampling can then be applied to stochastically decode a text T’ from P(LI:N | LI:N-1). To recover a pseudo-question-answer pair (Q’, A’) from the decoded text T', the structured format of the generation template in equation (1) is exploited to recover the generated question and answer (Q’, A’).”), these limitations are directed to a mathematical concepts. Regarding the “teacher model”, no details of the model or its training are recited, and the model is recited at a high level of generality and the model can be constructed by hand with pen and paper. Thus, the claimed “teacher model”, under the BRI, in light of the specification, could be any model including tokens of a natural language vocabulary, which could be constructed by hand with pen and paper based on a reasonable amount of observed data (i.e., a natural language vocabulary and its tokens). Given a sufficiently small set of tokens in the “natural language vocabulary”, nothing in the claims prohibit the producing pseudolabels for unlabeled images step/operation from being performed mentally or with pen and paper. The teacher model is recited at a high level of generality and therefore is being interpreted as performing an abstract idea on a generic computer. Therefore, the claims are directed to an abstract idea - mental processes combined with mathematical concepts. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. The claims recite the additional element of “tokens of a natural language vocabulary of the teacher model.” The “teacher model” is recited at a high level of generality as mere instructions to implement an abstract idea on a computer and amounts to the recitation of the words “apply it” (or an equivalent) or amount to no more than mere instructions to implement an abstract idea or other exception on a computer or merely uses a computer as a tool to perform an abstract idea (i.e., as generic computer components performing generic computer functions). See MPEP 2106.05(f). Claim 5 does not recite any additional elements and claim 12 only recites the additional element “wherein the computer program causes the hardware processor to pseudolabel unlabeled images Iu by obtaining logits”. This additional element is recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computer) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Step 2B Analysis: The claims do not recite additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, applying the generically-recited “teacher model” to perform the abstract idea amounts to no more than mere instructions to apply the exception (MPEP 2106.05(f). As also discussed above, the additional element “wherein the computer program causes the hardware processor to pseudolabel unlabeled images Iu by obtaining logits” in claim 12 amounts to no more than using generic computer components to implement the exception such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The additional element does not provide an inventive concept, and, therefore, the claims are not patent eligible. Regarding claims 6, 13 and 19, these claims are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 6 is directed to a method as depending from claim 1, claim 13 is directed to a system as depending from claim 8 and claim 19 is directed to a computer program product embodied in a non-transitory computer readable medium as depending from claim 16, thus the analysis for patent eligibilities of claims 1, 8 and 16 are incorporated herein. Step 2A Prong 1: The claims recite, using respective similar language, “wherein, given an image I, a question Q and answer A, the student model approximates P(A| Q, I), while the teacher model approximates P(Q, A | I), where P(A| Q, I) is conditional probability of A in Q, I and P(A, Q| I) is conditional probability of A and Q in I to enable using unlabeled data in training1.” The additional limitation added by these claims covers mathematical concepts - mathematical relationships, mathematical formulas or equations, or mathematical calculations (e.g., calculating approximations of P(Q, A | I) by generically-recited teacher and student models as calculated, conditional probabilities). Therefore, the claims are directed to an abstract idea - mental processes combined with mathematical concepts. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claims recite, using respective similar language, the additional element: “the student model approximates P(A| Q, I), while the teacher model approximates P(Q, A | I) … to enable using unlabeled data in training2.” The above-noted approximation steps/operations (i.e., abstract idea - mathematical concepts) are performed by the generically-recited “student model” and “teacher model”. The models are recited at a high level of generality as mere instructions to implement an abstract idea on a computer and amount to the recitation of the words “apply it” (or an equivalent) or amount to no more than mere instructions to implement an abstract idea or other exception on a computer or merely uses a computer as a tool to perform an abstract idea (i.e., as generic computer components performing generic computer functions). See MPEP 2106.05(f). While the claims recite “to enable using unlabeled data in training”, they do not positively recite any such use of the unlabeled data in any actual model training. Further, it is unclear which training is being referred to (i.e., training of the teacher model and/or student model, the VLM recited in base claims 1, 8 and 16, or to some other, unrecited model?). The “training” is generic machine learning/training to perform the abstract idea of calculating approximations as calculated, conditional probabilities (i.e., a mathematical concept) and amounts to mere instructions to apply the exception (MPEP 2106.05(f)). That is, the training is invoked as a tool to perform the underlying abstract idea of calculating approximations of P(Q, A | I) by generically-recited teacher and student models as calculated, conditional probabilities (i.e., a mathematical concept). Machine learning/training so recited cannot provide an improvement to the functioning of a computer or other technological field. MPEP § 2106.05(a)(I). These additional elements are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computer) such that they amount to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Step 2B Analysis: The claims do not recite additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, applying the generically-recited student and teacher models and generic training to perform the abstract idea (mathematical concept) amounts to no more than mere instructions to apply the exception (MPEP 2106.05(f). Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The additional elements do not provide an inventive concept, and, therefore, the claims are not patent eligible. Regarding claims 7, 14 and 20, these claims are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 7 is directed to a method as depending from claim 1, claim 14 is directed to a system as depending from claim 8 and claim 20 is directed to a computer program product embodied in a non-transitory computer readable medium as depending from claim 16, thus the analysis for patent eligibilities of claims 1, 8 and 16 are incorporated herein. Step 2A Prong 1: The claims recite, using respective similar language, “responding to medical inquiries with visual answers to assist in decision making for medical personnel3.” The additional limitation added by these claims covers mental processes – responding to received/observed medical inquiries/questions with visual answers/images. Given a sufficiently small set of medical inquiries/questions/queries, nothing in the claims prohibit the responding to medical inquiries with visual answers step/operation from being performed mentally or with pen and paper. Therefore, the claims are directed to an abstract idea - mental processes. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claims recite, using respective similar language, the additional element: “wherein the student model is trained on medical images and information”. The “student model” is recited at a high level of generality and therefore is being interpreted as performing a mental process on a generic computer. See MPEP 2106.04(a)(2) § III.C which states that “a concept that is performed in the human mind and applicant is merely claiming that concept performed 1) on a generic computer, or 2) in a computer environment, or 3) is merely using a computer as a tool to perform the concept” still recite a mental process. Also, “the student model is trained on medical images and information” is recited at a high level of generality as mere instructions to implement an abstract idea on a computer and amount to the recitation of the words “apply it” (or an equivalent) or amount to no more than mere instructions to implement an abstract idea or other exception on a computer or merely uses a computer as a tool to perform an abstract idea (i.e., as generic computer components performing generic computer functions). See MPEP 2106.05(f). Regarding “the student model is trained on medical images and information”, the training is generic machine learning/training to perform the abstract idea of responding to received/observed medical inquiries/queries/questions with visual answers (i.e., a mental process) and amounts to mere instructions to apply the exception (MPEP 2106.05(f)). That is, the training is invoked as a tool to perform the underlying abstract idea of responding to inquiries/questions (i.e., a mental process). Machine learning/training so recited cannot provide an improvement to the functioning of a computer or other technological field. MPEP § 2106.05(a)(I). These additional elements are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computer) such that they amount to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Step 2B Analysis: The claims do not recite additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, applying the generically-recited student model and generic training to perform the abstract idea (mental process) amounts to no more than mere instructions to apply the exception (MPEP 2106.05(f). Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The additional elements do not provide an inventive concept, and, therefore, the claims are not patent eligible. Regarding claim 15, this claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 15 is directed to a system as depending from claim 14, thus the analysis for patent eligibilities of claims 14 and of base claim 8 are incorporated herein. Step 2A Prong One: The claim does not recite any judicial exception. However, it is still directed to the same abstract ideas as identified in claims 8 and 14, discussed above. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recites this additional element: “wherein the computer program causes the hardware processor to display visual responses on a display device.” The additional element “wherein the computer program causes the hardware processor to display visual responses” is recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computer) such that it amounts to no more than mere instructions to apply the exception using generic computer components (i.e., the processor and display device) (MPEP 2106.05(b)). Also, the limitation “display visual responses on a display device” is post-solution activity that is not integrated into the claim as a whole and does not add a meaningful limitation to the above-noted mental processes specified in this claim. See MPEP § 2106.05(g); OIP Techs., Inc. v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1092-93 (Fed. Cir. 2015) (presenting offers and gathering statistics amounted to mere data gathering). The “display visual responses on a display device” amounts to insignificant extra-solution activity of data outputting for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. As such, the above-noted limitation amounts to an insignificant extra-solution activity, which does not integrate a judicial exception into a practical application. That is, this limitation is adding insignificant extra-solution activity (amounts to necessary data outputting) to the judicial exception, as discussed in MPEP § 2106.05(g). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element is an insignificant extra-solution activity. Insignificant extra-solution activities cannot provide an inventive concept. As also discussed above, the additional element “wherein the computer program causes the hardware processor to display visual responses” amounts to no more than using generic computer components to implement the exception such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)). Moreover, receiving, communicating, and storing data are insignificant extra-solution activities that are well-understood, routine, and conventional. See MPEP2106.05(d)(II) (“The courts have recognized the following computer functions as well‐understood, routine, and conventional functions… i. Receiving or transmitting data over a network…iv. Storing and retrieving information in memory”) (citing OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015)). Therefore, the recitation of “display visual responses on a display device” is the well-understood, routine, conventional activity of transmitting data over a network, as discussed in MPEP § 2106.05(d). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 4, 8, 11, 16 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over non-patent literature Yang, et al. ("Self-Training Vision Language BERTs with a Unified Conditional Model." arXiv preprint arXiv:2201.02010 (Jan 2022), hereinafter “Yang”) in view of Li et al. (U.S. Patent Application Pub. No. 2022/0391755 A1, hereinafter “Li”). With respect to claim 1, Yang discloses the invention as claimed including a computer-implemented method for training a visual question answer model (see, e.g., Abstract, “We propose a self-training approach that allows training VL-BERTs [visual language-bidirectional encoder representations from transformers] from unlabeled image data. … We then combine the labeled data and pseudo labeled data to train a student model. The process is iterated by putting the student model as a new teacher. By using the proposed self-training approach and only 300k unlabeled extra data” and page 4, Sect III B, “Our self-training approach is done by iterating three steps: train with labeled data, generate pseudo labeling on unlabeled data, mix labeled and unlabeled data, and re-train. … We first train UCM [Unified Conditional Model] using … captioning annotations in Visual Genome and questions in VQA” [i.e., a method for training a visual question answer/VQA/VL visual language model]), comprising: training a teacher model by performing image conditional visual question generation on a visual language model (VLM) and a targeted visual question answer dataset using images (see, e.g., Abstract, “We use the labeled image data to train a teacher model and use the trained model to generate pseudo captions on unlabeled image data.” and pages 1-2, Sect. I, “vision language pretraining requires paired image and language”, “Vision Language BERT Self-Training. Our self-training approach follows the self-training pipeline with extra optimization for vision language BERTs. … First, we use the labeled image data to train a teacher model and then use the trained model to generate pseudo labels on unlabeled image data.” and 4-5, sects. III B and IV. C, “Train UCM with labeled data. We first train UCM using captioning annotations in MSCOCO, dense captioning annotations in Visual Genome and questions in VQA [27] and GQA dataset” [i.e., targeted question answer/QA dataset] and “we separate the two images into two question and image pairs and process each pair using our proposed model.” [i.e., training a teacher model via image conditional visual question generation on a visual language/VL model/VLM and a targeted question answer/QA dataset to generate QA pairs]); pseudolabeling unlabeled images using the teacher model to decode synthetic question and answer pairs for the unlabeled images (see, e.g., Abstract, “We then combine the labeled data and pseudo labeled data to train a student model.” and pages 2-3, Sects. I and III A, “First, we use the labeled image data to train a teacher model and then use the trained model to generate pseudo labels on unlabeled image data. We then combine the labeled data and pseudo labeled data to train a student model … the pseudo labels are generated image captions and dense captions from UCM.”, “We process the language embeddings through the bi-directional language encoder and the one-directional language encoder. … The images are processed by the image encoder.” and 4-5, Sects. III B and IV B, “Train new model by mixing labeled and unlabeled data. After pseudo labels are generated for unlabeled images, we mix the pseudo labeled data and original labeled data to train a new model.” and “we use extra QA data from Visual Genome for data augmentation.” [i.e., use the teacher model for pseudo-labeling unlabeled images by decoding encoded synthetic/augmented QA/question-answer pairs]); merging the synthetic question and answer pairs for the unlabeled images with real data from the targeted visual question answer dataset to generate a self-augmented training set (see, e.g., FIG. 2 – depicting “During training, … merge visual features with the outputs from both language encoders” and pages 2-3, Sects I and III A, “use the trained model to generate pseudo labels on unlabeled image data. We then combine the labeled data and pseudo labeled data … We propose a self-training method for using unlabeled images in vision language pretraining”, “The images are processed by the image encoder. After that, the bi-directional output and one-directional output are merged with image output”, and 4-5 Sects. III B and IV B, “Train UCM with labeled data. We first train UCM using captioning annotations in MSCOCO, dense captioning annotations in Visual Genome and questions in VQA [27] and GQA dataset” and “we use extra QA data from Visual Genome for data augmentation.” [i.e., merging/combining the synthetic/augmented QA/question-answer pairs for unlabeled images with real data from the targeted QA dataset to generate self-augmented training data for self-training]); and training a student model using the VLM and the self-augmented training set to return visual answers to text queries (see, e.g., Abstract and pages 2, Sect. I, “We then combine the labeled data and pseudo labeled data to train a student model. … use the teacher model to generate pseudo labels for unlabeled data and finally use the labeled data and pseudo labeled data to jointly train a student model.” 3, Sect. III, “doing image-language understanding tasks, for example finetuning visual question answering” and 5, Sect. IV C, “VQAv2 dataset [27] is to answer questions given an image. The answering process is usually formatted as a classification task within all possible answers” [i.e., train a student model using the VL model and the combined/augmented training set to output answers to text questions/queries]). Although Yang substantially discloses the claimed invention, Yang is not relied on for explicitly disclosing training a teacher model … using images to generate question and answer pairs. However, in the same field, analogous art Li teaches training a teacher model … using images to generate question and answer pairs (see, e.g., paragraphs 18, “The VLP [Vison-and-learning pretraining] module 130 may generate an output 150 such as one or more output image-text pairs”, 41-42, “The momentum model is a continuously-evolving teacher model”, “During training, the visual-and-language base model can be trained so that its predictions match the predictions from the momentum model” 71, “Visual question answering ("VQA") requires the model to predict an answer given an image and a question. … VQA can be framed as an answer generation problem.” [i.e., training a teacher model by using images to generate QA pairs]). Yang and Li are analogous art because they are both directed to implementing, training and using visual language/VL, V+L models (see, e.g., Yang, Abstract and page 2, Sect. I, and Li, Abstract and paragraph 22). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yang to incorporate the teachings of Li to provide a “VLP module 130 [that] may generate an output 150 such as one or more output image-text pairs” and a technique for “training, the visual-and-language base model [i.e., a VLM] … so that its predictions match the predictions from the momentum model.” (see, e.g., Li, paragraph 42). Doing so would have allowed Yang to use Li’s model and training technique “in order to improve learning, such as in the presence of noisy input data for training the model” and “MoD can be considered as performing data augmentation to the original views, MoD [Momentum distillation] generates a diverse set of views that are absent in the original image-text pairs, which can improve the model's generalization performance”, as suggested by Li (See, e.g., Li, paragraphs 41 and 61). With respect to independent claim 8, claim 8 is substantially similar to claim 1 and therefore is rejected on the same grounds as claim 1, discussed above. In particular, claim 8 is a system claim with operations that correspond to the method steps of claim 1. Yang further discloses training a visual question answer model (see, e.g., Abstract, “We propose a self-training approach that allows training VL-BERTs [visual language-bidirectional encoder representations from transformers] from unlabeled image data. … We then combine the labeled data and pseudo labeled data to train a student model. The process is iterated by putting the student model as a new teacher. By using the proposed self-training approach and only 300k unlabeled extra data” and page 4, Sect III B, “Our self-training approach is done by iterating three steps: train with labeled data, generate pseudo labeling on unlabeled data, mix labeled and unlabeled data, and re-train. … We first train UCM [Unified Conditional Model] using … captioning annotations in Visual Genome and questions in VQA” [i.e., a method for training a visual question answer/VQA/VL visual language model]). Although Yang substantially discloses the claimed invention, Yang is not relied on for explicitly disclosing a system for training a visual question answer model, comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to: <perform operations>. However, in the same field, analogous art Li teaches a system for training a visual question answer model, comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to: <perform operations> (see, e.g., paragraphs 15, “a computing device for implementing a VLP system for training a vision-and-learning (V+L) model … computing device 100 includes a processor 110 coupled to memory 120 … Computing device 100 may be implemented as a stand-alone subsystem” and 18, “memory 120 may include non-transitory, tangible, machine readable media that includes executable code that when run by one or more processors (e.g., processor 110) may cause the one or more processors to perform the methods described in further detail herein. … memory 120 includes instructions for a VLP module 130 that may be used to implement and/or emulate the systems and models, and/or to implement any of the methods described further herein.”). Yang and Li are analogous art because they are both directed to implementing, training and using visual language/VL, V+L models (see, e.g., Yang, Abstract and page 2, Sect. I, and Li, Abstract and paragraph 22). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yang to incorporate the teachings of Li to provide a “VLP module 130 [that] may generate an output 150 such as one or more output image-text pairs” and a technique for “training, the visual-and-language base model [i.e., a VLM] … so that its predictions match the predictions from the momentum model.” (see, e.g., Li, paragraph 42). Doing so would have allowed Yang to use Li’s model and training technique “in order to improve learning, such as in the presence of noisy input data for training the model” and “MoD can be considered as performing data augmentation to the original views, MoD [Momentum distillation] generates a diverse set of views that are absent in the original image-text pairs, which can improve the model's generalization performance”, as suggested by Li (See, e.g., Li, paragraphs 41 and 61). With respect to independent claim 16, claim 16 is substantially similar to claim 1 and therefore is rejected on the same grounds as claim 1, discussed above. In particular, claim 16 is a computer program product claim with method steps that correspond to the method steps of claim 1. Yang further discloses training a visual question answer model (see, e.g., Abstract, “We propose a self-training approach that allows training VL-BERTs [visual language-bidirectional encoder representations from transformers] from unlabeled image data. … We then combine the labeled data and pseudo labeled data to train a student model. The process is iterated by putting the student model as a new teacher. By using the proposed self-training approach and only 300k unlabeled extra data” and page 4, Sect III B, “Our self-training approach is done by iterating three steps: train with labeled data, generate pseudo labeling on unlabeled data, mix labeled and unlabeled data, and re-train. … We first train UCM [Unified Conditional Model] using … captioning annotations in Visual Genome and questions in VQA” [i.e., training a visual question answer/VQA/VL visual language model]). Although Yang substantially discloses the claimed invention, Yang is not relied on for explicitly disclosing a computer program product for training a visual question answer model, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method. However, in the same field, analogous art Li teaches a computer program product for training a visual question answer model, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method (see, e.g., paragraphs 15, “a computing device for implementing a VLP system for training a vision-and-learning (V+L) model … computing device 100 includes a processor 110 coupled to memory 120” and 18, “memory 120 may include non-transitory, tangible, machine readable media that includes executable code that when run by one or more processors (e.g., processor 110) may cause the one or more processors to perform the methods described in further detail herein. … memory 120 includes instructions for a VLP module 130 that may be used to implement and/or emulate the systems and models, and/or to implement any of the methods described further herein.”). Yang and Li are analogous art because they are both directed to implementing, training and using visual language/VL, V+L models (see, e.g., Yang, Abstract and page 2, Sect. I, and Li, Abstract and paragraph 22). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yang to incorporate the teachings of Li to provide a “VLP module 130 [that] may generate an output 150 such as one or more output image-text pairs” and a technique for “training, the visual-and-language base model [i.e., a VLM] … so that its predictions match the predictions from the momentum model.” (see, e.g., Li, paragraph 42). Doing so would have allowed Yang to use Li’s model and training technique “in order to improve learning, such as in the presence of noisy input data for training the model” and “MoD can be considered as performing data augmentation to the original views, MoD [Momentum distillation] generates a diverse set of views that are absent in the original image-text pairs, which can improve the model's generalization performance”, as suggested by Li (See, e.g., Li, paragraphs 41 and 61). Regarding claims 4, 11 and 18, as discussed above, Yang in view of Li teaches the method of claim 1, the system of claim 8, and the computer program product of claim 16. Yang further discloses wherein training the teacher model includes learning through deep learning to maximize a conditional likelihood of a question-answer pair, given an image (see, e.g., FIG. 5 – depicting “the attention map at condition tokens” where a conditional/condition flag “[CND] affects more in deeper layers.” and pages 3-4, Sect III A, “The loss of bi-directional CMLM is defined as the negative log-likelihood of predicting the masked words given conditions … Image-Text Matching … done for the bi-directional prediction branch. At 50% possibility, we assign a fake sentence to the image. Specifically, we use the final feature at position [CLS] to represent the summary of current visual and language input. We use this feature to classify whether the current input text and image are matched. … Question Answering The qa task is only done for the bi-directional prediction branch. If the sampled text is a question, we use the feature at position [CLS] and a qa header to generate its answer and calculate its loss based on classification.”, “The loss of MOAM is defined as the negative log-likelihood of predicting the masked regions’ class and attributes” and 7, Sect. IV F, “[CND] affects deeper layers more. … the deeper layers tend to assign higher weights to [CND] position. This is because the deeper layers are more directly related to producing results, thus they rely more on the [CND] flag to condition” [i.e., training the teacher model includes deep learning to maximize a percentage possibility/conditional likelihood of a QA pair given an image]). Claims 7, 14, 15 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Yang in view of Li as applied to claims 1, 8 and 16 above, and further in view of non-patent literature Gong, et al. ("VQAmix: Conditional triplet mixup for medical visual question answering." IEEE Transactions on Medical Imaging 41.11 (June 2022): 3332-3343, hereinafter “Gong”). Regarding claims 7, 14 and 20, as discussed above, Yang in view of Li teaches the method of claim 1, the system of claim 8, and the computer program product of claim 16. Yang further discloses wherein the student model is trained on… images and information and further comprises responding to … inquiries with visual answers (see, e.g., Abstract and pages 2, Sect. I, “First, we use the labeled image data to train a teacher model and then use the trained model to generate pseudo labels on unlabeled image data. We then combine the labeled data and pseudo labeled data to train a student model. … use the teacher model to generate pseudo labels for unlabeled data and finally use the labeled data and pseudo labeled data to jointly train a student model.” 3, Sect. III, “doing image-language understanding tasks, for example finetuning visual question answering” and 5, Sect. IV C, “VQAv2 dataset [27] is to answer questions given an image. The answering process is usually formatted as a classification task within all possible answers” [i.e., train the student model using images and information/pseudo labels set to output visual answers responding to text questions/queries]). Although Yang in view of Li substantially teaches the claimed invention, Yang in view of Li is not relied on to teach wherein the student model is trained on medical images and information and further comprises responding to medical inquiries with visual answers to assist in decision making for medical personnel. However, in the same field, analogous art Gong teaches wherein the student model is trained on medical images and information and further comprises responding to medical inquiries with visual answers to assist in decision making for medical personnel (see, e.g., Abstract “Medical visual question answering (VQA) aims to correctly answer a clinical question related to a given medical image … VQAMix generates more labeled training samples by linearly combining a pair of VQA samples, which can be easily embedded into any visual-language model” [i.e., training a visual question answering/VQA model] and pages 1, Sect. I, “Medical VQA can assist clinicians to obtain a second opinion on diagnosis and enhance their confidence in interpreting complex medical images.”, 3, Sect. II, “a multi-task learning framework by generating the pseudo labels to the unlabeled data according to the modal of medical image.” and 8-9, Sect. V, “In this work, we use the student’s t-test to demonstrate the effectiveness of the proposed methods by comparing the overall accuracy of the best-performed”, “VQAMix works well on the medical visual question answering task.” [i.e., student model is trained on medical images and information/samples and responds to medical inquiries/clinical questions with visual answers to assist in decision making for medical personnel/clinicians]). Yang, Li and Gong are analogous art because they are each directed to implementing, training and using visual language/VL, V+L models (see, e.g., Yang, Abstract and page 2, Sect. I, Li, Abstract and paragraph 22, and Gong, Abstract). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yang in view of Li to incorporate the teachings of Gong to provide “a simple yet effective data augmentation method, VQAMix, to mitigate the data limitation problem.” for “Medical visual question answering (VQA) [that] aims to correctly answer a clinical question related to a given medical image.” (see, e.g., Gong, Abstract). Doing so would have allowed Yang in view of Li to use Gong’s VQAMix method that “significantly improves the performance of the baseline by about 7% and 5% on the averaging result of two backbones, respectively. More importantly, VQAMix could improve confidence calibration and model interpretability, which is significant for medical VQA models in practical applications”, as suggested by Gong (See, e.g., Gong, Abstract). Regarding claim 15, as discussed above, Yang in view of Li and Gong teaches the system of claim 14. Although Yang substantially discloses the claimed invention, and FIGs. 2, 4 and 6 of Yang illustrates displays of “visual output”/visual responses, “generated image descriptions” and “the captions are generated given different visual masks” on a screen/display device, Yang is not relied on for explicitly disclosing wherein the computer program causes the hardware processor to display visual responses on a display device. However, in the same field, analogous art Li teaches wherein the computer program causes the hardware processor to display visual responses on a display device (see, e.g., paragraphs 15, “Operation of computing device 100 is controlled by processor 110”, 18, “the VLP module 130, may receive several inputs, e.g., such as an image input 142 and a text input 144 from a user, via a data interface 115. The data interface 115 may be any of a user interface that receives an image input and text input from a user, … The VLP module 130 may generate an output 150 such as one or more output image-text pairs.” [i.e., the program causes the processor to display output/visual responses on a user interface/display device]). Yang and Li are analogous art because they are both directed to implementing, training and using visual language/VL, V+L models (see, e.g., Yang, Abstract and page 2, Sect. I, and Li, Abstract and paragraph 22). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yang to incorporate the teachings of Li to provide a “VLP module 130 [that] may generate an output 150 such as one or more output image-text pairs” and a technique for “training, the visual-and-language base model [i.e., a VLM] … so that its predictions match the predictions from the momentum model.” (see, e.g., Li, paragraph 42). Doing so would have allowed Yang to use Li’s model and training technique “in order to improve learning, such as in the presence of noisy input data for training the model” and “MoD can be considered as performing data augmentation to the original views, MoD [Momentum distillation] generates a diverse set of views that are absent in the original image-text pairs, which can improve the model's generalization performance”, as suggested by Li (See, e.g., Li, paragraphs 41 and 61). Conclusion The prior art made of record, listed on form PTO-892, and not relied upon, is considered pertinent to applicant's disclosure. The references listed on form PTO-892 are all generally related to techniques, methods and systems for implementing, training and using machine learning models, such as teacher and student models, using visual question-answer training data. For example, non-patent literature Pan et al. ("Muvam: A multi-view attention-based model for medical visual question answering." arXiv preprint arXiv:2107.03216 (July 2021), hereinafter “Pan”) discloses “a multi-view attention-based model (MuVAM) for medical visual question answering which integrates the high-level semantics of medical images on the basis of text description” and “The overall framework of MuVAM is shown in Fig. (2). Given a medical image v and a question Q related to the image, the correct answer is finally predicted. … to maximize the accuracy of the prediction results, the composite loss is composed of the classification loss and the IQC loss to train MuVAM together” [i.e., a method for training a visual question answering/VQA model] (see, Abstract and page 3, Sect. III). The examiner requests, in response to this office action, support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line no(s) in the specification and/or drawing figure(s). This will assist the examiner in prosecuting the application. When responding to this office action, Applicant is advised to clearly point out the patentable novelty which he or she thinks the claims present, in view of the state of the art disclosed by the reference cited or the objections made. He or she must also show how the amendments avoid such references or objections See 37 CFR 1.111 (c). Any inquiry concerning this communication or earlier communications from the examiner should be directed to RANDY K BALDWIN whose telephone number is (571)270-5222. The examiner can normally be reached on Mon - Fri 9:00-6:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RANDALL K. BALDWIN/Primary Examiner, Art Unit 2125 1 Also, the limitation includes intended use language with no patentable weight (e.g., “to enable using unlabeled data in training”). The claims do not positively recite any such use of the unlabeled data in actual model training, and it is unclear which training is being referred to (i.e., training of the VLM, teacher model and/or student model, or to some other, unrecited model?). 2 As noted above, the limitation includes intended use language with no patentable weight (e.g., “to enable using unlabeled data in training”). 3 The limitation includes intended use language with no patentable weight (e.g., “to enable using unlabeled data in training”).
Read full office action

Prosecution Timeline

Oct 30, 2023
Application Filed
Jul 24, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694291
AUTOMATED DRIFT DETECTION IN MULTIDIMENSIONAL DATA
3y 5m to grant Granted Jul 28, 2026
Patent 12688405
ANALOG NEUROMOPRHIC CIRCUIT WITH STACKS OF RESISTIVE MEMORY CROSSBAR CONFIGURATIONS
3y 12m to grant Granted Jul 21, 2026
Patent 12675694
USING EMBEDDING FUNCTIONS WITH A DEEP NETWORK
2y 4m to grant Granted Jul 07, 2026
Patent 12670383
SYSTEM AND METHOD OF USING FRACTIONAL ADAPTIVE LINEAR UNIT AS ACTIVATION IN ARTIFACIAL NEURAL NETWORK
4y 6m to grant Granted Jun 30, 2026
Patent 12664417
MAGNETIC EFFECT ARTIFICIAL INTELLIGENCE SYSTEM
3y 3m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
99%
With Interview (+28.0%)
3y 5m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 242 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month