Prosecution Insights
Last updated: October 02, 2026
Application No. 18/294,532

SIMILARITY RETRIEVAL

Non-Final OA §101§103§DOUBLEPATENT
Filed
Feb 01, 2024
Priority
Aug 02, 2021 — EU 21189078.5 +1 more
Examiner
KABIR, SAMIYAH
Art Unit
Tech Center
Assignee
Bayer Aktiengesellschaft
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
8 currently pending
Career history
5
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103 §DOUBLEPATENT
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C 102 and 103 (or as subject to pre-AIA 35 U.S.C 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Claims 1-15 are subject to review. Information Disclosure Statement The information disclosure statement (IDS) submitted on 02/01/2024 is being considered by the examiner. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1, 2, 4, 5, 11, 12, 14, and 15 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over Claims 1, 5, 6, 12, 13, and 14 of copending Application No. 18/280,160 (reference application) in view of Zhang, Zhenzhong (US 20220076052 A1), hereinafter referred to as Zhang. Although the claims at issue are not identical, they are not patentably distinct from each other because they are both directed towards the same inventive concept. Claims 3 and 9 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over Claim 1 of copending Application No. 18/280,160 (reference application) in of NPL reference Qiu et al. “Battling Alzheimer’s Disease through Early Detection: A Deep Battling Alzheimer’s Disease through Early Detection: A Deep Multimodal Learning Approach Multimodal Learning Approach” hereinafter referred to as Qiu. Although the claims at issue are not identical, they are not patentably distinct from each other because they are both directed towards the same inventive concept. Claims 7, 8, and 13 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over Claims 1 of copending Application No. 18/280,160 (reference application) in view of Lei et al. (US 20180211182 A1) hereinafter to referred to as Lei. Although the claims at issue are not identical, they are not patentably distinct from each other because they are both directed towards the same inventive concept. Claim 6 is provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over Claim 1 of copending Application No. 18/280,160 (reference application) in view of NPL reference Udandarao et al. (“COBRA: Contrastive Bi-Modal Representation Algorithm”), hereinafter referred to as Udandarao. Although the claims at issue are not identical, they are not patentably distinct from each other because they are both directed towards the same inventive concept. Claim 10 is provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over Claim 1 of copending Application No. 18/280,160 (reference application) in view of Udandarao and further in view of Chen et al. (US 20210319266 A1) hereinafter to referred to as Chenthey are both directed towards the same inventive concept. This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented. Instant Application Patent Application No. 18/280,160 Claim 1 A computer-implemented method, comprising: providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer; providing training data for training the machine learning model, wherein providing training data comprises: receiving, for each object of a multitude of objects, input data of at least two different modalities, first input data of a first modality and second input data of a second modality, generating first augmented input data from the first input data and second augmented input data from the second input data, and generating first masked input data from the first augmented input data and second masked input data from the second augmented input data; training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input layer, inputting the second masked input data into the second input layer, reconstructing the first augmented input data from the first masked input data via the first output layer, reconstructing the second augmented input data from the second masked input data via the second output layer, generating a joint representation of the first masked input data and the second masked input data via the third output layer, and discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data of different objects; receiving input data related to a first object; inputting the input data related to the first object into the trained machine learning model; receiving from the trained machine learning model a first representation of the first object via the third output layer; receiving at least one second representation of at least one second object; computing a similarity value, the similarity value indicating the similarity between the first representation and the at least one second representation; and outputting the similarity value and/or information related to the at least one second object. Claim 1 A computer-implemented method of training a machine learning model, the method comprising: receiving, for each object of a multitude of objects, input data, comprising first input data of a first modality and second input data of a second modality; generating first augmented input data from the first input data and second augmented input data from the second input data; generating first masked input data from the first augmented input data and second masked input data from the second augmented input data; providing a machine learning model, the machine learning model comprising: a first input, a second input, a first output, a second output, and a third output, training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input, inputting the second masked input data into the second input, reconstructing the first augmented input data from the first masked input data via the first output, reconstructing the second augmented input data from the second masked input data via the second output, generating a joint representation of the first masked input data and the second masked input data via the third output, and discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data of different objects; and storing and/or outputting the trained machine learning model and/or providing the trained machine learning model for predictive purposes. The patent application cited above fails to teach the above bolded limitations at this claim, however, Zhang teaches them at Paragraph [0022]. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with Zhang in order to increase the accuracy of similarity results involving multimodal data. Claim 2 The method of claim 1, wherein the first object, the at least one second object and each object of the multitude of objects is a human being. Claim 5 The method of claim 1, wherein the object is a living object, or a part thereof. Claim 3 The patent application cited above fails to teach the limitations at this claim, however, Qiu teaches them at Pg. 7, Experimental Study, Data. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with Qui in order to improve model performance using multi-modal data. Claim 4 The method of claim 1, wherein the first object, the at least one second object and each object of the multitude of objects is a plant or a plurality of plants or one or more parts of a plant. Claim 6 The method of claim 1, wherein the object is a crop or a collection of crops or an agricultural field or a part of the Earth's surface. Claim 5 The method of claim 1, wherein the first object, the at least one second object and each object of the multitude of objects is a part of the Earth's surface. Claim 6 The method of claim 1, wherein the object is a crop or a collection of crops or an agricultural field or a part of the Earth's surface. Claim 6 The method of claim 1, wherein the first input data of the first modality and the second input data of the second modality comprise or are derived from one or more images, text files and/or audio files, wherein the first modality is different from the second modality. Claim 7 The method of claim 1, wherein the second input data of the second modality comprise, for each object, text data. The patent application cited above fails to teach the above bolded limitations at this claim, however, Udandarao teaches them at Abstract. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with Udandarao in order to improve the quality of joint cross-modal embedding spaces. Claim 7 The patent application cited above fails to teach the limitations at this claim, however, Lei teaches them at Paragraph [0024]. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with Lei in order to improve model efficiency and reduce computational costs. Claim 8 The patent application cited above fails to teach the limitations at this claim, however, Lei teaches them at Paragraph [0031]. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with Lei in order to improve model efficiency and reduce computational costs. Claim 9 The patent application cited above fails to teach the limitations at this claim, however, Qui teaches them at Pg. 6, Paragraph 5. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with Qui in order to improve model performance using multi-modal data. Claim 10 The method of claim 1, wherein the machine learning model is or comprises a deep neural network, wherein the deep neural network comprises, at least for the training, a first encoder, a first decoder, a second encoder, a second decoder, a fusion component, an attention weighted pooling, and a projection head, wherein: the first encoder is configured to receive the first masked input data, and to generate the first representation from the first masked input data, the second encoder is configured to receive the second masked input data, and to generate the second representation from the second masked input data, the fusion component is configured to generate the joint representation from the first representation and the second representation, the first decoder is configured to reconstruct the first augmented input data from the joint representation, the second decoder is configured to reconstruct the second augmented input data from the joint representation, the attention weighted pooling is configured to reduce the dimensions of the joint representation, and the projection head is configured to map the dimensionally reduced joint representation to a space where contrastive loss is applied. Claim 10 The method of claim 1, wherein the machine learning model is or comprises a deep neural network, the deep neural network comprising a first encoder el(.), a second encoder e2(.),a second decoder d2(.), a fusion component f(.), an attention weighted pooling a(.), and a projection head p(.). The patent application cited above fails to teach the above bolded limitations at this claim, however, Udandarao teaches them at Fig. 2, and Chen teaches them at Paragraphs [0046-0047]. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with the combination of Udandarao and Chen in order to provide improved visual representations of multimodal data. Claim 11 The method of claim 1, wherein the training further comprises: computing a reconstruction loss for each reconstruction task, computing a contrastive loss for each discrimination task, computing a total loss on the basis of the reconstruction losses and the discrimination losses, and modifying parameters of the machine learning model so that the total loss is minimized. Claim 12 The method of claim 1,wherein training of the machine learning model is steered by a training loss LT, the loss LT comprising three parts: a reconstruction loss L; for the reconstruction of the data of first modality, a reconstruction loss L, for the reconstruction of the data of second modality, and a contrastive loss Lc: PNG media_image1.png 57 258 media_image1.png Greyscale wherein α, β, and ϒ are weighting factors. Claim 12 The patent application cited above fails to teach the limitations at this claim, however, Zhang teaches them at Paragraph [0022]. It would have been obvious to one of ordinary skill in the art before filing date to modify the teachings of Hoehne with Zhang in order increase the accuracy of similarity results involving multimodal data. Claim 13 The patent application cited above fails to teach the limitations at this claim, however, Lei teaches them at Paragraph [0024]. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with Lei in order to improve model efficiency and reduce computational costs. Claim 14 A computer system comprising: a processor; and a memory storing an application program configured to perform, when executed by the processor, an operation, the operation comprising: receiving input data related to a first object, inputting the input data into a trained machine learning model, receiving from the trained machine learning model a first representation of the first object, receiving at least one second representation of at least one second object, computing a similarity value, the similarity value indicating the similarity between the first representation and the at least one second representation, and outputting the similarity value and/or information related to the at least one second object; wherein the trained machine learning model was trained in a training process, the training process comprising the following steps: providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer; receiving training data for training the machine learning model, wherein providing training data comprises: receiving, for each object of a multitude of objects, input data of at least two different modalities, first input data of a first modality and second input data of a second modality, generating first augmented input data from the first input data and second augmented input data from the second input data, and generating first masked input data from the first augmented input data and second masked input data from the second augmented input data; training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input layer, inputting the second masked input data into the second input layer, reconstructing the first augmented input data from the first masked input data via the first output layer, reconstructing the second augmented input data from the second masked input data via the second output layer, generating a joint representation of the first masked input data and the second masked input data via the third output layer, and discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data of different objects. Claim 13 A computer system comprising: a processor; and a memory storing an application program configured to perform, when executed by the processor, an operation for training a machine learning model, the operation comprising: receiving, for each object of a multitude of objects, input data, comprising first input data of a first modality and second input data of a second modality, generating first augmented input data from the first input data and second augmented input data from the second input data, generating first masked input data from the first augmented input data and second masked input data from the second augmented input data, providing a machine learning model, the machine learning model comprising a first input a second input a first output; a second output; and a third output, training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input inputting the second masked input data into the second input; reconstructing the first augmented input data from the first masked input data via the first output; reconstructing the second augmented input data from the second masked input data via the second output; generating a joint representation of the first masked input data and the second masked input data via the third output; and discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data of different objects, and storing and/or outputting the trained machine learning model and/or providing the trained machine learning model for predictive purposes. The patent application cited above fails to teach the above bolded limitations at this claim, however, Zhang teaches them at Paragraph [0022]. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with Zhang in order to increase the accuracy of similarity results involving multimodal data. Claim 15 A non-transitory computer readable medium having stored thereon software instructions that, when executed by a processor of a computer system, cause the computer system to execute the following steps: receiving input data related to a first object, inputting the input data into a trained machine learning model, receiving from the trained machine learning model a first representation of the object, computing a similarity value, the similarity value indicating the similarity between the first representation and at least one second representation of at least one second object, outputting the similarity value and/or information related to the at least one second object, wherein the trained machine learning model was trained in a training process, the training process comprising the following steps: providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer; receiving training data for training the machine learning model, wherein providing training data comprises: receiving, for each object of a multitude of objects, input data of at least two different modalities, first input data of a first modality and second input data of a second modality, generating first augmented input data from the first input data and second augmented input data from the second input data, and generating first masked input data from the first augmented input data and second masked input data from the second augmented input data; training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input layer, inputting the second masked input data into the second input layer, reconstructing the first augmented input data from the first masked input data via the first output layer, reconstructing the second augmented input data from the second masked input data via the second output layer, generating a joint representation of the first masked input data and the second masked input data via the third output layer, and discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data of different objects. Claim 14 A non-transitory computer readable medium storing instructions that, when executed by a processor of a computer system, cause the computer system to: receive, for each object of a multitude of objects, input data of at least two different modalities, comprising first input data of a first modality and second input data of a second modality; generate first augmented input data from the first input data and second augmented input data from the second input data; generate first masked input data from the first augmented input data and second masked input data from the second augmented input data; provide a machine learning model, the machine learning model comprising; a first input, a second input, a first output, a second output, and a third output; train the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input, inputting the second masked input data into the second input, reconstructing the first augmented input data from the first masked input data via the first output, reconstructing the second augmented input data from the second masked input data via the second output, generating a joint representation of the first masked input data and the second masked discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data of different objects; and store and/or output the trained machine learning model and/or providing the trained machine learning model for predictive purposes. The patent application cited above fails to teach the above bolded limitations at this claim, however, Zhang teaches them at Paragraph [0022]. It would have been obvious to one of ordinary skill in the art before the filing date to modify the teachings of Hoehne with Zhang in order to increase the accuracy of similarity results involving multimodal data. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 1-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without reciting significantly more. Step 1 – is the claim directed to a process, machine, manufacture, or composition of matter? Claims 1-13 are directed to a “method” which describes one of the four statutory categories of patentable subject matter, i.e., a process. Claim 14 is directed to a “system” which describes one of the four statutory categories of patentable subject matter, i.e., a machine. Claim 15 is directed to a “non-transitory computer-readable medium” which describes one of the four statutory categories of patentable subject matter, i.e., a manufacture. Regarding Claim 1 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, Claim 1 recites an abstract idea, substantially as follows: “reconstructing the first augmented input data from the first masked input data via the first output layer,” – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)), as the process of “reconstructing” on data is an algorithm and is considered to be a mathematical calculation “reconstructing the second augmented input data from the second masked input data via the second output layer,” – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)), as the process of “reconstructing” on data is an algorithm and is considered to be a mathematical calculation “discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data of different objects;” – is directed to the abstract idea of a mental process i.e., discriminating is equivalent to classifying data after evaluation, which are concepts performed in the human mind at a high level (see MPEP 2106.04(a)(2)(III)(C)), and may be performed with the aid of pen and paper, or using a computer as a tool. “computing a similarity value, the similarity value indicating the similarity between the first representation and the at least one second representation; and” – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as it is describing computing a value, which is considered to be a mathematical calculation. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 1 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer; “ – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. “providing training data for training the machine learning model, wherein providing training data comprises: receiving, for each object of a multitude of objects, input data of at least two different modalities, first input data of a first modality and second input data of a second modality, “ – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. “generating first augmented input data from the first input data and second augmented input data from the second input data, and “ – is merely selecting a particular data source or type of data to be manipulated, considered insignificant extra-solution activity (see MPEP 2106.05(g)(3)). “generating first masked input data from the first augmented input data and second masked input data from the second augmented input data; “ – is merely selecting a particular data source or type of data to be manipulated, considered insignificant extra-solution activity (see MPEP 2106.05(g)(3)). “training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input layer, inputting the second masked input data into the second input layer,” – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “generating a joint representation of the first masked input data and the second masked input data via the third output layer, and” – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). “receiving input data related to a first object; “ – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “inputting the input data related to the first object into the trained machine learning model;” – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “receiving from the trained machine learning model a first representation of the first object via the third output layer;” – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). “receiving at least one second representation of at least one second object; “ – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “outputting the similarity value and/or information related to the at least one second object.” – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 1 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer; “ – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. “providing training data for training the machine learning model, wherein providing training data comprises: receiving, for each object of a multitude of objects, input data of at least two different modalities, first input data of a first modality and second input data of a second modality, “ – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “generating first augmented input data from the first input data and second augmented input data from the second input data, and “ – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). Further, the insignificant extra-solution data gathering is also WURC, see MPEP 2106.05(d)(III) “The courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. Iii. Electronic recordkeeping”. “generating first masked input data from the first augmented input data and second masked input data from the second augmented input data; “ – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). Further, the insignificant extra-solution data gathering is also WURC, see MPEP 2106.05(d)(III) “The courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. Iii. Electronic recordkeeping”. “training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input layer, inputting the second masked input data into the second input layer,” – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “generating a joint representation of the first masked input data and the second masked input data via the third output layer, and” – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “receiving input data related to a first object; “ – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “inputting the input data related to the first object into the trained machine learning model;” – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “receiving from the trained machine learning model a first representation of the first object via the third output layer;” – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “receiving at least one second representation of at least one second object; “ – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “outputting the similarity value and/or information related to the at least one second object.” – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 2 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? No, Claim 2 does not recite an abstract idea. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 2 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “wherein the first object, the at least one second object and each object of the multitude of objects is a human being.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 2 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “wherein the first object, the at least one second object and each object of the multitude of objects is a human being.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 3 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? No, Claim 3 does not recite an abstract idea. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 3 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “wherein the training data and the input data related to the first object comprise personal data, the personal data being selected from one or more of the following group: age, height, weight, gender, eye color, hair color, skin color, blood group, blood pressure, resting heart rate, heart rate variability, vagus nerve tone, hematocrit, sugar concentration in urine, existing illnesses, existing conditions, pre-existing illnesses, pre-existing conditions, eyesight, consumption of alcohol, smoking, exercise, diet, information from an electronic medical record, self-assessment data, medical image(s), sound(s) from: heartbeat, breathing noise, cough, swallow, sneeze, clear throat, scratch, voice, noises when knocking against part(s) of the body and/or joint noise.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 3 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “wherein the training data and the input data related to the first object comprise personal data, the personal data being selected from one or more of the following group: age, height, weight, gender, eye color, hair color, skin color, blood group, blood pressure, resting heart rate, heart rate variability, vagus nerve tone, hematocrit, sugar concentration in urine, existing illnesses, existing conditions, pre-existing illnesses, pre-existing conditions, eyesight, consumption of alcohol, smoking, exercise, diet, information from an electronic medical record, self-assessment data, medical image(s), sound(s) from: heartbeat, breathing noise, cough, swallow, sneeze, clear throat, scratch, voice, noises when knocking against part(s) of the body and/or joint noise.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 4 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? No, Claim 4 does not recite an abstract idea. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 4 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “wherein the first object, the at least one second object and each object of the multitude of objects is a plant or a plurality of plants or one or more parts of a plant.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 4 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “wherein the first object, the at least one second object and each object of the multitude of objects is a plant or a plurality of plants or one or more parts of a plant.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 5 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? No, Claim 5 does not recite an abstract idea. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 5 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “wherein the first object, the at least one second object and each object of the multitude of objects is a part of the Earth's surface.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 5 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “wherein the first object, the at least one second object and each object of the multitude of objects is a part of the Earth's surface.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 6 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? No, Claim 6 does not recite an abstract idea. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 6 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “wherein the first input data of the first modality and the second input data of the second modality comprise or are derived from one or more images, text files and/or audio files, wherein the first modality is different from the second modality.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 6 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “wherein the first input data of the first modality and the second input data of the second modality comprise or are derived from one or more images, text files and/or audio files, wherein the first modality is different from the second modality.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 7 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? No, Claim 7 does not recite an abstract idea. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 7 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “wherein the first object is different from the at least one second object.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 7 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “wherein the first object is different from the at least one second object.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 8 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? No, Claim 8 does not recite an abstract idea. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 8 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “wherein the first object is identical to the at least one second object, wherein the first representation is a representation of the first object at a first point in time and the at least one second representation represents the first object at least one second point in time. “ – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 8 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “wherein the first object is identical to the at least one second object, wherein the first representation is a representation of the first object at a first point in time and the at least one second representation represents the first object at least one second point in time. “ – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 9 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? No, Claim 9 does not recite an abstract idea. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 9 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “wherein the machine learning model comprises a number k of input layers, and a number k+1 output layers, wherein k is a natural number greater than two, and wherein each input layer is configured to receive input data of a different modality.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 9 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “wherein the machine learning model comprises a number k of input layers, and a number k+1 output layers, wherein k is a natural number greater than two, and wherein each input layer is configured to receive input data of a different modality.” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 10 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, Claim 10 recites an abstract idea, substantially as follows: “the first decoder is configured to reconstruct the first augmented input data from the joint representation, “ – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)), as the process of “reconstructing” on data is an algorithm and is considered to be a mathematical calculation “the second decoder is configured to reconstruct the second augmented input data from the joint representation, “ – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)), as the process of “reconstructing” on data is an algorithm and is considered to be a mathematical calculation “the attention weighted pooling is configured to reduce the dimensions of the joint representation, and” – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as weighted pooling is considered to be an algorithm, which is considered to be a mathematical calculation. “the projection head is configured to map the dimensionally reduced joint representation to a space where contrastive loss is applied. “ – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as mapping the joint representation to space is equivalent to a mapping function which is considered to be a mathematical relationship. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 10 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “wherein the machine learning model is or comprises a deep neural network, wherein the deep neural network comprises, at least for the training, a first encoder, a first decoder, a second encoder, a second decoder, a fusion component, an attention weighted pooling, and a projection head, wherein:” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. “the first encoder is configured to receive the first masked input data, and to generate the first representation from the first masked input data,“ – is merely a recitation of an insignificant extra-solution data receiving and outputting (see MPEP 2106.05(g)). “the second encoder is configured to receive the second masked input data, and to generate the second representation from the second masked input data,” – is merely a recitation of an insignificant extra-solution data receiving and outputting (see MPEP 2106.05(g)). “the fusion component is configured to generate the joint representation from the first representation and the second representation, “ – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 10 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “wherein the machine learning model is or comprises a deep neural network, wherein the deep neural network comprises, at least for the training, a first encoder, a first decoder, a second encoder, a second decoder, a fusion component, an attention weighted pooling, and a projection head, wherein:” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. “the first encoder is configured to receive the first masked input data, and to generate the first representation from the first masked input data,“ – the broadest reasonable interpretation of this imitation is found to be merely receiving and outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “the second encoder is configured to receive the second masked input data, and to generate the second representation from the second masked input data,” – the broadest reasonable interpretation of this imitation is found to be merely receiving and outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “the fusion component is configured to generate the joint representation from the first representation and the second representation, “ – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 11 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, Claim 11 recites an abstract idea, substantially as follows: “computing a reconstruction loss for each reconstruction task, “ – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as it is describing computing a loss, and is considered to be a mathematical calculation. “computing a contrastive loss for each discrimination task, “ – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as it is describing computing a loss, and is considered to be a mathematical calculation. “computing a total loss on the basis of the reconstruction losses and the discrimination losses, and” – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as it is describing computing a loss, and is considered to be a mathematical calculation. “modifying parameters of the machine learning model so that the total loss is minimized. “ – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as it is describing adjusting parameters in order to minimize loss, which is equivalent to optimization, and is considered to be a mathematical calculation. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 11 does not include additional limitations that integrate the judicial exception into a practical application. Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 11 does not include additional limitations that amount to significantly more than the judicial exception. Regarding Claim 12 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, Claim 12 recites an abstract idea, substantially as follows: “for each second object of a plurality of second objects, a similarity value is computed, the similarity value quantifying the similarity between a second representation of the second object and the first representation of the first object, wherein a number m of second objects is identified, the similarity values of the number m of second objects being greater than the similarity values of second objects not belonging to the number m of second objects, wherein m is a natural number greater than 0.” – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as it is describing computing a value, and is considered to be a mathematical calculation. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 12 does not include additional limitations that integrate the judicial exception into a practical application. Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 12 does not include additional limitations that amount to significantly more than the judicial exception. Regarding Claim 13 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, Claim 13 recites an abstract idea, substantially as follows: “for each second object of the number m of second objects, input data related to the second object is analyzed in order to identify data characterizing the second object that is not available for the first object. “ – is directed to the abstract idea of a mental process i.e., analysis or evaluation are concepts performed in the human mind (see MPEP 2106.04(a)(2)(III)(C)), and may be performed with the aid of pen and paper, or using a computer as a tool. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 13 does not include additional limitations that integrate the judicial exception into a practical application. Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 13 does not include additional limitations that amount to significantly more than the judicial exception. Regarding Claim 14 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, Claim 14 recites an abstract idea, substantially as follows: “computing a similarity value, the similarity value indicating the similarity between the first representation and the at least one second representation, and” – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as it is describing computing a value, and is considered to be a mathematical calculation. “reconstructing the first augmented input data from the first masked input data via the first output layer, reconstructing the second augmented input data from the second masked input data via the second output layer, “ – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)), as the process of “reconstructing” on data is considered to be an algorithm and is considered to be a mathematical calculation “discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data of different objects.” – is directed to the abstract idea of a mental process i.e., discriminating is equivalent to classifying data after evaluation, which are concepts performed in the human mind at a high level (see MPEP 2106.04(a)(2)(III)(C)), and may be performed with the aid of pen and paper, or using a computer as a tool. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 14 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “a processor; and a memory storing an application program configured to perform, when executed by the processor, an operation, the operation comprising: “ – is directed to merely applying an abstract idea using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.04(d)). “receiving input data related to a first object, inputting the input data into a trained machine learning model, receiving from the trained machine learning model a first representation of the first object, receiving at least one second representation of at least one second object, “ – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “outputting the similarity value and/or information related to the at least one second object;” – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). “wherein the trained machine learning model was trained in a training process, the training process comprising the following steps: providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer;” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. “receiving training data for training the machine learning model, wherein providing training data comprises: receiving, for each object of a multitude of objects, input data of at least two different modalities, first input data of a first modality and second input data of a second modality, “ – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “generating first augmented input data from the first input data and second augmented input data from the second input data, and generating first masked input data from the first augmented input data and second masked input data from the second augmented input data; – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). “training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input layer, inputting the second masked input data into the second input layer,” – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “generating a joint representation of the first masked input data and the second masked input data via the third output layer, and” – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 14 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “a processor; and a memory storing an application program configured to perform, when executed by the processor, an operation, the operation comprising: “ – is directed to merely applying an abstract idea using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.04(d)). “receiving input data related to a first object, inputting the input data into a trained machine learning model, receiving from the trained machine learning model a first representation of the first object, receiving at least one second representation of at least one second object, “ – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “outputting the similarity value and/or information related to the at least one second object;” – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “wherein the trained machine learning model was trained in a training process, the training process comprising the following steps: providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer;” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. “receiving training data for training the machine learning model, wherein providing training data comprises: receiving, for each object of a multitude of objects, input data of at least two different modalities, first input data of a first modality and second input data of a second modality, “ – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “generating first augmented input data from the first input data and second augmented input data from the second input data, and generating first masked input data from the first augmented input data and second masked input data from the second augmented input data; – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first masked input data into the first input layer, inputting the second masked input data into the second input layer,” – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “generating a joint representation of the first masked input data and the second masked input data via the third output layer, and” – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). Therefore, the additional elements, alone or in combination, do not amount to significantly more than the judicial exception (See MPEP 2106.05). Regarding Claim 15 Steps 2A Prong 1 – is the claim directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, Claim 15 recites an abstract idea, substantially as follows: “computing a similarity value, the similarity value indicating the similarity between the first representation and at least one second representation of at least one second object,” – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)) as it is describing computing a value, which is considered to be a mathematical calculation. “reconstructing the first augmented input data from the first masked input data via the first output layer, reconstructing the second augmented input data from the second masked input data via the second output layer, “ – is directed to the abstract idea of mathematical concepts (See MPEP 2106.04(a)(2)), as the process of “reconstructing” on data is considered to be an algorithm and is considered to be a mathematical calculation “discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data of different objects.” – is directed to the abstract idea of a mental process i.e., discriminating is equivalent to classifying data after evaluation, which are concepts performed in the human mind at a high level (see MPEP 2106.04(a)(2)(III)(C)), and may be performed with the aid of pen and paper, or using a computer as a tool. Step 2A Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No, Claim 15 does not include additional limitations that integrate the judicial exception into a practical application. The additional limitation(s): “receiving input data related to a first object, inputting the input data into a trained machine learning model, receiving from the trained machine learning model a first representation of the object, “ – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “outputting the similarity value and/or information related to the at least one second object,” – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). “wherein the trained machine learning model was trained in a training process, the training process comprising the following steps: providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer;” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. “receiving training data for training the machine learning model, wherein providing training data comprises: receiving, for each object of a multitude of objects, input data of at least two different modalities, first input data of a first modality and second input data of a second modality,” – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “generating first augmented input data from the first input data and second augmented input data from the second input data, and generating first masked input data from the first augmented input data and second masked input data from the second augmented input data;” – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). “inputting the first masked input data into the first input layer, inputting the second masked input data into the second input layer,” – is merely a recitation of an insignificant extra-solution data gathering (see MPEP 2106.05(g)). “generating a joint representation of the first masked input data and the second masked input data via the third output layer, and” – is merely a recitation of an insignificant extra-solution data outputting (see MPEP 2106.05(g)). Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application (See MPEP 2106.04). Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception? No, Claim 15 does not include additional limitations that amount to significantly more than the judicial exception. The additional limitation(s): “receiving input data related to a first object, inputting the input data into a trained machine learning model, receiving from the trained machine learning model a first representation of the object, “ – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “outputting the similarity value and/or information related to the at least one second object,” – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “wherein the trained machine learning model was trained in a training process, the training process comprising the following steps: providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer;” – is merely indicating a field of use or technological environment (see MPEP 2106.06(h)) and fails to integrate the judicial exception. “receiving training data for training the machine learning model, wherein providing training data comprises: receiving, for each object of a multitude of objects, input data of at least two different modalities, first input data of a first modality and second input data of a second modality,” – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “generating first augmented input data from the first input data and second augmented input data from the second input data, and generating first masked input data from the first augmented input data and second masked input data from the second augmented input data;” – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “inputting the first masked input data into the first input layer, inputting the second masked input data into the second input layer,” – the broadest reasonable interpretation of this imitation is found to be merely receiving data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). “generating a joint representation of the first masked input data and the second masked input data via the third output layer, and” – the broadest reasonable interpretation of this imitation is found to be merely outputting data, which is analogous to receiving or transmitting data over a network, considered WURC under MPEP2106.05(d)(II)(i). Therefore, the additional elements, alone or in combination, do not amount to more than the judicial exception (See MPEP 2106.05). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inqiuries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 2, 6, 11, 12, 14, and 15 are rejected under 35 U.S.C 103 as being unpatentable over NPL reference Udandarao et al. (“COBRA: Contrastive Bi-Modal Representation Algorithm”), hereinafter referred to as Udandarao, in view of Kim, Tae Hoon (US 11200494 B1), hereinafter referred to as Kim, and further in view of Zhang, Zhenzhong (US 20220076052 A1), hereinafter referred to as Zhang. Regarding Claim 1 Udandarao discloses: “providing a machine learning model, the machine learning model comprising: a first input layer, a second input layer, a first output layer, a second output layer, and a third output layer;” (Udandarao at Pg. 2, Figure 2 PNG media_image2.png 762 1417 media_image2.png Greyscale ) [Examiner Note: image encoder mapped to first input layer, text encoder mapped to second input layer, image decoder mapped to first output layer, text decoder mapped to second output layer, and orthogonal transform layer mapped to third output layer] “providing training data for training the machine learning model, wherein providing training data comprises: receiving, […], input data of at least two different modalities, first input data of a first modality and second input data of a second modality,” (Udandarao at Fig. 2: image modality, text modality; Udandarao at Abstract: In this paper, we present a novel framework COBRA that aims to train two modalities (i.e., image and text) in a joint fashion […]) [Examiner Note: Fig. 2 shows the model receives data of two modalities, image and text] “training the machine learning model to perform a combined reconstruction and discrimination task, the training comprising: inputting the first […] input data into the first input layer, inputting the second […] input data into the second input layer, reconstructing the first […] input data from the first […] input data via the first output layer, reconstructing the second […] input data from the second […] input data via the second output layer, a joint representation of the first […] input data and the second […] input data via the third output layer, and discriminating joint representations which were generated from input data of the same object from joint contrastive representations which were generated from input data […];” (Udandarao at Fig 2; Udandarao at 4.2.2 Evaluation Metrics: The figures clearly exhibit the high discrimination between samples of different classes in the joint embedding space) [Examiner Note: Fig. 2 shows the reconstruction of both the image and the text as outputs; Fig. 2 shows the discrimination of the joint representation in the joint embedding space] “receiving input data related to a first object;” (Udandarao at 4.1 Cross-Modal Retrieval: In the task of cross-modal retrieval, we use COBRA to retrieve an image given a text query, or a text sample given an image query.) “inputting the input data related to the first object into the trained machine learning model;” (Udandarao at Fig 2 [Examiner Note: both datasets related to the object are input into the model]) However, Udandarao does not disclose: “generating first augmented input data from the first input data and second augmented input data from the second input data, generating first masked input data from the first augmented input data and second masked input data from the second augmented input data;” “receiving at least one second representation of at least one second object;” “computing a similarity value, the similarity value indicating the similarity between the first representation and the at least one second representation; and” “outputting the similarity value and/or information related to the at least one second object.” On the other hand, Kim discloses: “generating first augmented input data from the first input data and second augmented input data from the second input data, generating first masked input data from the first augmented input data and second masked input data from the second augmented input data;” (Kim at Col 20, Lines 29-33: the obfuscation network O may be trained to obfuscate the input data with the augmentation such that a result O(T(x)) created by augmenting input data and then by obfuscating the augmented input data [Examiner Note: mapped to generating masked input data from augmented input data]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Udandarao with the above teachings of Kim by using a model with multiple input and output layers configured to receive data of different modalities, as taught by Udandarao, and performing augmentation and masking of data as taught by Kim. The modification would have been obvious because one of ordinary skill in the art would be motivated to obfuscate sensitive data in order to ensure data privacy as suggested by Kim at Col 2, Lines 34-37: “It is still another object of the present disclosure to protect privacy and security of original data by generating obfuscated data, e.g., anonymized data or concealed data, through irreversibly obfuscating the original data”. However, the combination of Udandarao and Kim does not disclose: “receiving at least one second representation of at least one second object;” “computing a similarity value, the similarity value indicating the similarity between the first representation and the at least one second representation; and” “outputting the similarity value and/or information related to the at least one second object.” On the other hand, Zhang discloses: “receiving at least one second representation of at least one second object;” (Zhang at [0022]: acquiring second data of a plurality of second objects;) “computing a similarity value, the similarity value indicating the similarity between the first representation and the at least one second representation; and” (Zhang at [0022]: calculating a similarity between the first object and each second object by using any data similarity determination method provided by at least one embodiment of the present disclosure to obtain a plurality of object similarities;) “outputting the similarity value and/or information related to the at least one second object.” (Zhang at [0022]: and outputting information of a second object corresponding to an object similarity with a largest value among the plurality of object similarities.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Udandarao and Kim with the above teachings of Zhang by using a model with multiple input and output layers configured to receive masked data of different modalities as taught by Udandarao and Kim, and using a method of computing a similarity score between two objects, as taught by Zhang. The modification would have been obvious because one of ordinary skill in the art would be motivated to increase the accuracy of relevant or most similar results involving multimodal data as suggested by Zhang at [0112]: “For example, by using any data similarity determination method provided by at least one embodiment of the present disclosure to calculate the similarity between the first object and each second object, the accuracy of the data similarity calculation result can be improved, and therefore, the searching performance of the similar object search method provided by at least one embodiment of the present disclosure can be improved (for example, more similar objects can be found).” Referring to independent Claims 14 and 15, they are rejected on the same basis as independent Claim 1, mutatis mutandis, since both are analogous claims. Regarding Claim 2 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. The combination of Udandarao, Kim, and Zhang further discloses: “wherein the first object, the at least one second object and each object of the multitude of objects is a human being.” (Udandarao at Fig. 1, Fig. 2) [Examiner Note: the figures in Udandarao show the input data represents a human being] Regarding Claim 6 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. The combination of Udandarao, Kim, and Zhang further discloses: “wherein the first input data of the first modality and the second input data of the second modality comprise or are derived from one or more images, text files and/or audio files, wherein the first modality is different from the second modality.” (Udandarao at Fig. 2 [the model receives two different modalities of data, image and text]; Udandarao Abstract: In this paper, we present a novel framework COBRA that aims to train two modalities (i.e., image and text) in a joint fashion) Regarding Claim 11 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. The combination of Udandarao, Kim, and Zhang further discloses: “computing a reconstruction loss for each reconstruction task, computing a contrastive loss for each discrimination task, computing a total loss on the basis of the reconstruction losses and the discrimination losses, and modifying parameters of the machine learning model so that the total loss is minimized. “ (Udandarao at 3.2.1 Reconstruction Loss; Udandarao at 3.2.4 Contrastive Loss; Udandarao at 3.2 Model Architecture: We define the loss function in COBRA as a weighted sum [Examiner Note: mapped to total loss] of the reconstruction loss, cross-modal loss, supervised loss and contrastive loss; Udandarao at 3.3 Optimization and Training Strategy: The objective function in Eq. 7 [Examiner Note: mapped to loss function] is optimized using stochastic gradient descent. The loss is summed over all modalities, and the corresponding gradient is propagated through all the components in the model. The optimization process of our proposed network is illustrated in Algorithm 1; Udandarao at Algorithm 1: PNG media_image3.png 527 398 media_image3.png Greyscale [Examiner Note: update model weights is mapped to modifying parameters of the machine learning model]) Regarding Claim 12 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. The combination of Udandarao, Kim, and Zhang further discloses: “for each second object of a plurality of second objects, a similarity value is computed, the similarity value quantifying the similarity between a second representation of the second object and the first representation of the first object, wherein a number m of second objects is identified, the similarity values of the number m of second objects being greater than the similarity values of second objects not belonging to the number m of second objects, wherein m is a natural number greater than 0.” (Zhang at [0022]: calculating a similarity between the first object and each second object by using any data similarity determination method provided by at least one embodiment of the present disclosure to obtain a plurality of object similarities; Zhang at [0194]: S440: outputting information of a second object corresponding to an object similarity with a largest value [Examiner Note: mapped to identifying a second object where the similarity value is greater than the objects that were not identified, in this case outputted] among the plurality of object similarities) Claims 4 and 5 are rejected under 35 U.S.C 103 as being unpatentable over NPL reference Udandarao in view of Kim in view of Zhang and further in view of NPL reference Gadiraju et al. “Multimodal Deep Learning Based Crop Classification Using Multispectral and Multitemporal Satellite Imagery”, hereinafter referred to as Gadiraju. Regarding Claim 4 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. However, the combination of Udandarao, Kim, and Zhang does not disclose: “wherein the first object, the at least one second object and each object of the multitude of objects is a plant or a plurality of plants or one or more parts of a plant.” On the other hand, Gadiraju discloses: “wherein the first object, the at least one second object and each object of the multitude of objects is a plant or a plurality of plants or one or more parts of a plant.” (Gadiraju at Pg. 3, 3 DATASET DESCRIPTION: We use the Normalized Difference Vegetation Index (NDVI) information from MODIS. Over 60,000 data points corresponding to the following crops: corn, cotton, soy, spring wheat, winter wheat and barley were collected from multiple locations at different time stamps across the USA.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Udandarao, Kim, and Zhang with the above teachings of Gadiraju by using a method that computes a similarity value between objects comprising of multiple modalities, as taught by Udandarao, Kim, and Zhang, and using data related to plants, as taught by Gadiraju. The modification would have been obvious because one of ordinary skill in the art would be motivated to improve computational performance with multimodal data as suggested by Gadiraju at Pg. 1, Col 2, Introduction: “Machine learning solutions for crop classification using the readily available large scale remote sensing imagery could provide a quick and cost effective solution for crop monitoring and global food estimation.” Regarding Claim 5 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. However, the combination of Udandarao, Kim, and Zhang does not disclose: “wherein the first object, the at least one second object and each object of the multitude of objects is a part of the Earth's surface.” On the other hand, Gadiraju discloses: “wherein the first object, the at least one second object and each object of the multitude of objects is a part of the Earth's surface.” (Gadiraju at Pg. 3, 3 DATASET DESCRIPTION: High temporal resolution: 250m spatial resolution, bi-weekly temporal resolution Moderate Resolution Imaging Spectro radiometer (MODIS) imagery. [Examiner Note: mapped to objects part of the Earth’s surface] We use the Normalized Difference Vegetation Index (NDVI) information from MODIS.) The same motivation that utilized for combining Udandarao, Kim, and Zhang with Gadiraju, as set forth in Claim 4, is equally applicable to Claim 5. Claims 7, 8, and 13 are rejected under 35 U.S.C 103 as being unpatentable over NPL reference Udandarao in view of Kim in view of Zhang and further in view of Lei et al. (US 20180211182 A1) hereinafter to referred to as Lei. Regarding Claim 7 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. However, the combination of Udandarao, Kim, and Zhang does not disclose: “wherein the first object is different from the at least one second object.” On the other hand, Lei discloses: “wherein the first object is different from the at least one second object.” (Lei at [0024]: In an aspect, the matrix generation component 104 can compare data associated with a corresponding time value from a first stream of time series data [Examiner Note: mapped to first object] and a second stream of time series data [Examiner Note: mapped to first object]) [Examiner Note: the two streams of data, first and second, are mapped to being different] It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Udandarao, Kim, and Zhang with the above teachings of Lei by using a method that computes a similarity value between objects comprising of multiple modalities, as taught by Udandarao, Kim, and Zhang, and using data related to different objects, as taught by Lei. The modification would have been obvious because one of ordinary skill in the art would be motivated to improve model efficiency and reduce computational costs as suggested by Lei at [0020]: “Accordingly, flexibility for processing and/or analyzing time series data can be provided, effectiveness of a machine learning model can be improved, efficiency of one or more processors that execute a machine learning model can be improved, and/or a greater number of machine learning models can be employed for time series data. Moreover, an amount of time to perform a machine learning process can be reduced and/or amount of processing required for a machine learning process can be reduced.” Regarding Claim 8 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. However, the combination of Udandarao, Kim, and Zhang does not disclose: “wherein the first object is identical to the at least one second object, wherein the first representation is a representation of the first object at a first point in time and the at least one second representation represents the first object at least one second point in time.” On the other hand, Lei discloses: “wherein the first object is identical to the at least one second object, wherein the first representation is a representation of the first object at a first point in time and the at least one second representation represents the first object at least one second point in time.” ([0031]: or example, the data matrix 202 can compare similarities between corresponding data pairs from the time series data 114.sub.1-N. A corresponding data pair can be, for example, a first data value from a first stream [Examiner Note: first object at a first point in time] of time series data and a second data value from a first stream [Examiner Note: first object at a least one second point in time] of time series data with a corresponding time value.) The same motivation that utilized for combining Udandarao, Kim, and Zhang with Lei, as set forth in Claim 7, is equally applicable to Claim 8. Regarding Claim 13 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. However, the combination of Udandarao, Kim, and Zhang does not disclose: “wherein, for each second object of the number m of second objects, input data related to the second object is analyzed in order to identify data characterizing the second object that is not available for the first object.” On the other hand, Lei discloses: “wherein, for each second object of the number m of second objects, input data related to the second object is analyzed in order to identify data characterizing the second object that is not available for the first object.” (Lei at [0024]: For instance, the matrix generation component 104 can determine a difference between first data associated a first stream of time series data that corresponds to a particular time value and second data associated with a second stream of time series data that corresponds to the particular time value. Furthermore, a data element of the data matrix can correspond to the difference between the first data associated with the first stream of time series data and the second data associated with the second stream of time series data). The same motivation that utilized for combining Udandarao, Kim, and Zhang with Lei, as set forth in Claim 7, is equally applicable to Claim 13. Claims 3 and 9 are rejected under 35 U.S.C 103 as being unpatentable over Udandarao in view of Kim in view of Zhang and further in view of NPL reference Qiu et al. “Battling Alzheimer’s Disease through Early Detection: A Deep Battling Alzheimer’s Disease through Early Detection: A Deep Multimodal Learning Approach Multimodal Learning Approach” hereinafter referred to as Qiu. Regarding Claim 3 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. However, the combination of Udandarao, Kim, and Zhang further does not disclose: “wherein the training data and the input data related to the first object comprise personal data, the personal data being selected from one or more of the following group: age, height, weight, gender, eye color, hair color, skin color, blood group, blood pressure, resting heart rate, heart rate variability, vagus nerve tone, hematocrit, sugar concentration in urine, existing illnesses, existing conditions, pre-existing illnesses, pre-existing conditions, eyesight, consumption of alcohol, smoking, exercise, diet, information from an electronic medical record, self-assessment data, medical image(s), sound(s) from: heartbeat, breathing noise, cough, swallow, sneeze, clear throat, scratch, voice, noises when knocking against part(s) of the body and/or joint noise.” On the other hand, Qiu discloses: “wherein the training data and the input data related to the first object comprise personal data, the personal data being selected from one or more of the following group: age, height, weight, gender, eye color, hair color, skin color, blood group, blood pressure, resting heart rate, heart rate variability, vagus nerve tone, hematocrit, sugar concentration in urine, existing illnesses, existing conditions, pre-existing illnesses, pre-existing conditions, eyesight, consumption of alcohol, smoking, exercise, diet, information from an electronic medical record, self-assessment data, medical image(s), sound(s) from: heartbeat, breathing noise, cough, swallow, sneeze, clear throat, scratch, voice, noises when knocking against part(s) of the body and/or joint noise.” (Qiu at Pg. 7, Experimental Study, Data: PNG media_image4.png 272 801 media_image4.png Greyscale ) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Udandarao, Kim, and Zhang with the above teachings of Qiu by using a method that computes a similarity value between objects comprising of multiple modalities, as taught by Udandarao, Kim, and Zhang, and using medical and demographic data related to patients, as taught by Qiu. The modification would have been obvious because one of ordinary skill in the art would be motivated to improve predictive performance using multi-modal data as suggested by Qiu at Pg. 2, Introduction: “Our preliminary results show that (1) incorporating all three modalities (image, text and genotype) leads to better predictive performance compared to using only a pair of modalities, confirming the complementary value derived from different modalities; (2) learning both modality-invariant representation and modality specific representations together (which is the novel architectural element in MIMSL) leads to improved performance compared to learning modality-invariant representation alone (as done in previous multimodal neural architectures); (3) the predictive performance of MIMSL surpasses that of state-of the-art multimodal learning models.” Regarding Claim 9 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. However, the combination of Udandarao, Kim, and Zhang does not disclose: “wherein the machine learning model comprises a number k of input layers, and a number k+1 output layers, wherein k is a natural number greater than two, and wherein each input layer is configured to receive input data of a different modality.” On the other hand, Qiu discloses: “wherein the machine learning model comprises a number k of input layers, and a number k+1 output layers, wherein k is a natural number greater than two, and wherein each input layer is configured to receive input data of a different modality.” (Qiu at Pg. 6, Paragraph 5: “ The entire network, comprising two encoders and decoders per modality, and the additional classification layer is trained as a single feedforward network. The network has three input layers (one per modality) and four output layers (three for reconstructing each modality and one for classification); Qiu at Pg. 5, Fig. 2 PNG media_image5.png 394 807 media_image5.png Greyscale ; Qiu at Pg. 6: Our architecture can be used for any number of input modalities.) The same motivation that utilized for combining Udandarao, Kim, and Zhang with Qiu, as set forth in Claim 3, is equally applicable to Claim 9. Claim 10 is rejected under 35 U.S.C 103 as being unpatentable over NPL reference Udandarao in view of Kim in view of Zhang and further in view of Chen et al. (US 20210319266 A1) hereinafter to referred to as Chen. Regarding Claim 10 The combination of Udandarao, Kim, and Zhang discloses, “The method of claim 1,” and the limitations are shown in the rejection above. The combination of Udandarao, Kim, and Zhang further discloses: “wherein the machine learning model is or comprises a deep neural network, wherein the deep neural network comprises, at least for the training, a first encoder, a first decoder, a second encoder, a second decoder, a fusion component, […], wherein: the first encoder is configured to receive the first masked input data, and to generate the first representation from the first masked input data, the second encoder is configured to receive the second masked input data, and to generate the second representation from the second masked input data, the fusion component is configured to generate the joint representation from the first representation and the second representation, the first decoder is configured to reconstruct the first augmented input data from the joint representation, the second decoder is configured to reconstruct the second augmented input data from the joint representation, […]” (Udandarao at Fig. 2) However, the combination of Udandarao, Kim, and Zhang does not disclose: “attention weighted pooling, and a projection head […] wherein the attention weighted pooling is configured to reduce the dimensions of the joint representation, and wherein the projection head is configured to map the dimensionally reduced joint representation to a space where contrastive loss is applied.” On the other hand, Chen discloses: “attention weighted pooling, and a projection head […] wherein the attention weighted pooling is configured to reduce the dimensions of the joint representation, and wherein the projection head is configured to map the dimensionally reduced joint representation to a space where contrastive loss is applied.” (Chen at [0046] Some example implementations opt for simplicity and adopt the ResNet architecture (He et al., 2016) to obtain […] the output after the average pooling layer [Examiner Note: average pooling layer is mapped to attention weight pooling as its function is equivalent to reducing dimensions of a representation]; Chen at [0047]: A projection head neural network 206 (represented in the notation herein as g(⋅)) that maps the intermediate representations to final representations within the space where contrastive loss is applied). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Udandarao, Kim, and Zhang with the above teachings of Chen by utilizing a fusion component to generate a joint representation with encoders and decoders taught by Udandarao, Kim, and Zhang, and using attention weighted pooling in combination with a project head component, as taught by Chen. The modification would have been obvious because one of ordinary skill in the art would be motivated to provide improved visual representations of data as suggested by Chen at [0002]: “More particularly, the present disclosure relates to contrastive learning frameworks that leverage data augmentation and a learnable nonlinear transformation between the representation and the contrastive loss to provide improved visual representations.” Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20150310177 A1 – Recites methods of computing similarities between medical cases to retrieve relevant/similar medical cases to aid in patient diagnosis. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAMIYAH KABIR whose telephone number is (571)270-0722. The examiner can normally be reached Monday-Thursday 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SAMIYAH KABIR/Examiner, Art Unit 2126 /DAVID YI/Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Feb 01, 2024
Application Filed
Aug 20, 2026
Non-Final Rejection mailed — §101, §103, §DOUBLEPATENT (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month