DETAILED ACTION
This Office Action is in response to the amendments filed on 02/10/2026.
Claim 15 is currently amended.
Claims 1-15 are currently pending in this application and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
In reference to Applicant’s arguments on page(s) 9-13 regarding rejections made under 35 U.S.C. 101:
Claims 1-15 are rejected under 35 U.S.C. § 101 because the claimed invention is allegedly directed to non-statutory subject matter. Applicant respectfully disagrees.
Applicant respectfully disagrees. MPEP 2106.05(g) describes insignificant extra solution activity as "activities incidental to the primary process or product that are merely a nominal or tangential addition to the claim". In contrast, the above-emphasized claim features provide meaningful limits to the claim.
With regard to the independent claims, Applicant notes that the claims are directed to server-directed requests of specific metadata from plural sites, and that further determines a measure of the variation of such data to ensure that the model is trained from the metadata from these sites in a way that increases the heterogeneity of the dataset. In other words, the requesting and training of the data are not incidental to the claims, but instead, are used with the heterogeneity features of the claim to provide improved robustness to the model while preserving data privacy and reduced latency. As set forth in the August 4, 2025 memorandum ("Reminders on evaluating subject matter eligibility of claims under 35 U.S.C. 101"), the Step 2A Prong Two analysis is to consider the claims as a whole, and that the way that the additional elements use or integrate with the judicial exception into a practical application. As Applicant submits that the additional elements impose meaningful limits on the judicial exception, the rejection should be withdrawn under Step 2A Prong Two. That is, Applicant respectfully submits that the independent claims 1-15 are directed to a practical application under Step 2A, Prong Two, and requests that the rejection of claims 1-15 be withdrawn.
Examiner’s response:
Applicant’s arguments have been fully considered but are found to be not persuasive.
Applicant argues that the limitation of requesting metadata from clinical sites is not insignificant extra solution activity. Examiner disagrees. The limitation in question falls under one of the specified examples of extra-solution activity of receiving data i.e., pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory").
Applicant argues that the requesting of data and training of the model are not incidental to the claims. Examiner disagrees. Since the independent claims are directed toward a method for selecting training data to be used for training a model, the requesting of the data is seen as extra-solution activity, as mentioned above, and the limitation of using the selected data to train a model is directed toward using a generic computer component to apply the judicial exception, as no information about how the model is trained is presented in the independent claims.
Applicant argues that the claims present a practical application of the judicial exceptions. Examiner disagrees. The present disclosure does not present any novel way of selecting data with which to train a machine learning model, distributed or not. Since the claims rely on abstract ideas of making determinations and selecting data based on a determined measure, there is no practical application of the judicial exception.
In light of the arguments presented, the rejections made under 35 U.S.C. 101 are maintained and updated below.
In reference to Applicant’s arguments on page(s) 13-16 regarding rejections made under 35 U.S.C. 102:
Claims 1, 2, 4-5, 14, and 15 are rejected under 35 U.S.C. 102(a)(1) as allegedly anticipated by Chang et al.'s "Distributed deep learning networks among institutions for medical imaging" (hereinafter, "Chang"). Applicant respectfully traverses this rejection.
Applicant respectfully submits that independent claim 1 is allowable for at least the reason that Chang fails to disclose, teach, or suggest at least "determining, from the metadata, a measure of variation of the features of the matching data" and "based on the measure of variation, selecting training data for the training dataset from the matching data using the metadata so as to increase the heterogeneity of the training datase'. Though unclear, the non-final Office Action appears to equate the labels binarized to healthy and diseased to the metadata (see page 27 of the non-final Office Action), and despite referencing the paragraph from Chang associated with introducing variability, appears to equate the variation to the image resolution (see page 28 of the non-final Office Action). Applicant respectfully disagrees. Initially, it is noted that the claimed "determining from the metadata" refers to is metadata from each clinical site. That is, determining, from the metadata, a measure of variation is determining the variation within the metadata from each site among plural sites. The resolutions, on the other hand, appear to be the same amongst all of the sites. Further, Chang explicitly discloses that the variability is introduced in only the data of one of the sites, which is not the same as determining from the metadata the measure of variation (for each of plural sites).
For at least the reason that a prima facie case of anticipation has not been established for the above-emphasized claim features, Applicant respectfully requests that the rejection be withdrawn and the claim allowed.
Claims 2 and 4-5 depend from claim 1 and inherit all of the respective features of claim 1. Thus, claims 2 and 4-5 are patentable over Chang for at least the same reasons discussed above with respect to claim 1, with claims 2 and 4-5 containing further distinguishing patentable features.
For similar reasons presented above for claim 1, Applicant respectfully submits that independent claim 14 is allowable for at least the reason that Chang fails to disclose, teach, or suggest at least "determine, from the metadata, a measure of variation of the features of the matching data" and "based on the measure of variation, select training data for the training dataset from the matching data using the metadata so as to increase the heterogeneity of the training dataset'. For at least the reason that a prima facie case of anticipation has not been established for the above-emphasized claim features, Applicant respectfully requests that the rejection be withdrawn and the claim allowed.
For similar reasons presented above for claim 1, Applicant respectfully submits that independent claim 15, as amended, is allowable for at least the reason that Chang fails to disclose, teach, or suggest at least "determine, from the metadata, a measure of variation of the features of the matching data" and "based on the measure of variation, select training data for the training dataset from the matching data using the metadata so as to increase the heterogeneity of the training dataset'. For at least the reason that a prima facie case of anticipation has not been established for the above- emphasized claim features, Applicant respectfully requests that the rejection be withdrawn and the claim allowed.
Examiner’s response:
Applicant’s arguments have been fully considered but are found to be not persuasive.
Applicant argues that the sited reference of Chang does not teach the claims relating to determining a measure of variation of the metadata and selecting training data based on the determined variation. Examiner disagrees. While resolution might not be the exact metadata measured in the instant application, it is a type of metadata. The rejection of Claim 1 recites various types of metadata (labels, resolution, conditions under which the images were acquired, etc.). Chang also presents a measure of variation of the features from the metadata in Table 1 regarding the introduction of an institution with variability, specifically that in the experiment there are four institutions presented, three of which have good resolution, and the final having bad, low resolution. Based on the variation of resolution among the institutions, we can select only the training data of the sites with good resolution, and not the training data associated with the site with bad resolution. Similar reasoning applies to independent claims 14 and 15 as well.
In light of the arguments presented, the rejections made under 35 U.S.C. 102 are maintained and updated below.
In reference to Applicant’s arguments on page(s) 16-20 regarding rejections made under 35 U.S.C. 103:
Claims 3 and 7-10 have been rejected under §103(a) as allegedly obvious over Chang in view of Srinivasa. Applicant respectfully traverses this rejection. The addition of Srinivasa does not cure the deficiencies of Chang discussed above in connection with
independent claim 1. Therefore, claim 1 is considered patentable under any combination of these references. Furthermore, since independent claim 1 is allowable for at least the reasons discussed above, Applicant respectfully submits that claims 3 and 7-10 are allowable for the additional and separate reason that each depends from an allowable claim (See, e.g., In re Fine, 837 F.2d 1071, 5 U.S.P.Q. 2d 1596 (Fed. Cir. 1988)), with claims 3 and 7-10 containing further distinguishing patentable features. Therefore, Applicant respectfully requests that the rejection of claims 3 and 7-10 be withdrawn and the claims allowed.
Claim 6 has been rejected under §103(a) as allegedly obvious over Chang in view of Luca. Applicant respectfully traverses this rejection. The addition of Luca does not cure the deficiencies of Chang discussed above in connection with independent claim 1. Therefore, claim 1 is considered patentable under any combination of these references. Furthermore, since independent claim 1 is allowable for at least the reasons discussed above, Applicant respectfully submits that claim 6 is allowable for the additional and separate reason that claim 6 depends from an allowable claim (See, e.g., In re Fine, 837 F.2d 1071, 5 U.S.P.Q. 2d 1596 (Fed. Cir. 1988)), with claim 6 containing further distinguishing patentable features. Therefore, Applicant respectfully requests that the rejection of claim 6 be withdrawn and the claim allowed.
Claims 11 and 12 have been rejected under §103(a) as allegedly obvious over Chang in view of Srinivasa and Mathworks. Applicant respectfully traverses this rejection. The addition of Srinivasa and Mathworks does not cure the deficiencies of Chang discussed above in connection with independent claim 1. Therefore, claim 1 is considered patentable under any combination of these references. Furthermore, since independent claim 1 is allowable for at least the reasons discussed above, Applicant respectfully submits that claims 11-12 are allowable for the additional and separate reason that each depends from an allowable claim (See, e.g., In re Fine, 837 F.2d 1071, 5 U.S.P.Q. 2d 1596 (Fed. Cir. 1988)), with claims 11-12 containing further distinguishing patentable features. Therefore, Applicant respectfully requests that the rejection of claims 11-12 be withdrawn and the claims allowed.
Claim 13 has been rejected under §103(a) as allegedly obvious over Chang in view of Tong. Applicant respectfully traverses this rejection. The addition of Tong does not cure the deficiencies of Chang discussed above in connection with independent claim 1. Therefore, claim 1 is considered patentable under any combination of these references. Furthermore, since independent claim 1 is allowable for at least the reasons discussed above, Applicant respectfully submits that claim 13 is allowable for the additional and separate reason that claim 13 depends from an allowable claim.
Examiner’s response:
Applicant’s arguments have been fully considered but are found to be not persuasive.
Applicant argues that the dependent claims are deemed to be allowable since they depend on an allowable independent claim. Examiner disagrees. As mentioned above in the section relating to rejections made under 35 U.S.C. 102, Chang is found to teach all the limitations of the independent claims, therefore the dependent claims do not depend on an allowable independent claim.
In light of the arguments presented, the rejections made under 35 U.S.C. 103 are maintained and updated below.
Claim Rejections - 35 USC § 101
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claims 1-15 rejected under 35 U.S.C. 101 because they are directed toward an abstract idea without significantly more.
Step 1 analysis:
Independent Claim 1 recites, in part, a computer implemented method, therefore falling into the statutory category of process. Independent Claim 14 recites, in part, an apparatus, therefore falling into the statutory category of machine. Independent Claim 15 recites, in part, a non-transitory computer readable medium storing computer readable code, therefore falling into the statutory category of manufacture.
Regarding Claim 1:
Step 2A: Prong 1 analysis:
Claim 1 recites in part:
“selecting a training dataset with which to train a model using a distributed machine learning process, wherein the training dataset comprises medical data that satisfies one or more clinical requirements and wherein training data in the training dataset is located at a plurality of clinical sites”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses selecting data to train a model.
“determining, from the metadata, a measure of variation of the features of the matching data”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses determining variance in metadata.
“based on the measure of variation, selecting training data for the training dataset from the matching data using the metadata so as to increase the heterogeneity of the training dataset”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses selecting data to train a model, based on measure metadata.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of:
“requesting from each of the clinical sites, metadata describing features of matching data at the respective clinical site that satisfies the one or more clinical requirements”. This additional elements is recited at a high level of generality and amounts to extra-solution activity of gathering data i.e. pre-solution activity of gathering data for use in the claimed process.
“training the model using the selected training data according to the distributed learning process”. This additional element is recited at a high level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (machine learning model) (See MPEP 2106.05(f)).
Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
The additional element(s) of “requesting from each of the clinical sites, metadata describing features of matching data at the respective clinical site that satisfies the one or more clinical requirements” is/are recited at a high level of generality and amount(s) to extra-solution activity of receiving data i.e., pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory").
As discussed above, the additional element(s) of “training the model using the selected training data according to the distributed learning process” is/are recited at a high-level of generality such that it/they amount(s) to no more than mere instructions to apply the exception using generic computer components (machine learning model) (See MPEP 2106.05(f)).
Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception.
Regarding Claim 2:
Step 2A: Prong 1 analysis:
Claim 2 recites in part:
“selecting training data from the matching data at the plurality of clinical sites so as to increase the measure of variation of the features in the resulting training dataset”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses selecting data to train a model.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding Claim 3:
Step 2A: Prong 1 analysis:
Claim 3 recites in part:
“wherein the measure of variation of the features of the matching data is determined for matching data at each respective clinical site”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses measuring variance in data.
“selecting the training data from the matching data at the respective clinical site so as to increase the measure of variation of the training data selected from the respective clinical site”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses selecting data to train a model.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding Claim 4:
Step 2A: Prong 1 analysis:
Claim 4 recites in part:
“wherein the measure of variation of the features of the matching data is determined for matching data across all of the plurality of clinical sites”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses measuring variance in data.
“selecting the training data from the matching data across all the plurality of clinical sites so as to increase the measure of variation across the training dataset as a whole”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses selecting data to train a model.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding Claim 5:
Step 2A: Prong 1 analysis:
Claim 5 recites in part:
“selecting the training data from the matching data at the plurality of clinical sites to contain an even representation of different data types in the training dataset”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses selecting data to train a model.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding Claim 6:
Step 2A: Prong 1 analysis:
Claim 6 recites in part:
“supplementing the training dataset with augmented training data so as to increase the variation of the selected training dataset”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses adding data to a dataset.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding Claim 7:
Step 2A: Prong 1 analysis:
Claim 7 recites in part:
“wherein the measure of variation comprises a measure of heterogeneity”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses determining a measure in data.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding Claim 8:
Step 2A: Prong 2 analysis:
The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of:
“wherein the measure of heterogeneity is determined using a second machine learning model that takes the features in the metadata as input and outputs the measure of heterogeneity”. This additional element is recited at a high level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (machine learning model) (See MPEP 2106.05(f)).
Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
As discussed above, the additional element(s) of “wherein the measure of heterogeneity is determined using a second machine learning model that takes the features in the metadata as input and outputs the measure of heterogeneity” is/are recited at a high-level of generality such that it/they amount(s) to no more than mere instructions to apply the exception using generic computer components (machine learning model) (See MPEP 2106.05(f)).
Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception.
Regarding Claim 9:
Step 2A: Prong 2 analysis:
The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of:
“wherein the second machine learning model outputs a list comprising a subset of the matching data that optimizes or maximizes the heterogeneity compared to other possible subsets of the matching training data”. This additional elements is recited at a high level of generality and amounts to extra-solution activity of receiving data i.e. post-solution activity of outputting/displaying data for use in the claimed process.
“wherein the second machine learning model outputs a list comprising a subset of the matching data for which the heterogeneity is within a predetermined tolerance limit”. This additional elements is recited at a high level of generality and amounts to extra-solution activity of receiving data i.e. post-solution activity of outputting/displaying data for use in the claimed process.
Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
The additional element(s) of “wherein the second machine learning model outputs a list comprising a subset of the matching data that optimizes or maximizes the heterogeneity compared to other possible subsets of the matching training data” and “wherein the second machine learning model outputs a list comprising a subset of the matching data for which the heterogeneity is within a predetermined tolerance limit” is/are recited at a high level of generality and amount(s) to extra solution activity because it/they is/are a mere nominal or tangential addition to the claim, amounting to mere data output (see MPEP 2106.05(g)). The courts have similarly found limitations directed to displaying/outputting a result, recited at a high level of generality, to be well-understood, routine, and conventional. See (MPEP 2106.05(d)(II), "presenting offers and gathering statistics.", “determining an estimated outcome and setting a price”).
Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception.
Regarding Claim 10:
Step 2A: Prong 1 analysis:
Claim 10 recites in part:
“wherein the second machine learning model takes as input feature values in the metadata, xn, and determines a combination of xn that optimally flattens a linear function f(x)=x'3+b, wherein p comprises the gradient of the function and b comprises an offset”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses optimizing a linear function.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding Claim 11:
Step 2A: Prong 1 analysis:
Claim 11 recites in part:
“the second machine learning model is configured to determine f(x) with the minimal norm value (β′β) according to a convex optimization problem wherein the function: J(β)=1/2(β′β) is to be minimized”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses optimizing a linear function.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding Claim 12:
Step 2A: Prong 2 analysis:
The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of:
“wherein the second machine learning model comprises a support vector regression model”. This additional element is recited at a high level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (support vector machine) (See MPEP 2106.05(f)).
Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
As discussed above, the additional element(s) of “wherein the second machine learning model comprises a support vector regression model” is/are recited at a high-level of generality such that it/they amount(s) to no more than mere instructions to apply the exception using generic computer components (support vector machine) (See MPEP 2106.05(f)).
Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception.
Regarding Claim 13:
Step 2A: Prong 1 analysis:
Claim 13 recites in part:
“combining the results of the training according to the distributed learning process”. Under the broadest reasonable interpretation, this limitation covers a mental process including an observation, evaluation, judgment or opinion that could be performed in the human mind or with the aid of pencil and paper. See MPEP 2106.04(a)(2)(III). As drafted, this limitation encompasses combining data.
Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea.
Step 2A: Prong 2 analysis:
The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of:
“instructing each clinical site in the plurality of clinical sites to create a local copy of the model and train the local copy of the model using the training data in the training dataset selected from the respective clinical site”. This additional elements is recited at a high level of generality and amounts to extra-solution activity of gathering data i.e. pre-solution activity of gathering data for use in the claimed process.
Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
The additional element(s) of “instructing each clinical site in the plurality of clinical sites to create a local copy of the model and train the local copy of the model using the training data in the training dataset selected from the respective clinical site” is/are recited at a high level of generality and amount(s) to extra-solution activity of receiving data i.e., pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory").
Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception.
Regarding Claim 14:
Due to claim language similar to that of Claim 1, Claim 14 is rejected for the same reasons as presented above in the rejection of Claim 1, with the exception of the limitation(s) covered below.
Step 2A: Prong 2 analysis:
The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of:
“a memory comprising instruction data representing a set of instructions”. This additional element is recited at a high level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (memory) (See MPEP 2106.05(f)).
“a processor configured to communicate with the memory and to execute the set of instructions”. This additional element is recited at a high level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (processor) (See MPEP 2106.05(f)).
Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B analysis:
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
As discussed above, the additional element(s) of “a memory comprising instruction data representing a set of instructions” and “a processor configured to communicate with the memory and to execute the set of instructions” is/are recited at a high-level of generality such that it/they amount(s) to no more than mere instructions to apply the exception using generic computer components (memory and processor) (See MPEP 2106.05(f)).
Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception.
Regarding Claim 15:
Due to claim language similar to that of Claims 1 and 14, Claim 15 is rejected for the same reasons as presented above in the rejection of Claims 1 and 14.
Claim Rejections - 35 USC § 102
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claim(s) 1, 2, 4, 5, 14, and 15 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chang et al (Chang, Ken & Balachandar, Niranjan & Lam, Carson & Yi, Darvin & Brown, James & Beers, Andrew & Rosen, Bruce & Rubin, Daniel & Kalpathy-Cramer, Jayashree. (2018). Distributed deep learning networks among institutions for medical imaging. Journal of the American Medical Informatics Association. 25. 945-954. 10.1093/jamia/ocy017., hereinafter Chang).
Regarding Claim 1:
Chang teaches
A computer implemented method in a central server of selecting a training dataset with which to train a model using a distributed machine learning process, wherein the training dataset comprises medical data that satisfies one or more clinical requirements and wherein training data in the training dataset is located at a plurality of clinical sites (Chang [Page 946, Introduction, par. 5]: “In this study, we simulate the distribution of deep learning models across institutions using various nonparallel training heuristics”; [Page 946, Preprocessing, par. 1]: “We obtained 35 126 color digital retinal fundus (interior surface of the eye) photos from the Kaggle Diabetic Retinopathy competition.13 Each image was rated for disease severity by a licensed clinician on a scale of 0–4 (absent, mild, moderate, severe, and proliferative retinopathy, respectively). The images came from 17 563 patients of multiple primary care sites throughout California and elsewhere”)
requesting from each of the clinical sites, metadata describing features of matching data at the respective clinical site that satisfies the one or more clinical requirements (Chang [Page 946, Preprocessing, par. 1]: “The acquisition conditions were varied, with a range of camera models, levels of focus, and exposures. In addition, the resolutions ranged from 433 289 pixels to 5184 3456 pixels…To simplify training of the network, the labels were binarized to Healthy (scale 0) and Diseased (scale 2, 3, or 4)… It is also known that there is a correlation between the disease status of the left eye and the status of the right eye. To remove this as a confounding factor in our study, only images from left eye were utilized”)
determining, from the metadata, a measure of variation of the features of the matching data (Chang [Page 948, Introduction of an institution with variability, par. 1]: “In our initial division of the different institutions, we assumed that each institution had the same number of patients, ratio of healthy to diseased patients, and image quality. However, in a real scenario, there will likely be variability within institutions that may compromise the predictive performance of the model. To simulate this possibility, we introduced variability into one of the 4 institutions and assessed the performance of the different training heuristics. We simulated two scenarios: In the first, we decreased the resolution of the images by a factor of 16. In the second, we significantly decreased the number of patients (from n = 1500 to n = 150) and introduced class imbalance (ratio of healthy to diseased was 9:1)”)
based on the measure of variation, selecting training data for the training dataset from the matching data using the metadata so as to increase the heterogeneity of the training dataset (Chang [Page 950, Introduction of an institution with variability, par. 1]: “We next addressed what would happen if variability was introduced into one of the institutions. The modes of variability were either an institution with low-resolution images or an institution with few patients and class-imbalance…We also assessed performance of cyclical weight transfer when the variable institution was skipped. The resulting testing accuracy was 74.4%, which is comparable to cyclical weight transfer that included the variable institution”);
training the model using the selected training data according to the distributed learning process (Chang [Page 948, Model training heuristics with 4 institutions, par. 2]: “The last heuristic was training a model at each institution for a predetermined number of epochs (weight transfer frequency) before transferring the model to the next institution (cyclical weight transfer, Figure 2D)”).
Regarding Claim 2:
Chang teaches
wherein the step of selecting training data for the training dataset from the matching data using the metadata comprises: selecting training data from the matching data at the plurality of clinical sites so as to increase the measure of variation of the features in the resulting training dataset (Chang [Page 948, Introduction of an institution with variability, par. 1]: “In our initial division of the different institutions, we assumed that each institution had the same number of patients, ratio of healthy to diseased patients, and image quality. However, in a real scenario, there will likely be variability within institutions that may compromise the predictive performance of the model. To simulate this possibility, we introduced variability into one of the 4 institutions and assessed the performance of the different training heuristics. We simulated two scenarios: In the first, we decreased the resolution of the images by a factor of 16. In the second, we significantly decreased the number of patients (from n = 1500 to n = 150) and introduced class imbalance (ratio of healthy to diseased was 9:1)”)
Regarding Claim 4:
Chang teaches
The method of claim 1 wherein the measure of variation of the features of the matching data is determined for matching data across all of the plurality of clinical sites (Chang [Page 948, Introduction of an institution with variability, par. 1]: “In our initial division of the different institutions, we assumed that each institution had the same number of patients, ratio of healthy to diseased patients, and image quality. However, in a real scenario, there will likely be variability within institutions that may compromise the predictive performance of the model. To simulate this possibility, we introduced variability into one of the 4 institutions and assessed the performance of the different training heuristics. We simulated two scenarios: In the first, we decreased the resolution of the images by a factor of 16. In the second, we significantly decreased the number of patients (from n = 1500 to n = 150) and introduced class imbalance (ratio of healthy to diseased was 9:1)”)
and wherein the step of selecting training data for the training dataset comprises: selecting the training data from the matching data across all the plurality of clinical sites so as to increase the measure of variation across the training dataset as a whole (Chang [Page 948, Introduction of an institution with variability, par. 1]: “In our initial division of the different institutions, we assumed that each institution had the same number of patients, ratio of healthy to diseased patients, and image quality. However, in a real scenario, there will likely be variability within institutions that may compromise the predictive performance of the model. To simulate this possibility, we introduced variability into one of the 4 institutions and assessed the performance of the different training heuristics. We simulated two scenarios: In the first, we decreased the resolution of the images by a factor of 16. In the second, we significantly decreased the number of patients (from n = 1500 to n = 150) and introduced class imbalance (ratio of healthy to diseased was 9:1)”).
Regarding Claim 5:
Chang teaches
The method of claim 1 wherein the step of selecting training data for the training dataset from the matching data using the metadata so as to increase the heterogeneity of the training dataset compared to the matching data set comprises: selecting the training data from the matching data at the plurality of clinical sites to contain an even representation of different data types in the training dataset (Chang [Page 946, Preprocessing, par. 1]: “To simplify training of the network, the labels were binarized to Healthy (scale 0) and Diseased (scale 2, 3, or 4)… It is also known that there is a correlation between the disease status of the left eye and the status of the right eye. To remove this as a confounding factor in our study, only images from left eye were utilized”; [Page 952, Cyclical weight transfer with 20 institutions, par. 1]: “We next addressed whether cyclical weight transfer can improve model performance when the performance of any individual institution is no better than random classification. To do this, we divided 6000 patient samples from the Kaggle Diabetic Retinopathy dataset into 20 institutions (n = 300 per institution) with equal class distributions”)
Regarding Claim 14:
Due to claim language similar to that of Claim 1, Claim 14 is rejected for the same reasons as presented above in the rejection of Claim 1, with the exception of the limitation(s) covered below.
Chang teaches
a memory comprising instruction data representing a set of instructions (Chang [Page 946, Convolutional neural network, par. 1]: “We utilized the 34-layer residual network (ResNet34) architecture (Figure 1A). Our implementation was based on the Keras package with Theano backend. The convolutional neural networks were run on a NVIDIA Tesla P100 Graphics Processing Unit”)
a processor configured to communicate with the memory and to execute the set of instructions (Chang [Page 946, Convolutional neural network, par. 1]: “We utilized the 34-layer residual network (ResNet34) architecture (Figure 1A). Our implementation was based on the Keras package with Theano backend. The convolutional neural networks were run on a NVIDIA Tesla P100 Graphics Processing Unit”)
Regarding Claim 15:
Due to claim language similar to that of Claims 1 and 14, Claim 15 is rejected for the same reasons as presented above in the rejection of Claims 1 and 14.
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claim(s) 3 and 7-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chang as applied to claims 1, 14, and 15 above, and further in view of Srinivasa et al (Srinivasa, R. S., Qian, C., Theodorou, B., Spaeder, J., Xiao, C., Glass, L., & Sun, J. (2022). Clinical trial site matching with improved diversity using fair policy learning. arXiv [Cs.LG]. Retrieved from http://arxiv.org/abs/2204.06501, hereinafter Srinivasa).
Regarding Claim 3:
Chang teaches
and wherein the step of selecting training data for the training dataset comprises:
selecting the training data from the matching data at the respective clinical site so as to increase the measure of variation of the training data selected from the respective clinical site (Chang [Page 948, Introduction of an institution with variability, par. 1]: “In our initial division of the different institutions, we assumed that each institution had the same number of patients, ratio of healthy to diseased patients, and image quality. However, in a real scenario, there will likely be variability within institutions that may compromise the predictive performance of the model. To simulate this possibility, we introduced variability into one of the 4 institutions and assessed the performance of the different training heuristics. We simulated two scenarios: In the first, we decreased the resolution of the images by a factor of 16. In the second, we significantly decreased the number of patients (from n = 1500 to n = 150) and introduced class imbalance (ratio of healthy to diseased was 9:1)”).
Chang does not distinctly disclose
wherein the measure of variation of the features of the matching data is determined for matching data at each respective clinical site
However, Srinivasa teaches
wherein the measure of variation of the features of the matching data is determined for matching data at each respective clinical site (Srinivasa [Section 3.1, par. 1]: “Assume a list of M potential trial sites or investigators {d1,··· ,dM}. Each site/ investigator di is characterized by their medical expertise and specialty areas. Further, we assume that each site provides access to a set of patients, each of whom belongs to one of L groups. For example, these could be ethnicity-based sub-populations. Hence, each site di is associated with a distribution of patients over the L groups. We represent these distributions for all the sites together using matrix P of size M ×L. Here, P[i,l] represents the percentage of the patients at site i that belong to group l.”)
Chang and Srinivasa are both analogous to the claimed invention because they are both in the same field of measuring variation within clinical data for multiple sites. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have used the measure of variation taught by Chang et al. and determined that variation across each clinical site as taught by Srinivasa et al. Doing so would enable a better demographic composition of clinical data (Srinivasa et al. Introduction)
Regarding Claim 7:
Chang does not distinctly disclose
The method of claim 1 wherein the measure of variation comprises a measure of heterogeneity.
However, Srinivasa teaches
The method of claim 1 wherein the measure of variation comprises a measure of heterogeneity (Srinivasa [Section 1, par. 4]: “To formalize the problem, we pose the task of trial site matching as a fair ranking problem, where we rank the list of potential trial sites in order to maximize expected patient enrollment and patient diversity. In the subsequent text, we use the terms ‘investigator’ and ‘site’ interchangeably for convenience. Towards this end, we propose a machine learning algorithm that learns to generate a Top-K, trial-specific ranking of a given list of investigators. Given a new clinical trial and a list of M investigators (M > K), the algorithm selects the top K investigators from this list. The set of Top-K investigators are chosen by simultaneously optimizing for patient enrollment and patient diversity…For the scope of this paper, we define diversity based on the race and ethnicity of the patients. In particular, we assume the following groups: White, Hispanic, African-American, Asian, Mixed, and Others.”; (EN): Section 3.2 clarifies the “patient enrollment and patient diversity” is maximized by stating “We use a policy learning based framework to learn the Top-K selection from a given list. In particular, our model learns to assign a score to each site in the input list and chooses the Top-K sites based on the scores).
Chang and Srinivasa are both analogous to the claimed invention because they are both in the same field of measuring variation within clinical data for multiple sites. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced class imbalance and image resolution as the measure of variation as taught by Chang et al. with a measure of heterogeneity as taught by Srinivasa et al. Hence, this would be a simple substitution of one known element (class imbalance and resolution) with another (heterogeneity) to obtain predictable results (measure of variation) (MPEP 2141(III)(B) Simple substitution of one known element for another to obtain predictable results).
Regarding Claim 8:
Chang does not distinctly disclose
The method of claim 7 wherein the measure of heterogeneity is determined using a second machine learning model that takes the features in the metadata as input and outputs the measure of heterogeneity.
However, Srinivasa teaches
The method of claim 7 wherein the measure of heterogeneity is determined using a second machine learning model that takes the features in the metadata as input and outputs the measure of heterogeneity (Srinivasa [Section 1, par. 4]: “To formalize the problem, we pose the task of trial site matching as a fair ranking problem, where we rank the list of potential trial sites in order to maximize expected patient enrollment and patient diversity. In the subsequent text, we use the terms ‘investigator’ and ‘site’ interchangeably for convenience. Towards this end, we propose a machine learning algorithm that learns to generate a Top-K, trial-specific ranking of a given list of investigators. Given a new clinical trial and a list of M investigators (M > K), the algorithm selects the top K investigators from this list. The set of Top-K investigators are chosen by simultaneously optimizing for patient enrollment and patient diversity; [Section 4.1, par. 1]: “For a given trial t, our model first predicts a score for each of the M investigators. The scores are generated using a deep neural network, denoted as fθ, where θ is the set of parameters of the deep neural network. The input to the neural network is the trial features (of dimension p), along with the set of investigator features. We denote the set of M outputs of the deep neural network containing the scores as h(t)∈RM.”)
Chang and Srinivasa are both analogous to the claimed invention because they are both in the same field of measuring variation within clinical data for multiple sites. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced class imbalance and image resolution as the measure of variation as taught by Chang et al. with a measure of heterogeneity determined using a machine learning model as taught by Srinivasa et al. Hence, this would be a simple substitution of one known element (class imbalance and resolution) with another (heterogeneity) to obtain predictable results (measure of variation) (MPEP 2141(III)(B) Simple substitution of one known element for another to obtain predictable results).
Regarding Claim 9:
Chang teaches
The method of claim 8 wherein the second machine learning model outputs a list comprising a subset of the matching data that optimizes or maximizes the heterogeneity compared to other possible subsets of the matching training data (Chang [Page 946, Preprocessing, par. 1]: “We obtained 35 126 color digital retinal fundus (interior surface of the eye) photos from the Kaggle Diabetic Retinopathy competition.13 Each image was rated for disease severity by a licensed clinician on a scale of 0–4 (absent, mild, moderate, severe, and proliferative retinopathy, respectively). The images came from 17,563 patients of multiple primary care sites throughout California and elsewhere. The acquisition conditions were varied, with a range of camera models, levels of focus, and exposures. In addition, the resolutions ranged from 433 289 pixels to 5184 3456 pixels…To simplify training of the network, the labels were binarized to Healthy (scale 0) and Diseased (scale 2, 3, or 4)… It is also known that there is a correlation between the disease status of the left eye and the status of the right eye. To remove this as a confounding factor in our study, only images from left eye were utilized…The dataset was randomly sampled, with equal class distributions, into 4 “institutions,” each institution having n = 1500 patients. In addition, the dataset was sampled to create a single validation cohort (n = 3000 patients) and a single testing (n = 3000 patients) cohort, again with equal class probabilities (Figure 1B). Sampling was without replacement such that there are no overlapping patients in any of the cohorts. The image intensity was normalized within each channel across all patients within each cohort. Because model performance plateaus as the number of training patient samples increases, the number of patients per institution was limited to 1500 to prevent saturation of learning for models trained in single institutions. We tested several different training heuristics (Figure 2) and compared the results. The first heuristic is training a neural network for each institution individually, assuming there is no collaboration between the institutions”)
Srinivasa further teaches
The method of claim 8 wherein the second machine learning model outputs a list comprising a subset of the matching data that optimizes or maximizes the heterogeneity compared to other possible subsets of the matching training data (Srinivasa [Section 1, par. 4]: “To formalize the problem, we pose the task of trial site matching as a fair ranking problem, where we rank the list of potential trial sites in order to maximize expected patient enrollment and patient diversity. In the subsequent text, we use the terms ‘investigator’ and ‘site’ interchangeably for convenience. Towards this end, we propose a machine learning algorithm that learns to generate a Top-K, trial-specific ranking of a given list of investigators. Given a new clinical trial and a list of M investigators (M > K), the algorithm selects the top K investigators from this list. The set of Top-K investigators are chosen by simultaneously optimizing for patient enrollment and patient diversity…For the scope of this paper, we define diversity based on the race and ethnicity of the patients. In particular, we assume the following groups: White, Hispanic, African-American, Asian, Mixed, and Others.”)
Chang and Srinivasa are both analogous to the claimed invention because they are both in the same field of measuring variation within clinical data for multiple sites. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced class imbalance and image resolution as the measure of variation as taught by Chang et al. with a measure of heterogeneity determined using a machine learning model as taught by Srinivasa et al. Hence, this would be a simple substitution of one known element (class imbalance and resolution) with another (heterogeneity) to obtain predictable results (measure of variation) (MPEP 2141(III)(B) Simple substitution of one known element for another to obtain predictable results).
Regarding Claim 10:
Chang does not distinctly disclose
The method of claim 8 wherein the second machine learning model takes as input feature values in the metadata, xn, and determines a combination of xn that optimally flattens a linear function f(x)=x'3+b, wherein p comprises the gradient of the function and b comprises an offset.
However, Srinivasa teaches
The method of claim 8 wherein the second machine learning model takes as input feature values in the metadata, xn, and determines a combination of xn that optimally flattens a linear function f(x)=x'3+b, wherein p comprises the gradient of the function and b comprises an offset (Srinivasa [Section 4.3]: Section 4.3 showcases policy optimization via Monte Carlo sampling; in Equation 18, “P” represents the trial features, thus corresponding to “input feature values in the metadata xn” and “x’” in f(x)=x′β+b; Equation 25 shows that the gradient [symbolized by an inverse triangle] of F(r|t) is taken which includes lf(r, P) in Equation 18 and taking the gradient means taking the derivative, gradient of F(r|t) corresponds to “ x’ ”; “λ” in Equations 18-20 and Equation 25 correspond to “the gradient of the function” and “β” f(x)=x′β+b ; “U(r|t)” in Equations 18-20 and Equation 25 corresponds to “an offset” and “b” in f(x)=x′β+b).
Chang and Srinivasa are both analogous to the claimed invention because they are both in the same field of measuring variation within clinical data for multiple sites. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced class imbalance and image resolution as the measure of variation as taught by Chang et al. with a measure of heterogeneity determined using a machine learning model that flattens a linear function as taught by Srinivasa et al. Hence, this would be a simple substitution of one known element (class imbalance and resolution) with another (heterogeneity) to obtain predictable results (measure of variation) (MPEP 2141(III)(B) Simple substitution of one known element for another to obtain predictable results).
Claim Rejections - 35 USC § 103
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chang as applied to claims 1, 14, and 15 above, and further in view of de Luca et al (de Luca, A. B., Zhang, G., Chen, X., & Yu, Y. (2022). Mitigating Data Heterogeneity in Federated Learning with Data Augmentation. arXiv [Cs.LG]. Retrieved from http://arxiv.org/abs/2206.09979, hereinafter de Luca).
Regarding Claim 6:
Chang does not distinctly disclose
The method of claim 1 further comprising: supplementing the training dataset with augmented training data so as to increase the variation of the selected training dataset.
However, de Luca teaches
The method of claim 1 further comprising: supplementing the training dataset with augmented training data so as to increase the variation of the selected training dataset (de Luca [Section 4, par. 3]: “Data augmentation is a widely used strategy in machine learning for increasing the diversity of samples. It often allows us to obtain more robust models with better performance effectively”; [Section 5.4, par. 1]: “By augmenting the clients’ data…local models can be trained in isolation for longer without suffering from weight divergence, as reported in”).
Chang and Luca are both analogous to the claimed invention because they are both in the same field of increasing variation in training data. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the left eye image training data disclosed by Chang et al. by augmenting that training data as taught by Luca et al. Doing so would prevent weight divergence in machine learning models (Luca et al. Section 5.4).
Claim Rejections - 35 USC § 103
Claim(s) 11 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chang and Srinivasa as applied to claims 1, 14, and 15 above, and further in view of MathWorks (MathWorks, "Understanding Support Vector Machine Regression", 2016 [WayBack Machine] (Year: 2016), hereinafter MathWorks).
Regarding Claim 11:
Chang + Srinivasa does not distinctly disclose
The method of claim 10 wherein the second machine learning model is configured to determine f(x) with the minimal norm value (β′β) according to a convex optimization problem wherein the function: J(β)=1/2(β′β) is to be minimized
However, MathWorks teaches
The method of claim 10 wherein the second machine learning model is configured to determine f(x) with the minimal norm value (β′β) according to a convex optimization problem wherein the function: J(β)=1/2(β′β) is to be minimized (MathWorks [Primal Formula]: “To find the linear function f(x)=x′β+b, and ensure that it is as flat as possible, find f(x) with the minimal norm value (β′β). This is formulated as a convex optimization problem to minimize J(β)=1/2(β′β)”)
Chang et al., Srinivasa et al. and MathWorks are all analogous to the claimed invention because they are all in the same field of using machine learning models to minimize a loss function. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the neural network disclosed by Srinivasa et al. with a support vector as taught by MathWorks which uses the minimal norm value to minimize a loss function. Hence, this would be a simple substitution of one known element (neural network) with another (support vector machine using the minimal norm value) to obtain predictable results (minimizing a loss function) (MPEP 2141(III)(B) Simple substitution of one known element for another to obtain predictable results).
Regarding Claim 12:
Chang + Srinivasa does not distinctly disclose
The method of claim 8 wherein the second machine learning model comprises a support vector regression model.
However, MathWorks teaches
The method of claim 8 wherein the second machine learning model comprises a support vector regression model (MathWorks [Overview]: “Support vector machine (SVM) analysis is a popular machine learning tool for classification and regression, first identified by Vladimir Vapnik and his colleagues in 1992[5]. SVM regression is considered a nonparametric technique because it relies on kernel functions. Statistics and Machine Learning Toolbox™ implements linear epsilon-insensitive SVM (ε-SVM) regression, which is also known as L1 loss. In ε-SVM regression, the set of training data includes predictor variables and observed response values. The goal is to find a function f(x) that deviates from yn by a value no greater than ε for each training point x, and at the same time is as flat as possible”).
Chang et al., Srinivasa et al. and MathWorks are all analogous to the claimed invention because they are all in the same field of using machine learning models to minimize a loss function. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the neural network disclosed by Srinivasa et al. with a support vector as taught by MathWorks which uses the minimal norm value to minimize a loss function. Hence, this would be a simple substitution of one known element (neural network) with another (support vector machine using the minimal norm value) to obtain predictable results (minimizing a loss function) (MPEP 2141(III)(B) Simple substitution of one known element for another to obtain predictable results).
Claim Rejections - 35 USC § 103
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chang as applied to claims 1, 14, and 15 above, and further in view of Tong et al (Tong, J., Luo, C., Islam, M.N. et al. Distributed learning for heterogeneous clinical data with application to integrating COVID-19 data across 230 sites. npj Digit. Med. 5, 76 (2022). https://doi.org/10.1038/s41746-022-00615-8, hereinafter Tong).
Regarding Claim 13:
Chang does not distinctly disclose
The method of claim 1 further comprising instructing each clinical site in the plurality of clinical sites to create a local copy of the model and train the local copy of the model using the training data in the training dataset selected from the respective clinical site;
and combining the results of the training according to the distributed learning process.
However, Tong teaches
The method of claim 1 further comprising instructing each clinical site in the plurality of clinical sites to create a local copy of the model and train the local copy of the model using the training data in the training dataset selected from the respective clinical site (Tong [Page 6, Methods, par. 1]: “To handle the site-specific effect, we develop a privacy-preserving distributed pairwise conditional logistic regression (dCLR) algorithm. As shown in Fig. 5, there are two steps required to implement the proposed algorithm: initialization and surrogate estimator estimation. In the first step (i.e., initialization), each site fits a conditional logistic regression model with its own local patient-level data. Then, the sites shared the initial estimates of the parameters of interest (i.e., regression coefficients) across the collaborative sites within the network. With all the initial values, an overall initial estimate, β(i.e., average or weighted average of the initial values) is calculated. In Step 2 (i.e., surrogate estimator estimation), each site first shares the intermediate results (i.e., first and second gradients of the local log likelihood function, which are calculated using the overall initial estimate β). Then, the intermediate results are assembled to construct the surrogate pairwise likelihood, which is maximized by the surrogate estimator”);
and combining the results of the training according to the distributed learning process (Tong [Page 6, Methods, par. 1]: “To handle the site-specific effect, we develop a privacy-preserving distributed pairwise conditional logistic regression (dCLR) algorithm. As shown in Fig. 5, there are two steps required to implement the proposed algorithm: initialization and surrogate estimator estimation. In the first step (i.e., initialization), each site fits a conditional logistic regression model with its own local patient-level data. Then, the sites shared the initial estimates of the parameters of interest (i.e., regression coefficients) across the collaborative sites within the network. With all the initial values, an overall initial estimate, β(i.e., average or weighted average of the initial values) is calculated. In Step 2 (i.e., surrogate estimator estimation), each site first shares the intermediate results (i.e., first and second gradients of the local log likelihood function, which are calculated using the overall initial estimate β). Then, the intermediate results are assembled to construct the surrogate pairwise likelihood, which is maximized by the surrogate estimator”).
Chang and Tong are both analogous to the claimed invention because they are both in the same field of utilizing distributed learning for multiple clinical sites. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the distributed learning architecture which uses cyclical weight transfer as disclosed by Chang et al. with a distributed learning architecture that uses local models for each site as disclosed by Tong et al. Thus, this would be a simple substitution of one known element (cyclical weight transfer) with another (distributed learning using local models) to obtain predictable results (aggregating results from multiple clinical sites using distributed learning) (MPEP 2141(III)(B) Simple substitution of one known element for another to obtain predictable results).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Fadila Zerka et al. Systematic Review of Privacy-Preserving Distributed Machine Learning From Federated Databases in Health Care. JCO Clin Cancer Inform 4, 184-200(2020). DOI:10.1200/CCI.19.00047 – utilizing clinical data from multiple sites in distributed learning
US 20200134446 A1 – training local models then aggregating results to a central model
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to COREY M SACKALOSKY whose telephone number is (703)756-1590. The examiner can normally be reached M-F 7:30am-3:30pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at (571) 272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/COREY M SACKALOSKY/Examiner, Art Unit 2128
/BRIAN M SMITH/Primary Examiner, Art Unit 2122