Prosecution Insights
Last updated: August 16, 2026
Application No. 18/579,328

METHOD AND APPARATUS FOR TRAINING MACHINE LEARNING MODELS, COMPUTER DEVICE, AND STORAGE MEDIUM

Non-Final OA §102§103§112
Filed
Jan 12, 2024
Priority
Jul 12, 2021 — nonprovisional of PCTCN2021105777
Examiner
CHEN, KUANG FU
Art Unit
Tech Center
Assignee
Shanghai United Imaging Healthcare Co., Ltd.
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
216 granted / 270 resolved
+20.0% vs TC avg
Strong +68% interview lift
Without
With
+68.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
26 currently pending
Career history
295
Total Applications
across all art units

Statute-Specific Performance

§101
16.9%
-23.1% vs TC avg
§103
50.3%
+10.3% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
15.2%
-24.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 270 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to the claims filed 1/12/2024. Claims 1-13, 15-17, 19-21, and 26 are presented for examination. Priority Applicant’s claimed benefit of PCT/CN2021/105777 filed 7/12/2021 is acknowledged. Information Disclosure Statement The information disclosure statement (IDS) submitted on 1/12/2024, 10/6/2024, and 6/25/2026 were considered by the examiner. Specification The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: Method and Apparatus for Training Machine Learning Models Using Shared Parameters. Claim Interpretation - 35 U.S.C. 112(f) The claims in this application are given their broadest reasonable interpretation in light of the specification as it would be interpreted by one of ordinary skill in the art. Because this application is a national stage application, the claims are examined in accordance with 35 U.S.C. 112(f). This application includes one or more claim limitations that do not use the word "means," but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because each such claim limitation uses a generic placeholder (a nonce term) that is coupled with functional language without reciting sufficient structure to perform the recited function, and the generic placeholder is not preceded by a structural modifier. See Williamson v. Citrix Online, LLC, 792 F.3d 1339, 1349-51 (Fed. Cir. 2015) (en banc). Such claim limitations are: a sample set acquisition module configured to acquire a first training sample set and a second training sample set, recited in claim 26; and one or more training modules configured to: perform, based on the first training sample set, multiple rounds of model training, to obtain a first machine learning model; and perform, based on the second training sample set, multiple rounds of model training, to obtain a second machine learning model, recited in claim 26. For each of the above limitations, the term "module" is a generic placeholder (a nonce term) that operates as a substitute for the word "means." The word "module" is not preceded by a structural modifier; the qualifiers "sample set acquisition" and "training" describe what the module does rather than denoting any definite structure. Each "module" is coupled with functional language introduced by "configured to" and does not itself recite sufficient structure to perform the entirety of the recited function. Accordingly, all three prongs of the analysis set forth in MPEP 2181, subsection I, are met, and each limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f), applicant may: (1) amend the claim limitations so that they will no longer be interpreted under 35 U.S.C. 112(f) (for example, by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f). Sample set acquisition module (claim 26). The claimed function is acquiring a first training sample set and a second training sample set, the training samples comprising medical images obtained by scanning a scan subject with a medical scanning device. The corresponding structure disclosed in the specification is the processor of the computer device of FIG. 12, executing a computer program stored on a memory (paragraphs [0215] and [0216]), programmed to perform the acquisition algorithm described in the specification, namely acquiring medical images from one or more sources (for example, dividing a plurality of medical images obtained from the same medical imaging equipment into two image sets, or acquiring medical images from a first hospital and from a second, different hospital) and generating the first training sample set and the second training sample set from the acquired medical images (paragraphs [0086] through [0091] and [0198]). This disclosure clearly links a structure, a programmed processor performing the described acquisition algorithm, to the entirety of the claimed function. The broadest reasonable interpretation of this limitation is the programmed processor of the computer device performing the acquisition algorithm described at paragraphs [0086] through [0091] and [0198], and equivalents thereof. One or more training modules (claim 26). The claimed function is (1) performing, based on the first training sample set, multiple rounds of model training to obtain a first machine learning model, and (2) performing, based on the second training sample set, multiple rounds of model training to obtain a second machine learning model, at least a part of the model parameters of each machine learning model being used when the other machine learning model is trained. The corresponding structure disclosed in the specification is the processor of the computer device of FIG. 12, executing a computer program stored on a memory (paragraphs [0215] and [0216]), programmed to perform the training algorithm shown in the flow diagrams of FIG. 2 and FIGS. 4 through 9 and described in prose at paragraphs [0084] through [0195]. That algorithm includes performing, based on first initial model parameters and the first training sample set, a first round of model training to obtain a first initial model; performing, based on the first training sample set and at least a part of the model parameters of the present second machine learning model, an Nth round of model training to obtain the first machine learning model; terminating the training of the first machine learning model when a model index of the first machine learning model meets a first preset index, and otherwise performing a further round of training; performing corresponding rounds of model training of the second machine learning model using at least a part of the model parameters of the present first machine learning model and terminating when a model index of the second machine learning model meets a second preset index; and, during each round, determining a descent gradient using a batch gradient algorithm when an output result does not meet a preset convergence condition according to a preset loss function, and continuing until the preset convergence condition is met (paragraphs [0188] through [0195]). This disclosure clearly links a structure, a programmed processor performing the described training algorithm, to the entirety of the claimed function. The broadest reasonable interpretation of this limitation is the programmed processor of the computer device performing the training algorithm described at paragraphs [0084] through [0195] and depicted in FIG. 2 and FIGS. 4 through 9, and equivalents thereof. Because the specification discloses corresponding structure, a programmed processor executing the disclosed acquisition and training algorithms, that is clearly linked to and performs the entirety of each claimed function, the limitations interpreted under 35 U.S.C. 112(f) do not render claim 26 indefinite. No rejection under 35 U.S.C. 112(b) is made on the basis of these limitations, and no rejection under 35 U.S.C. 112(a) arising from these limitations is warranted. This interpretation affects claim construction only. Claim Rejections - 35 U.S.C. 112(a) The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claim 20 is rejected under 35 U.S.C. 112(a) as failing to comply with the written description requirement. The claim contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, at the time the application was filed, had possession of the claimed invention. Claim 20 depends from claim 1 and further recites combining the first machine learning model and the second machine learning model, to obtain a target machine learning model. The written description does not reasonably convey possession of this limitation commensurate with its full scope. The only disclosure of the combining feature appears at paragraphs [0028], [0056], and [0166]-[0167]. Paragraph [0167] states that the model training terminal combines the first machine learning model and the second machine learning model to obtain the combined target machine learning model, and gives a single illustration in which, by combining a dose prediction model with an automatic sketching model, a target machine learning model that performs dose prediction first followed by automatic sketching may be obtained, making the function of the model more powerful. Paragraphs [0028] and [0056] restate the same result in the summary and apparatus contexts without adding detail. This disclosure recites only the result to be achieved, namely a combined, more powerful target machine learning model, together with one application-specific instance, namely a sequential pipeline of a dose prediction model followed by an automatic sketching model. It does not describe any structure, architecture, or algorithm by which two separately trained machine learning models are combined into a single target machine learning model, such as ensembling, weight merging, sequential composition, or knowledge distillation. Because independent claim 1, from which claim 20 depends, permits the first and second machine learning models to be identical in structure and application (see claim 17 and paragraph [0163], describing two dose prediction models) as well as different in application (see claim 17 and paragraph [0165]), the claimed scope of combining the two models to obtain a target machine learning model reaches combinations for which the specification provides no descriptive support. For example, combining two dose prediction models trained on medical images from different hospitals into a single target machine learning model is not addressed anywhere in the disclosure, and the sequential dose-prediction-then-sketching example does not reach that scenario. A claim that is defined by the function or result to be achieved, without a description of the structure or steps that accomplish that function sufficient to show that the inventor possessed the full claimed scope, does not satisfy the written description requirement (MPEP 2163; Ariad Pharmaceuticals, Inc. v. Eli Lilly and Co., 598 F.3d 1336, 1349-51 (Fed. Cir. 2010) (en banc); Abbvie Deutschland GmbH and Co. v. Janssen Biotech, Inc., 759 F.3d 1285, 1300-01 (Fed. Cir. 2014)). Although claim 20 was present in the application as originally filed, the presumption that an original claim serves as its own written description is rebutted where, as here, the broadest reasonable scope of the claim exceeds what the specification as a whole conveys the inventor possessed (In re Koller, 613 F.2d 819, 823 (CCPA 1980)). Applicant is required to point to a disclosure that reasonably conveys possession of the full scope of the combining limitation or to amend the claim to recite only the disclosed manner of combination. Claim 20 is rejected under 35 U.S.C. 112(a) as failing to comply with the enablement requirement. The claim contains subject matter which was not described in the specification in such a way as to enable one skilled in the art to which it pertains, or with which it is most nearly connected, to make and/or use the invention. The limitation combining the first machine learning model and the second machine learning model, to obtain a target machine learning model is not enabled across its full scope. Enablement is assessed as of the effective filing date and must be commensurate with the full scope of the claim. Whether a person of ordinary skill in the art would be required to engage in undue experimentation to make and use the full scope of this limitation is evaluated under the factors set forth in In re Wands, 858 F.2d 731, 737 (Fed. Cir. 1988): (a) The breadth of the claims. The limitation is broad and functional. It reads on any manner of combining any two machine learning models of any architecture, and on models that are identical in structure and application or different in application, to obtain any target machine learning model. Broad functional scope demands correspondingly greater enabling disclosure. (b) The nature of the invention. The invention is a computer-implemented method for training machine learning models on medical images; the combining feature of claim 20 is an added step that produces a single downstream target model from two separately trained models. (c) The state of the prior art. Techniques for combining models, such as ensembling and model composition, existed in the art, but the specification does not invoke or adapt any such technique. It ties the combining feature only to a bespoke sequential dose-prediction-then-sketching pipeline. (d) The level of one of ordinary skill. The level of ordinary skill is high, typically an artisan holding an advanced degree or equivalent experience in machine learning and medical image processing. Such an artisan could implement known combination techniques, but could not, from this disclosure, determine how to combine two arbitrary trained models, in particular two models of the same application or of differing structures, into a single more powerful target model. (e) The level of predictability in the art. Software and machine learning are generally predictable arts. However, whether an ad hoc combination of two independently trained networks yields a functional and improved target model is not predictable from the single application-specific example provided; the effect of combining models trained on different data or for different tasks is not foreseeable from the disclosure. (f) The amount of direction or guidance provided. The direction is minimal. A single sentence in paragraph [0166] states the result and gives one application-specific illustration. No parameters, architecture, interface, ordering rule, or method for performing the combination is provided. (g) The presence or absence of working examples. The specification provides one prophetic illustration, framed in terms of a result that may be obtained, rather than an actual worked example. No example is given for combining two models of the same application, and no example reduces the combination to a described sequence of operations. (h) The quantity of experimentation needed. For the full claimed scope, a person of ordinary skill would have to independently devise a combination scheme for each pairing of models, including pairings of differing structures and pairings of the same application, with no disclosed starting algorithm. This amounts to more than routine experimentation. Weighing the Wands factors as a whole, and giving due weight to the generally predictable nature of the software art, the specification nonetheless does not enable a person of ordinary skill to make and use the full scope of the combining limitation without undue experimentation. Merely stating that a computer combines two models to obtain a target model, without disclosing how the combination is performed, fails to enable the full scope of the claimed function (MPEP 2164.01(a); Williamson v. Citrix Online, LLC, 792 F.3d 1339, 1351-52 (Fed. Cir. 2015); Vasudevan Software, Inc. v. MicroStrategy, Inc., 782 F.3d 671, 681-83 (Fed. Cir. 2015)). This same absence of a disclosed algorithm supports the written description rejection above, as the two requirements are separate and independent. Claim Rejections - 35 USC 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 12-13, 17, 19-21, and 26 are rejected under 35 U.S.C. 102(a)(1) as anticipated by Daykin et al. (hereinafter Daykin), US 2021/0097381 A1. Regarding independent claim 1, Daykin discloses a method for training machine learning models (Daykin: [0048], "Certain embodiments provide a method for training a model comprising : training a model by a respective apparatus at each of a plurality of entities"; Daykin discloses a multi-entity method for training machine-learning models), the method comprising: acquiring a first training sample set and a second training sample set, training samples in the first training sample set and the second training sample set comprising medical images obtained by scanning a scan subject with a medical scanning device (Daykin: [0064-0065] FIGs.3 and 4, "FIG. 4, the data store 60 stores a data cohort comprising medical imaging data…FIG. 3, each institution 40A, 40B, 40C, 40D has a respective one or more apparatuses 50…Each of the apparatuses has access to a respective data cohort at its respective institution . The data cohort for an institution may comprise data acquired by scanners located at the institution. The data cohort for an institution may comprise data for patients treated by the institution", [0070] “FIG. 5 is a flow chart illustrating…a training method…performed by apparatuses 50 at a plurality of institutions 40A, 40B, 40C, 40D”; Daykin discloses at least two respective institutional cohorts containing patient data acquired by institutional scanners comprising medical images acquired by scanners located at the institutions for patients treated by the institution); performing, based on the first training sample set, multiple rounds of model training, to obtain a first machine learning model (Daykin: [0126-0127], "An output of FIG. 5 is a respective updated model for each institution…cycle may be repeated multiple times…until the models converge"; Daykin discloses that each institution obtains its own updated model and repeats the training cycle multiple times on the institutional cohorts); and performing, based on the second training sample set, multiple rounds of model training, to obtain a second machine learning model (Daykin: [0127], "FIG . 5 shows only one training cycle. In practice, multiple training cycles may be performed…Each of the institutions…obtained by training the updated model on data held at that institution"; Daykin applies the repeated-cycle procedure at every institution, including a second institution with its respective cohort and updated model), wherein at least a part of the first machine learning model has a same structure as at least a part of the second machine learning model (Daykin: [0073], "The same model is provided to each institution 40A , 40B , 40C , 40D"; the copies of the same model necessarily have the same organization in at least corresponding parts), at least a part of model parameters of the second machine learning model is used when the first machine learning model is trained (Daykin: [0126], "At each institution , the updated model is formed by aggregating model parameters from all of the institutions in accordance with the influence values"; the first institution's updated-model training expressly uses model parameters from the second institution), and at least a part of model parameters of the first machine learning model is used when the second machine learning model is trained (Daykin: [0127], "each institution sends its updated model to all of the institutions for training. Each of the institutions returns further model parameters that were obtained by training the updated model on data held at that institution"; the all to all operation is reciprocal, so the second institution trains with the first institution's updated parameter state). Regarding dependent claim 12, Daykin discloses the method of claim 1, wherein the acquiring medical images from a first hospital, and generating, based on the medical images from the first hospital, the first training sample set; and acquiring medical images from a second hospital, and generating, based on the medical images from the second hospital, the second training sample set, wherein the first hospital is different from the second hospital (Daykin: [0048] “wherein the training at each of the plurality of entities is performed is performed on a respective data cohort held at said entity”, [0049-0050], "In the present embodiment, the institutions are hospitals", [0064-0065] FIGs.3 and 4, "FIG. 4, the data store 60 stores a data cohort comprising medical imaging data…FIG. 3, each institution 40A, 40B, 40C, 40D has a respective one or more apparatuses 50…Each of the apparatuses has access to a respective data cohort at its respective institution . The data cohort for an institution may comprise data acquired by scanners located at the institution. The data cohort for an institution may comprise data for patients treated by the institution", [0070] “FIG. 5 is a flow chart illustrating…a training method…performed by apparatuses 50 at a plurality of institutions 40A, 40B, 40C, 40D”; Daykin's plurality of institutions have respective scanner-derived cohorts used for training, wherein the data cohort comprising medical imaging data, and expressly identifying those institutions as hospitals provides different first and second hospitals). Regarding dependent claim 13, Daykin discloses the method of claim 1, wherein the acquiring the first machine learning model and the second machine learning model ([0064-0065] FIGs.3 and 4, "FIG. 4, the data store 60 stores a data cohort comprising medical imaging data…FIG. 3, each institution 40A, 40B, 40C, 40D has a respective one or more apparatuses 50…Each of the apparatuses has access to a respective data cohort at its respective institution . The data cohort for an institution may comprise data acquired by scanners located at the institution. The data cohort for an institution may comprise data for patients treated by the institution", [0070] “FIG. 5 is a flow chart illustrating…a training method…performed by apparatuses 50 at a plurality of institutions 40A, 40B, 40C, 40D”) comprise at least one of a dose prediction model, an automatic sketching model, an efficacy evaluation model, a survival index evaluation model, a cancer screening model or a deformation registration model (Daykin: [0076], [0132], "the trained model performs the task for which it is trained , for example classification and / or segmentation and / or location"; Daykin's automatically trained medical-image segmentation/location model produces the region delineation recited by the automatic-sketching alternative). Regarding dependent claim 17, Daykin discloses the method of claim 1, wherein the first machine learning model and the second machine learning model are identical in structure and application, or the first machine learning model and the second machine learning model are different in application (Daykin: [0073], "The same model is provided to each institution 40A , 40B , 40C , 40D"; copies of the same task model necessarily have identical structure and application). Regarding dependent claim 19, Daykin discloses the method of claim 1, wherein only model parameters are transferred during the training of the first machine learning model and the second machine learning model (Daykin: [0066] “Typically, the apparatus at an institution only has access to the data cohort for that institution”, [0075] "a reference to a model being transferred between institutions may be taken to refer to a set of model parameters being transferred between institutions", [0081] “At stage 71, the training apparatus…also sends a copy of the trained parameters… that it has obtained by training the model on its own institution’s data cohort to apparatuses at each of the other institutions”, [0141] “The soft federated learning method of FIGS. 5 and 7”; Daykin prevents transfer of the local data cohorts and defines the models sent during training as parameter sets). Regarding dependent claim 20, Daykin discloses the method of claim 1, further comprising: combining the first machine learning model and the second machine learning model, to obtain a target machine learning model (Daykin: [0052], "The apparatus 50 is configured to train a model (for example , a neural network) and to combine trained model parameters obtained from training on the apparatus 50 itself and from training on other apparatuses 50 at other institutions"; combining the two trained parameter sets produces the local updated target model). Regarding independent claim 21, Daykin discloses a method for training machine learning models (Daykin: paragraph 0048, "Certain embodiments provide a method for training a model comprising: training a model by a respective apparatus at each of a plurality of entities"; Daykin discloses a multi-entity machine-model training method), the method comprising: acquiring at least two training sample sets, training samples in the training sample set comprising medical images obtained by scanning a scan subject with a medical scanning device (Daykin: [0055-0057] “at least one scanner 54 may comprise any scanner that is configured to perform medical scanning…Image data sets obtained by the at least one scanner 54 are stored in the data store 60 and subsequently provided to computing apparatus 52…data store 60 stores a cohort of training data comprising a plurality of training image data sets”, [0065] "Image data sets obtained by the at least one scanner 54 are stored in the data store 60 and subsequently provided to computing apparatus 52"; Daykin performs the scanner-to-computer acquisition for each of multiple institutions having respective patient cohorts); and performing, based on each of the at least two training sample sets, multiple rounds of model training, to obtain a machine learning model corresponding to a respective training sample set (Daykin: [0126-0127], "An output of FIG . 5 is a respective updated model for each institution…FIG. 5 shows only one training cycle. In practice, multiple training cycles may be performed…Each of the institutions…obtained by training the updated model on data held at that institution"; each institutional set has its own updated model and Daykin repeats the training cycle multiple times); wherein at least a part of each of at least two machine learning models has a same structure as at least a part of each of other machine learning models (Daykin: [0073], "The same model is provided to each institution 40A, 40B, 40C, 40D"; the common architecture necessarily provides same-structured corresponding parts); when one of the at least two machine learning models is trained, at least a part of model parameters of the part of another of the at least two machine learning models having the same structure is used (Daykin: [0126-0127], "each institution sends its updated model to all of the institutions for training"; each receiver trains with another site's compatible updated model and locally aggregates the institutional parameter vectors). Regarding independent claim 26 Daykin discloses an apparatus for training machine learning models (Daykin: [0052], [0054], "apparatus 50 is configured to train a model…The apparatus 50 comprises a computing apparatus 52 , in this case a personal computer (PC) or workstation"; Daykin discloses a programmed computing apparatus for machine-model training), the apparatus comprising: a sample set acquisition module (the sample set acquisition module construed under 35 U.S.C. 112(f) as disclosed programmed processor structures or equivalents that execute the claimed algorithms) configured to acquire a first training sample set and a second training sample set, training samples in the first training sample set and the second training sample set comprising medical images obtained by scanning a scan subject with a medical scanning device (Daykin: [0054-0057] “at least one scanner 54 may comprise any scanner that is configured to perform medical scanning…Image data sets obtained by the at least one scanner 54 are stored in the data store 60 and subsequently provided to computing apparatus 52…data store 60 stores a cohort of training data comprising a plurality of training image data sets”; the scanner, data store, and programmed computing apparatus execute the same acquisition algorithm for each respective institutional image cohort and are equivalent structure); one or more training modules (training modules construed under 35 U.S.C. 112(f) as disclosed programmed processor structures or equivalents that execute the claimed algorithms) configured to: perform, based on the first training sample set, multiple rounds of model training, to obtain a first machine learning model; and perform, based on the second training sample set, multiple rounds of model training, to obtain a second machine learning model (Daykin: [0061-0062] "the circuitries 64, 66, 68, 69 are each implemented in the CPU and / or GPU by means of a computer program having computer-readable instructions that are executable to perform the method of the embodiment", [0126-0130] "An output of FIG . 5 is a respective updated model for each institution…FIG. 5 shows only one training cycle. In practice, multiple training cycles may be performed…Each of the institutions…obtained by training the updated model on data held at that institution"; the programmed processor executes the repeated local training, testing, aggregation, and convergence algorithm for the respective cohorts and models and is equivalent training-module structure), wherein at least a part of the first machine learning model has a same structure as at least a part of the second machine learning model, at least a part of model parameters of the second machine learning model is used when the first machine learning model is trained, and at least a part of model parameters of the first machine learning model is used when the second machine learning model is trained (Daykin: [0073], "The same model is provided to each institution 40A, 40B, 40C, 40D", [0126-0127], "At each institution, the updated model is formed by aggregating model parameters from all of the institutions in accordance with the influence values"; the programmed common-architecture modules reciprocally include each peer's compatible parameters in the respective updated models over repeated cycles). Claim Rejections - 35 U.S.C. 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2-9, 11, 15, and 16 are rejected under 35 U.S.C. 103 as unpatentable over Daykin, as applied in the rejection of claim 1 above, in view of Balachandar et al. (hereinafter Balachandar), US 2021/0049473 A1. Regarding dependent claim 2, Daykin teaches the method of claim 1, wherein the performing, based on the first training sample set, multiple rounds of model training, to obtain the first machine learning model (Daykin: [0126-0127], "An output of FIG. 5 is a respective updated model for each institution…cycle may be repeated multiple times…until the models converge"; Daykin discloses that each institution obtains its own updated model and repeats the training cycle multiple times on the institutional cohorts) comprises: performing, based on first initial model parameters and the first training sample set, a first round of model training, to obtain a first initial model, at least a part of model parameters of the first initial model being used to train the second machine learning model (Daykin: [0073], [0081-0082], "The training apparatus 50A also sends a copy of the trained parameters 20A that it has obtained by training the model on its own institution's data cohort to apparatuses at each of the other institutions 40B, 40C, 40D. Corresponding training processes for the model are performed by apparatuses 50B, 50C, 50D at each of the other institutions 40B, 40C, 40D"; Daykin teaches common initial model parameters, first-site local training, and delivery of the first trained parameter state to the other sites such as second site for corresponding training at the second institution); performing, based on the first training sample set and at least a part of the model parameters of the present second machine learning model, to obtain the first machine learning model (Daykin: [0127], "Each of the institutions returns further model parameters that were obtained by training the updated model on data held at that institution"; in a later cycle the first site trains (performing based on the first training sample set) a peer-derived updated model (and at least a part of the model parameters of the present second machine learning model) on its own cohort (to obtain the first machine learning model)); terminating, if it is determined that a model index of the first machine learning model meets a first preset index, the training of the first machine learning model (Daykin: [0130], "At stage 104, the training circuitry 64 determines whether training is complete. For example, the training circuitry 64 may determine a measure of convergence"; Daykin teaches a measured convergence index and a completion condition that ends training when satisfied). Daykin does not expressly label the two-site sequence as a first model followed by a present second model in cyclic serial order of an Nth round of model training, N being a positive integer greater than 1. However, Balachandar teaches performing, based on a first training sample set and at least a part of model parameters of a second machine learning model, an Nth round of model training, to obtain a first machine learning model, N being a positive integer greater than 1 (Balachandar: [0045] and Fig. 2, "CWT, in accordance with many embodiments, involves starting training at one institution for a certain number of iterations, transferring the updated model weights to a subsequent institution, training the model at the subsequent institution for a certain number of iterations, then transferring the updated weights to the next institution, and so on until model convergence"; Balachandar supplies the express serial first site to second site initialization and repeated cyclic rounds). Because Daykin and Balachandar are analogous art concerning privacy-preserving distributed neural-network training across medical institutions using local medical-image datasets and exchanged parameter values, addressing the same problem of improving models without pooling sensitive institutional image data, accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention, to apply Balachandar's serial CWT schedule to Daykin's compatible all to all parameter updates to teach performing, based on the first training sample set and at least a part of the model parameters of the present second machine learning model, an Nth round of model training, to obtain the first machine learning model, N being a positive integer greater than 1. This modification would have been motivated by the desire to provide specialized approaches for distributed deep learning for medical applications (Balachandar: [0004]). Regarding dependent claim 3, Daykin, in view of Balachandar, teach the method of claim 2, wherein a part of the first machine learning model has the same structure as a part of the second machine learning model (Daykin: [0073], "The same model is provided to each institution 40A , 40B , 40C , 40D"; the copies of the same model necessarily have the same organization in at least corresponding parts), and the performing, based on the first training sample set and at least a part of the model parameters of the present second machine learning model, the Nth round of model training, to obtain the first machine learning model (Balachandar: [0045] and Fig. 2, "CWT, in accordance with many embodiments, involves starting training at one institution for a certain number of iterations, transferring the updated model weights to a subsequent institution, training the model at the subsequent institution for a certain number of iterations, then transferring the updated weights to the next institution, and so on until model convergence"; Balachandar supplies the express serial first site to second site initialization and repeated cyclic rounds) comprises: performing, based on the first training sample set, model parameters of the part of the present second machine learning model having the same structure, and a part of the model parameters of the first machine learning model obtained from a previous round of model training, the Nth round of model training, to obtain the first machine learning model (Daykin: [0123], [0126], "At each institution, the updated model is formed by aggregating model parameters from all of the institutions in accordance with the influence values"; Daykin teaches an aggregate containing the site's own prior trained vector and compatible peer vectors, and selecting corresponding parts satisfies the at-least-part formulation; wherein Balachandar: [0045] and Fig. 2, supplies the express serial first site to second site initialization and repeated cyclic rounds). Regarding dependent claim 4, Daykin, in view of Balachandar, teach the method of claim 2, wherein the whole of the first machine learning model has the same structure as the whole of the second machine learning model (Daykin: [0073], "The same model is provided to each institution 40A , 40B , 40C , 40D"; the copies of the same model necessarily have the same organization and whole corresponding parts), and the performing, based on the first training sample set and at least a part of the model parameters of the present second machine learning model (Daykin: [0126-0127], " An output of FIG. 5 is a respective updated model for each institution…FIG . 5 shows only one training cycle. In practice, multiple training cycles may be performed…Each of the institutions…obtained by training the updated model on data held at that institution"; Daykin applies the repeated-cycle procedure at every institution, including the first and the second institutions with their respective cohort and updated model), the Nth round of training, to obtain the first machine learning model (Balachandar: [0045] and Fig. 2, "CWT, in accordance with many embodiments, involves starting training at one institution for a certain number of iterations, transferring the updated model weights to a subsequent institution, training the model at the subsequent institution for a certain number of iterations, then transferring the updated weights to the next institution, and so on until model convergence"; Balachandar supplies the express serial first site to second site initialization and repeated cyclic rounds) comprises: performing, based on the first training sample set and all of the model parameters of the present second machine learning model, the Nth round of model training, to obtain the first machine learning model (Daykin: [0123], "the aggregation circuitry 68 aggregates the model parameters obtained from each of the institutions"; Daykin teaches aggregation of the full identically structured peer parameter vector into the first updated model; wherein Balachandar: [0045] and Fig. 2, supplies the express serial first site to second site initialization and repeated cyclic rounds). Regarding dependent claim 5, Daykin, in view of Balachandar, teach the method of claim 2, further comprising: performing, if it is determined that the model index of the first machine learning model does not meet the first preset index, an (N+1)th round of model training of the first machine learning model (Daykin: [0130], "If the training is not complete, the flow chart returns to stage 90. The training circuitry 64 receives the updated model to train"; failure of the convergence/completion criterion begins the next training cycle; wherein Balachandar: [0045] and Fig. 2, supplies the express serial first site to second site initialization and repeated cyclic rounds). Regarding dependent claim 6, Daykin, in view of Balachandar, teach the method of claim 2, wherein the performing, based on the second training sample set, and at least a part of the model parameters of the present first machine learning model, an Mth round of model training, to obtain the second machine learning model, M being a positive integer greater than 0 (Daykin: [0127], "each institution sends its updated model to all of the institutions for training"; the second institution trains a current first site updated model on the second cohort during a training; wherein Balachandar: [0045] and Fig. 2, "transferring the updated model weights to a subsequent institution, training the model at the subsequent institution for a certain number of iterations, then transferring the updated weights to the next institution"; Balachandar expressly teaches the second site's received weight training and onward update); and terminating, if it is determined that a model index of the second machine learning model meets a second preset index, the training of the second machine learning model (Daykin: [0130], "At stage 104 , the training circuitry 64 determines whether training is complete"; Daykin's identical site-local algorithm applies the convergence/completion index to the second model; wherein Balachandar: [0045] and Fig. 2, supplies the express serial second site's update). Regarding dependent claim 7, Daykin, in view of Balachandar, teach the method of claim 6, wherein a part of the first machine learning model has the same structure as a part of the second machine learning model (Daykin: [0073], "The same model is provided to each institution 40A , 40B , 40C , 40D"; the copies of the same model necessarily have the same organization in at least corresponding parts), and the performing, based on the second training sample set and at least a part of the model parameters of the present first machine learning model, the Mth round of model training, to obtain the second machine learning model (Daykin: [0127], "each institution sends its updated model to all of the institutions for training"; the second institution trains a current first site updated model on the second cohort during a training; wherein Balachandar: [0045] and Fig. 2, "transferring the updated model weights to a subsequent institution, training the model at the subsequent institution for a certain number of iterations, then transferring the updated weights to the next institution"; Balachandar expressly teaches the second site's received weight training and onward update) comprise: performing, based on the second training sample set, model parameters of the part of the first initial model having the same structure, and second initial model parameters, a first round of model training, to obtain a second initial model, at least a part of model parameters of the second initial model being used to train the first machine learning model (Daykin: [0081-0083] and [0126], "At each institution, the updated model is formed by aggregating model parameters from all of the institutions in accordance with the influence values"; the second site's initial local update combines its own trained parameters and compatible first-site parameters, and the all to all cycle returns that updated state to the first site); and continuing to perform, based on the second training sample set, model parameters of the part of the present first machine learning model having the same structure, and a part of the model parameters of the second machine learning model obtained from a previous round of model training, the model training, to obtain the second machine learning model (Daykin: [0127], "The updated influence values are then aggregated. The cycle may be repeated multiple times, for example until models converge"; repeated local aggregation retains the second site's own prior state and incorporates the compatible current first site state). Regarding dependent claim 8, Daykin, in view of Balachandar, teach the method of claim 6, wherein the whole of the first machine learning model has the same structure as the whole of the second machine learning model (Daykin: [0073], "The same model is provided to each institution 40A , 40B , 40C , 40D"; the copies of the same model necessarily have the same organization and whole corresponding parts), and the performing, based on the second training sample set and at least a part of the model parameters of the present first machine learning model, the Mth round of model training, to obtain the second machine learning model (Daykin: [0127], "each institution sends its updated model to all of the institutions for training"; the second institution trains a current first site updated model on the second cohort during a training; wherein Balachandar: [0045] and Fig. 2, "transferring the updated model weights to a subsequent institution, training the model at the subsequent institution for a certain number of iterations, then transferring the updated weights to the next institution"; Balachandar expressly teaches the second site's received weight training and onward update) comprises: performing, based on the second training sample set and all of the model parameters of the present first machine learning model, the Mth round of model training, to obtain the second machine learning model (Daykin: [0126], "At each institution, the updated model is formed by aggregating model parameters from all of the institutions in accordance with the influence values"; the full first-site vector is one of the vectors aggregated into the second updated model; wherein Balachandar: [0045] and Fig. 2, "transferring the updated model weights to a subsequent institution, training the model at the subsequent institution for a certain number of iterations, then transferring the updated weights to the next institution"; Balachandar expressly teaches the second site's received weight training and onward update). Regarding dependent claim 9, Daykin, in view of Balachandar, teach the method of claim 6, further comprising performing, if it is determined that the model index of the second machine learning model does not meet the second preset index, an (M+1)th round of model training of the second machine learning model (Daykin: [0130], "If the training is not complete, the flow chart returns to stage 90. The training circuitry 64 receives the updated model to train"; at the second institution, failure of the completion index likewise begins the next cycle; wherein Balachandar: [0045] and Fig. 2, "transferring the updated model weights to a subsequent institution, training the model at the subsequent institution for a certain number of iterations , then transferring the updated weights to the next institution"; Balachandar expressly teaches the second site's received weight training and onward update). Regarding dependent claim 11, Daykin, in view of Balachandar, teach the method of claim 6, wherein the method further comprises: determining, during each round of model training, a descent gradient using a batch gradient algorithm if it is determined, according to a preset loss function, that an output result of the machine learning model does not meet a preset convergence condition (Daykin: [0079], "Any suitable model training process may be used . For example, stochastic gradient descent may be used"; Daykin teaches descent-gradient optimization during local model training; Balachandar: [0042], "Various embodiments use minibatch sampling with an appropriate batch size. In some of these embodiments, the batch size is 32. Many embodiments use an optimization algorithm for model weight optimization. Certain embodiments use the Adam optimization algorithm"; minibatch Adam determines descent gradients for model weight optimization, and the same paragraph specifies cross entropy loss and a validation loss stopping condition), and continuing to perform the model training, and stopping, until it is determined that the output result of the machine learning model meets the preset convergence condition, the present round of model training (Daykin: paragraph 0130, "If the training is not complete, the flow chart returns to stage 90"; Daykin teaches continuing model training while the convergence/completion condition is unsatisfied; Balachandar: [0042], "Additionally, some embodiments terminate model learning early, if an amount of iterations or epochs pass without an improvement in validation loss"; absence of validation-loss improvement is a preset loss-convergence condition that terminates the local round after continued minibatch optimization). Regarding dependent claim 15, Daykin discloses the method of claim 1, wherein training performed on two independent networks, respectively (Daykin: [0059] “In general, access to the training image data sets and other image data sets stored in the data store 60 is restricted to a single institution. The apparatus 50 only has access to data from its own institution”, [0067-0068] "In other embodiments, any suitable number or type of communications networks may be used to connect the apparatuses at institutions 40A, 40B, 40C, 40D", [0072] “Similarity is calculated by training a model at each institution; Daykin teaches respective hospital local apparatuses and cohorts connected through one or more network domains wherein training a model at each institution with access to data sets stored in the data store restricted to each institution). Daykin does not expressly teach training of the first machine learning model and the training of the second machine learning model are performed on two independent networks, respectively. However, Balachandar teaches the training of the first machine learning model and the training of the second machine learning model are performed independently (Balachandar: [0037], [0040] "federated learning of many embodiments train Al models on local patient data , and numeric model parameters ( weights ) are transferred between institutions instead of patient data"; Balachandar teaches maintaining patient data and training locally at separate institutions while only parameter values cross the boundary). Because Daykin and Balachandar are analogous art concerning privacy-preserving distributed neural-network training across medical institutions using local medical-image datasets and exchanged parameter values, addressing the same problem of improving models without pooling sensitive institutional image data, accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention, to apply Balachandar's training of a first and second model locally at separate institutions to Daykin's training local apparatus on two independent networks respectively to teach wherein the training of the first machine learning model and the training of the second machine learning model are performed on two independent networks, respectively. This modification would have been motivated by the desire to provide specialized approaches for distributed deep learning for medical applications (Balachandar: [0004]). Regarding dependent claim 16, Daykin teaches all the elements of claim 1. Daykin does not expressly teach wherein the training of the first machine learning model and the training of the second machine learning model are performed alternately. However, Balachandar teaches wherein the training of the first machine learning model and the training of the second machine learning model are performed alternately (Balachandar: [0043], [0045] and Fig. 2, "CWT allows for synchronous , non-parallel training"; Balachandar's CWT trains at one institution and then a subsequent institution in repeated cyclic order). Because Daykin and Balachandar are analogous art because both concern privacy-preserving distributed neural-network training across medical institutions using local medical-image datasets and exchanged parameter values and their disclosures address the same problem of improving models without pooling sensitive institutional image data, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, to apply Balachandar's serial CWT schedule to Daykin's compatible all to all parameter updates to teach wherein the training of the first machine learning model and the training of the second machine learning model are performed alternately. This modification would have been motivated by the desire to provide specialized approaches for distributed deep learning for medical applications (Balachandar: [0004]). Claim 10 is rejected under 35 U.S.C. 103 as unpatentable over Daykin, in view of Balachandar, as applied in the rejection of claim 6 above, and further in view of KALE et al. (hereinafter Kale), US 2020/0242511 A1. Regarding dependent claim 10, Daykin, in view of Balachandar, teach all the elements of claim 6. Daykin and Balachandar do not expressly teach wherein the model index comprises an accuracy of an output result, the first preset index comprises a first preset accuracy, and the second preset index comprises a second preset accuracy. However, Kale teaches wherein the model index comprises an accuracy of an output result, the first preset index comprises a first preset accuracy, and the second preset index comprises a second preset accuracy (Kale: [0056] "it can be determined whether the accuracy calculated for the trained machine learning model is greater than a threshold accuracy (e.g., 25%, 50%, 75%, and the like)"; Kale expressly uses model output accuracy as the measured index and a predetermined accuracy threshold as the preset index which suggests applying same rule to respective models to supply first and second thresholds). Because Daykin, in view of Balachandar, and Kale are analogous art concerned with iterative machine learning training, model-performance evaluation, and the decision whether to continue or retrain a model, Kale’s accuracy based control is reasonably pertinent to implementing the convergence decisions in the medical federated training combination of Daykin, in view of Balachandar, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use Kale's accuracy threshold for each of Daykin and Balachandar’s local model to teach wherein the model index comprises an accuracy of an output result, the first preset index comprises a first preset accuracy, and the second preset index comprises a second preset accuracy. This modification would have been motivated by the desire to balance accuracy with resource efficiency and practicality via selective retraining (Kale: [0004]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Chang et al. “Distributed deep learning networks among institutions for medical imaging” (2018) (Abstract Objective: Deep learning has become a promising approach for automated support for clinical diagnosis. When medical data samples are limited, collaboration among multiple institutions is necessary to achieve high algorithm performance. However, sharing patient data often has limitations due to technical, legal, or ethical concerns. In this study, we propose methods of distributing deep learning models as an attractive alternative to sharing patient data. Methods: We simulate the distribution of deep learning models across 4 institutions using various training heuristics and compare the results with a deep learning model trained on centrally hosted patient data. The training heuristics investigated include ensembling single institution models, single weight transfer, and cyclical weight transfer. We evaluated these approaches for image classification in 3 independent image collections (retinal fundus photos, mammography, and ImageNet). Results: We find that cyclical weight transfer resulted in a performance that was comparable to that of centrally hosted patient data. We also found that there is an improvement in the performance of cyclical weight transfer heuristics with a high frequency of weight transfer. Conclusions: We show that distributing deep learning models is an effective alternative to sharing patient data. This finding has implications for any collaborative deep learning study). Roy et al. “A peer-to-peer environment for decentralized federated learning” (2019) (Abstract Access to sufficient annotated data is a common challenge in training deep neural networks on medical images. As annotating data is expensive and time-consuming, it is difficult for an individual medical center to reach large enough sample sizes to build their own, personalized models. As an alternative, data from all centers could be pooled to train a centralized model that everyone can use. However, such a strategy is often infeasible due to the privacy-sensitive nature of medical data. Recently, federated learning (FL) has been introduced to collaboratively learn a shared prediction model across centers without the need for sharing data. In FL, clients are locally training models on site-specific datasets for a few epochs and then sharing their model weights with a central server, which orchestrates the overall training process. Importantly, the sharing of models does not compromise patient privacy. A disadvantage of FL is the dependence on a central server, which requires all clients to agree on one trusted central body, and whose failure would disrupt the training process of all clients. In this paper, we introduce BrainTorrent, a new FL framework without a central server, particularly targeted towards medical applications. BrainTorrent presents a highly dynamic peer-to-peer environment, where all centers directly interact with each other without depending on a central body. We demonstrate the overall effectiveness of FL for the challenging task of whole brain segmentation and observe that the proposed server-less BrainTorrent approach does not only outperform the traditional server-based one but reaches a similar performance to a model trained on pooled data). Any inquiry concerning this communication or earlier communications from the examiner should be directed to KUANG FU CHEN whose telephone number is (571)272-1393. The examiner can normally be reached M-F 9:00-5:30pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached on (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KC CHEN/Primary Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Jan 12, 2024
Application Filed
Jul 27, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682047
Machine Learning Time Series Anomaly Detection
4y 1m to grant Granted Jul 14, 2026
Patent 12675771
System for Online Interaction with Content
5y 8m to grant Granted Jul 07, 2026
Patent 12664448
AVERAGE TREATMENT EFFECT FOR PAIRED DATA
4y 11m to grant Granted Jun 23, 2026
Patent 12657260
SIMULATING TRAINING DATA TO MITIGATE BIASES IN MACHINE LEARNING MODELS
4y 1m to grant Granted Jun 16, 2026
Patent 12657494
LEARNING SYSTEM, LEARNING METHOD, AND STORAGE MEDIUM
2y 12m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+68.4%)
2y 11m (~4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 270 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month