DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner's Note
The Examiner respectfully requests of the Applicant in preparing responses, to fully consider the entirety of the reference(s) as potentially teaching all or part of the claimed invention. It is noted, REFERENCES ARE RELEVANT AS PRIOR ART FOR ALL THEY CONTAIN. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain.” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)). A reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art, including non-preferred embodiments (see MPEP 2123). The Examiner has cited particular locations in the reference(s) as applied to the claim(s) above for the convenience of the Applicant. Although the specified citations are representative of the teachings of the art and are applied to the specific limitations within the individual claim(s), typically other passages and figures will apply as well.
Response to Argument
Applicant’s arguments, see REMARKS page 6-8 filed 01/20th/2026, regarding the rejection of claims 1-20 under 35 U.S.C. §103 have been considered and they are not persuasive.
Applicant argument #1
The applicant argues that “the Office Action took an official notice that "[t]he number of
data samples from each collaborator are aggregated into a single central model to obtain the third sum." (Page 4.)” Further, the Applicant argues “Yoo nowhere discloses that the number of data points are transferred to the aggregation server 121. Thus, the Examiner is respectfully requested to provide a documentary evidence that Yoo discloses transferring the third sum of the number of data points to the aggregation server. Unless the documentary evidence is provided in support of an officially-noticed fact, Yoo fails to teach or suggest the above-recited features of independent claims 1 and 11.”
Examiner Response #1
As an initial matter, in the Office Action dated 28 October 2025, the examiner does not invoke Official Notice. The portions of the rejection cited by the applicant neither indicate that the examiner has taken official notice nor assert a fact that the examiner considers to be well-known or common knowledge.
The applicant appears to interpret the “Examiner’s Note” that “The number of data samples from each collaborator are aggregated into a single central model to obtain the third sum” as invoking Official Notice. Instead, the “Examiner’s Note” is merely providing the examiner’s interpretation of YOO. Specifically, YOO [0033] further explains that In an embodiment, the aggregation weights may also depend on a number of data points from each collaborator 131. Collaborators 131 with more data may be accorded a higher weight when calculating the aggregated parameters for the central model. The examiner notes that the aggregation of the weights that depend on a number of data points from each collaborator is the equivalent to the claimed thirst sum.
Finally, although the applicant argues that YOO fails to teach that the number of data points are transferred to the aggregation server. However, the claimed invention does not claim transferring data points to the aggregation server but rather it claims tracking a third sum of a number of data points to the local model. Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yoo et al. (US 20230101741 A1) in view of Laszlo et al. (US 20220202348 A1) in further view of Reyes et al. (US 20240127114 A1), hereinafter referred to as Yoo, Laszlo, and Reyes, respectively.
Regarding Claim 1:
A method comprising: tracking, for a local model operating on a local node, a first sum of inputs to a batch normalization layer, a second sum of squares of the inputs to the batch normalization layer, and a third sum of a number of data points to the local model; ([0020]
“Aggregation weights are computed for each collaborator 131 based on the estimated divergence. Additionally, aggregation weights may be further adjusted by class imbalance ratio (for example a ratio between positive and negative samples) and number of data samples of each collaborator 131.”
[0031] “The model divergence is estimated by measuring aggregation validation metrics for each collaborator 131 at each communication round… The divergence may be calculated, for example, using a L1 norm or L2 norm.”
[0032] “The L1 norm is calculated as the sum of the absolute values of a vector for the respective model parameters. The L2 norm is calculated as the square root of the sum of the squared vector values…”
[0038] “The 3D input tensor is fed into a 3D 1 3 3 convolutional layer followed by batch normalization and leaky ReLU.”
Examiner’s Note: The aggregation weights computed for each collaborator is read as the tracked three sums of inputs. More specifically, the number of data samples is read as the third sum of data points to the local model. The L1 norm which is the sum of absolute values of a vector is read as the first sum of inputs to a batch normalization layer. The L2 norm which is the square root of the sum of the squared vector values is read as the second sum of square of inputs.
securely aggregating, by a central node, the first sum, the second sum, and the third sum from an instance of the local model operating on the local node with first sums, second sums, and third sums from other instances of local models operating on other local nodes participating in federated learning to generate an aggregated first sum, an aggregated second sum, and an aggregated third sum; (([0020] “Aggregation weights are computed for each collaborator 131 based on the estimated divergence. Additionally, aggregation weights may be further adjusted by class imbalance ratio (for example a ratio between positive and negative samples) and number of data samples of each collaborator 131.”
[0021] FIG. 2 depicts a method for aggregating model parameters from a plurality of collaborators 131 in a federated learning system that trains a model over multiple rounds of training. For each round, the collaborators 131 or remote devices train a local model with local data. The parameters are then sent to an aggregation server 121/aggregator that is configured to aggregate the parameters from multiple collaborators 131 into a single central model.
[0031] “Referring back to FIG. 2, at act A120, the aggregation server 121 calculates for each of the two or more collaborators 131 a model divergence value that approximates how much an updated collaborator model for a respective collaborator 131 of the two or more collaborators 131 deviates from a prior aggregated model. The model divergence is estimated by measuring aggregation validation metrics for each collaborator 131 at each communication round… The divergence may be calculated, for example, using a L1 norm or L2 norm.”
Examiner’s Note: Yoo teaches a federated learning system in par. 21 and the parameters obtained from remote devices with local models and data are sent to an aggregation server, which is where all the sums are aggregated. The computed aggregated statistics are labeled as L1 norm and L2 norm which are read as the first and second sums, respectively. The number of data samples from each collaborator are aggregated into a single central model to obtain the third sum.)
and updating, by the local node and the other local nodes, the batch normalization layer of the local model and batch normalization layers of the other instances of the local model, respectively, with the global mean and the global variance. ([0041] “The aggregation server 121 is configured to transmit the aggregated parameters (global parameters) to the one or more collaborators 131 using the transceiver 127.”
Examiner’s Note: The aggregated parameters, AKA global parameters, are read as the global mean and global variance. The transmitting of these global parameters is read as updating the batch normalization layers of each local model.
Yoo fails to teach: determining a global mean and a global variance from the first, second, and third aggregated sums;
However, Laszlo teaches: determining, by the central node, a global mean… from the first, second, and third aggregated sums; ([0153] “The global parameter updating engine 616 is configured to obtain the respective locally-updated parameter values 608 from each of one or more user devices 602a-n and use the sets of locally updated parameter values 608 to determine an update to the parameter values of the neural network stored in the global model parameter store 618. The global parameter updating engine 616 can combine i) the current version of the parameter values stored in the global model parameter store 618 and ii) the one or more sets of locally-updated parameter values 608, in any appropriate way. For example, the global parameter updating engine 616 can determine a weighted mean of different versions of the parameter values. As a particular example, the global parameter updating engine can weight the version stored in the global model parameter store 618 more than each of the locally-updated versions 608.”
Examiner’s Note: The locally-updated parameter values from one or more devices are read as the aggregated sums. The weighted mean of different versions of the parameter values is read as a global mean based on the aggregated sums.)
Yoo and Laszlo are considered to be analogous to each other as they are all in the field of machine learning. Therefore, before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the federated learning method taught by Yoo with the feature of determining a global mean based on aggregated data taught by Laszlo in order to account for the updated values from local devices based on the appropriate weight, which takes into account number of samples – AKA the third sum. ([0153] “The global parameter updating engine 616 can combine i) the current version of the parameter values stored in the global model parameter store 618 and ii) the one or more sets of locally-updated parameter values 608, in any appropriate way. For example, the global parameter updating engine 616 can determine a weighted mean of different versions of the parameter values… As another particular example, the global parameter updating engine 616 can weight respective versions based on how many training examples were used to generate them.”)
Laszlo fails to teach: determining… and a global variance from the first, second, and third aggregated sums;
However, Reyes teaches: determining… and a global variance from the first, second, and third aggregated sums; ([0213] “In this section, the relationship between the variance of weights across clients and the accuracy of the aggregated model was analyzed. FIG. 9 shows the average of the variances across weights and clients that was used to update the central weights of the model, as well as, the test accuracy against communication rounds. Interestingly, abrupt decreases in the accuracy of the method are correlated with a sudden increase in the aggregated variance.”
Examiner’s Note: The average of the variance across clients, which have been aggregated into a central model, is read as the global variance. This average of average is also referred to as the aggregated variance.)
Yoo and Reyes are considered to be analogous to each other as they are all in the field of machine learning. Therefore, before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the federated learning method taught by Yoo with the feature of determining a global variance based on aggregated data taught by Reyes in order to, similar to Laszlo’s teaching, attribute weight according to the relative contribution of local datasets from the local models, which accounts for number of samples – AKA the third sum. ([0010] “Developer(s) have also appreciated that by using a training parameter including the second raw moment or uncentered variance of the stochastic gradient, the individual intra-variability expressed during the training on local data may be estimated, and the weighted average of the trained models may be computed. By using such a training parameter, the relative contribution of the local datasets in the global training may be at least approximated.)
Regarding Claim 2:
Yoo fails to teach: The method of claim 1, further comprising updating a batch normalization layer of a central model with the global mean and the global variance.
However, Reyes teaches: The method of claim 1, further comprising updating a batch normalization layer of a central model with the global mean and the global variance. ([0207] “The second network was used for image recognition tasks from the CIFAR-10 dataset. The architecture consisted of one 3×3 convolutional layer (with 32 convolution filter using a ReLu activation), followed with a 2×2 max pooling, a batch normalization layer;”
[0010] Developer(s) have also appreciated that by using a training parameter including the second raw moment or uncentered variance of the stochastic gradient, the individual intra-variability expressed during the training on local data may be estimated, and the weighted average of the trained models may be computed. By using such a training parameter, the relative contribution of the local datasets in the global training may be at least approximated.
[0213] “In this section, the relationship between the variance of weights across clients and the accuracy of the aggregated model was analyzed. FIG. 9 shows the average of the variances across weights and clients that was used to update the central weights of the model, as well as, the test
Examiner’s Note: [0207] teaches that batch normalization layers are present within the neural network. The parameters obtained from the local models are then stored in the central server in order to update the neural network, which includes batch normalization layers. [0010] and [0213] show the presence of the weighted average which is read as the global mean and the average of the variances which is read as the global variance.
Yoo and Reyes are considered to be analogous to each other as they are all in the field of machine learning. Therefore, before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the federated learning method taught by Yoo with the feature of updating a batch normalization layer of a central model specifically using the global mean and global variance taught by Reyes in order to attribute weight according to the relative contribution of local datasets from the local models, which accounts for number of samples – AKA the third sum. ([0010] “Developer(s) have also appreciated that by using a training parameter including the second raw moment or uncentered variance of the stochastic gradient, the individual intra-variability expressed during the training on local data may be estimated, and the weighted average of the trained models may be computed. By using such a training parameter, the relative contribution of the local datasets in the global training may be at least approximated.”)
Regarding Claim 3:
Yoo further teaches: The method of claim 1, further comprising determining a sum of gradients for layers in the local model. ([0018] “The aggregation of the parameters typically uses an averaging mechanism. For example, one popular approach is Federated Averaging (FedAvg) where the model weights of the different local models are averaged by the aggregation server 121 to provide new model weights and, thus, a new aggregated model. In an example, at each iteration, FedAvg first locally performs epochs of stochastic gradient descent (SGD) on the devices.
[0047] “Different optimization algorithms may be used to minimize the loss function, such as, for example, gradient descent, Stochastic gradient descent, Batch gradient descent, Mini-Batch gradient descent, among others.”
Examiner’s Note: Batch gradient descent is read as determining a sum of gradients for layers in a local model.
Regarding Claim 4:
Yoo further teaches: The method of claim 3, further comprising aggregating the sum of gradients with sums of gradients from the other instances of the local model. ([0018] “The aggregation of the parameters typically uses an averaging mechanism. For example, one popular approach is Federated Averaging (FedAvg) where the model weights of the different local models are averaged by the aggregation server 121 to provide new model weights and, thus, a new aggregated model. In an example, at each iteration, FedAvg first locally performs epochs of stochastic gradient descent (SGD) on the devices.
[0045] “The processor 123 is configured to generate a global model that may eventually be applied for a certain task. The global model uses the aggregated weights and may be trained over multiple rounds of communications between the aggregation server 121 and the collaborators 131.”
[0047] “Different optimization algorithms may be used to minimize the loss function, such as, for example, gradient descent, Stochastic gradient descent, Batch gradient descent, Mini-Batch gradient descent, among others.”
Examiner’s Note: Batch gradient descent is read as determining a sum of gradients for layers in a local model. [0045] also teaches aggregating those gradients from multiple collaborators.
Regarding Claim 5:
Yoo teaches: The method of claim 1, further comprising determining the first sum across each input dimension. ([0038] “The 3D input tensor is fed into a 3D 1 3 3 convolutional layer followed by batch normalization and leaky ReLU. The feature maps were then propagated to 5 DenseNet blocks.
[0039] “The network includes a ResNet50 as the backbone axial feature extractor, that takes a series of CT in-plane slices as input and generated feature maps for the corresponding slices. The extracted features from all slices are then combined by a max-pooling operation. The global feature is fed to a fully connected layer that produces a COVID-19 prediction score per case by softmax operation.”
Examiner’s Note: The inputs are fed through multiple layers of the neural network. The extracted features from these inputs are then combined at the end to obtain a sum of inputs, which is read as the first sum.
Regarding Claim 6:
Yoo teaches: The method of claim 1, further comprising determining the second sum across each input dimension. ([0032] “The L2 norm is calculated as the square root of the sum of the squared vector values… By estimating model divergence for each collaborator 131 at each round using a preserved model divergence testing dataset, embodiments perform adaptive aggregation that can be robust to non-IID datasets between participating collaboration sites.”
[0045] “The global model uses the aggregated weights and may be trained over multiple rounds of communications between the aggregation server 121 and the collaborators 131… The model(s) may include a neural network that is defined as a plurality of sequential feature units or layers… Rather than pre-programming the features and trying to relate the features to attributes, the deep architecture is defined to learn the features at different levels of abstraction based on the input data.”
Examiner’s Note: The specification lists features as an alternative term to input dimension in [0022]. The model can consist of sequential layers and architecture is defined to train on input data to recognize features. The L2 norm calculations are performed on the models. Therefore, it is read that the sum of the squared vector values are aggregated across each input dimension.)
Regarding Claim 7:
Yoo further teaches: The method of claim 1, further comprising determining the global mean and the global variance without sharing the inputs from the local model or the other local models. ([0020] “Embodiments described herein provide for distributed processing of data while maintaining privacy and transmission concerns. The training occurs in a decentralized manner with multiple collaborators 131 with only the local data available to each collaborator 131. The multiple collaborators 131 do not share data. The aggregation of model parameters occurs on an aggregation server 121. Aggregation is performed by weighted average, in which the aggregation weight is adaptively computed. Embodiments utilize a model divergence value that approximates how much an updated collaborator model deviates from a previous aggregated model and adjusts aggregation weights based on the approximated divergence.”
Examiner’s Note: Each collaborator has a local model which conducts its own training without sharing data with other collaborators. The local data is read as the “inputs”. The aggregated model parameters that are performed on the aggregation server is read to include the global mean and global variance since these are “global” parameters computed based on all of the aggregated information.)
Regarding Claim 8:
Yoo further teaches: The method of claim 1, wherein updating the batch normalization layer comprises synchronizing the central model to each of the nodes. ([0006] “In a second aspect, a system is provided for federated learning. The system includes a plurality of collaborators and an aggregation server… The aggregation server is configured to receive the updated model weights from the plurality of collaborators, calculate a model divergence value for each collaborator from respective updated model weights and a prior model, calculate aggregated model weights based at least in part on the model divergence values, and transmit the aggregated model weights to the plurality of collaborators to update the local machine learned model.
[0038] “In an embodiment, the segmentation network includes a U-Net resembling architecture with 3D convolution blocks containing either 1 3 3 or 3 3 3 CNN kernels to deal with anisotropic resolutions. The 3D input tensor is fed into a 3D 1 3 3 convolutional layer followed by batch normalization and leaky ReLU.”
Examiner’s Note: Each of the local machines/collaborators are read as node. The aggregation model on the aggregation server is read as the central model. Every collaborator has their machine learned model updated, which includes layers of batch normalization.)
Regarding Claim 9:
Yoo fails to teach: The method of claim 1, wherein the nodes are associated with a domain where statistics change faster than a threshold change rate.
However, Reyes teaches: The method of claim 1, wherein the nodes are associated with a domain where statistics change faster than a threshold change rate. ([0008] “Nevertheless, developers have acknowledged that there is evidence that demonstrates a misleading interpretation of results and a reduction of statistical power when combining data from different sources without accounting for variation across sources. Therefore, despite the encouraging results produced with the Federated Averaging algorithm, developers believe that this technique underestimates the full extent of heterogeneity on domains where data is complex with a large diversity of features in its composition.
[0010] “Developer(s) have also appreciated that by using a training parameter including the second raw moment or uncentered variance of the stochastic gradient, the individual intra-variability expressed during the training on local data may be estimated, and the weighted average of the trained models may be computed. By using such a training parameter, the relative contribution of the local datasets in the global training may be at least approximated.”
[0011] “The present technology enables improving training of machine learning models by penalizing the model uncertainty at the client level to improve the robustness of the aggregated model, regardless of the data distribution: independent and identically distributed (IID) or non-identical and non-independent (Non-IID). The weighted average at each communication step may be computed by using the uncentered variance of the gradient estimator from the Adam optimizer.”
Examiner’s Note: The Specification does not describe in detail what the threshold change rate is. It does describe that data from domains with frequent influx of sensitive data that are not too heterogenous is beneficial for training. The calculated stochastic gradient for each node is read as the threshold rate of change. Reyes describes improving training by penalizing models based on their calculated uncertainty level which is computed by using gradient estimator.)
Yoo and Reyes are considered to be analogous to each other as they are all in the field of machine learning. Therefore, before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the federated learning method taught by Yoo with the feature of nodes being associated with domains with fast-changing statistics in order to have better training and therefore better accuracy ([0008] “Meta-analysis is a quantitative method that combines results from different studies on the same topic in order to draw a general conclusion and to evaluate the consistency among study findings. Nevertheless, developers have acknowledged that there is evidence that demonstrates a misleading interpretation of results and a reduction of statistical power when combining data from different sources without accounting for variation across sources.”
[0210] “The improved accuracy on CIFAR-10 could indicate that there is greater heterogeneity in models trained on natural images than in models trained on grayscale images, even in an IID setting.”)
Regarding Claim 10:
Yoo fails to teach: The method of claim 1, wherein the global mean and the global variance are configured to federate the batch normalization layer in the local models.
However, Reyes teaches: The method of claim 1, wherein the global mean and the global variance are configured to federate the batch normalization layer in the local models. ([0009] “More specifically, developers have appreciated that a federated learning system could be used to obtain a robust global or aggregated machine learning model generated via combination of machine learning models trained locally on local datasets (i.e. using different processing devices holding respective training datasets),”
[0010] “Developer(s) have also appreciated that by using a training parameter including the second raw moment or uncentered variance of the stochastic gradient, the individual intra-variability expressed during the training on local data may be estimated, and the weighted average of the trained models may be computed. By using such a training parameter, the relative contribution of the local datasets in the global training may be at least approximated.”
[0132] “As another non-limiting example, to accomplish the predictive task on the CIFAR-10 dataset, the architecture of the initial model 310 may be based on a CNN Keras sequential model comprising one 3×3 convolutional layer (with 32 convolution filter using a ReLu activation), followed with a 2×2 max pooling, a batch normalization layer;”
[0178] During the training procedure iteration, each of the first node 220, the second node 230 and the third node 240 train the aggregated trained model 350 on a respective training dataset to obtain an “updated” respective trained model.
[0213] “In this section, the relationship between the variance of weights across clients and the accuracy of the aggregated model was analyzed. FIG. 9 shows the average of the variances across weights and clients that was used to update the central weights of the model, as well as, the test accuracy against communication rounds. Interestingly, abrupt decreases in the accuracy of the method are correlated with a sudden increase in the aggregated variance.”
Examiner’s Note: In [0009], there’s an explicit mentioning of a federated learning system. In [0010], the average of the trained models is read as the global mean. In [0213], the average of the variances across the weights and clients is read as the global variance. [0132] teaches a model architecture that has at least one batch normalization layer. [0178] teaches the federated learning system, specifically updating the local models based on the aggregated statistics.)
Yoo and Reyes are considered to be analogous to each other as they are all in the field of machine learning. Therefore, before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the federated learning method taught by Yoo with the feature of updating a batch normalization layer of a central model specifically using the global mean and global variance taught by Reyes in order to attribute weight according to the relative contribution of local datasets from the local models, which accounts for number of samples – AKA the third sum. It also contributes to a robust global machine learning model through combination of locally trained learning models. ([0009] “More specifically, developers have appreciated that a federated learning system could be used to obtain a robust global or aggregated machine learning model generated via combination of machine learning models trained locally on local datasets (i.e. using different processing devices holding respective training datasets)…” [0010] “Developer(s) have also appreciated that by using a training parameter including the second raw moment or uncentered variance of the stochastic gradient, the individual intra-variability expressed during the training on local data may be estimated, and the weighted average of the trained models may be computed. By using such a training parameter, the relative contribution of the local datasets in the global training may be at least approximated.”)
Regarding Claim 11:
The claim is rejected on the same grounds as Claim 1 for reciting substantially similar limitations, with the exception of the limitation:
Yoo teaches: A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising: ([0042] “The instructions for implementing the processes, methods and/or techniques discussed herein are provided on non-transitory computer-readable storage media or memories, such as a cache, buffer, RAM, removable media, hard drive, or other computer readable storage media.”)
Regarding Claim 12:
The claim is rejected on the same grounds as Claim 2 for reciting substantially similar limitations.
Regarding Claim 13:
The claim is rejected on the same grounds as Claim 3 for reciting substantially similar limitations.
Regarding Claim 14:
The claim is rejected on the same grounds as Claim 4 for reciting substantially similar limitations.
Regarding Claim 15:
The claim is rejected on the same grounds as Claim 5 for reciting substantially similar limitations.
Regarding Claim 16:
The claim is rejected on the same grounds as Claim 6 for reciting substantially similar limitations.
Regarding Claim 17:
The claim is rejected on the same grounds as Claim 7 for reciting substantially similar limitations.
Regarding Claim 18:
The claim is rejected on the same grounds as Claim 8 for reciting substantially similar limitations.
Regarding Claim 19:
The claim is rejected on the same grounds as Claim 9 for reciting substantially similar limitations.
Regarding Claim 20:
The claim is rejected on the same grounds as Claim 10 for reciting substantially similar limitations.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
LI – (Privacy-preserving Spatiotemporal Scenario Generation of Renewable Energies: A Federated Deep Generative Learning Approach)
“LI teaches a novel federated deep generative learning framework, called Fed-LSGAN that integrates federated learning and least square generative adversarial networks (LSGANs) for renewable scenario generation. Specifically, federated learning learns a shared global model in a central server from renewable sites at network edges, which enables the Fed-LSGAN to generate scenarios in a privacy-preserving manner without sacrificing the generation quality by transferring model parameters, rather than all data”
IDRISSI – (FEDBS: Learning on Non-IID Data in Federated Learning using Batch Normalization)
“IDRISSI teaches FedBS, a new efficient strategy to handle global models having batch normalization layers, in the presence of Non-IID data. FedBS modifies FedAvg by introducing a new aggregation rule at the server-side, while also retaining full compatibility with Batch Normalization (BN)”
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHAMCY ALGHAZZY whose telephone number is (571)272-8824. The examiner can normally be reached on M-F 7:30am-5:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, OMAR FERNANDEZ RIVAS can be reached on (571) 272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHAMCY ALGHAZZY/Examiner, Art Unit 2128
/KYLE R STORK/Primary Examiner, Art Unit 2128