DETAILED ACTION
This communication is in response to Application No. 18/325,374 filed on May 30, 2023, in which claims 1-25 are presented for examination.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement submitted on 05/30/2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement was considered by the examiner.
Specification
The contents of the specification are sufficient for examination purposes.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 6-7, 13, and 18 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Regarding Claim 6, the claim recites the limitation “removing at least one of the one or more machine learning models” (ln. 9). There is insufficient antecedent basis for this limitation in the claim. Specifically, the claim recites both “one or more machine learning models in the model collection” (ln. 5) and “one or more other machine learning models in the model collection” (ln. 7). As a result, it is not clear which of the two sets of “one or more machine learning models” is subject to the “removing”. Therefore, the scope of the claim is indefinite. As a result, the claim is rejected. The claim should be amended to clarify which set of “one or more machine learning models” is subject to the “removing”.
Regarding Claim 7, the claim is rejected because it is dependent on a rejected claim.
Regarding Claim 13, the claim recites “removing at least one of the one or more machine learning models” (ln. 9), which is indefinite for substantially the same reasoning as discussed in regard to the rejection of claim 6. As a result, the claim is similarly rejected and should be amended in a similar manner.
Regarding Claim 18, the claim recites “removing at least one of the one or more machine learning models” (ln. 8), which is indefinite for substantially the same reasoning as discussed in regard to the rejection of claim 6. As a result, the claim is similarly rejected and should be amended in a similar manner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 6-7, 10-13, 15-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (hereinafter Zhang) (“Towards a Federated Learning Framework for Heterogeneous Devices of Internet of Things”) in view of Chen et al. (hereinafter Chen) (Patent No. US 12,438,786 B2).
Regarding Claim 1, Zhang teaches a computer-implemented method, comprising (Pg. 1, Col. 2 Fig. 1 and Pg. 1, Col. 1, Abstract, “we propose an FL framework targeting the heterogeneity of IoT devices . . . We conduct preliminary experiments to illustrate that our framework can facilitate the design of IoT-aware FL”, where “conduct[ing] preliminary experiments” using a “FL framework” is performance of a method, which is implemented by a computer, see Pg. 3, Col. 2, Para. 1, “All experiments are conducted in a Lenovo ThinkBook laptop (8 2.4GHz CPU cores and 16GB memory)”):
PNG
media_image1.png
206
404
media_image1.png
Greyscale
Zhang, Figure 1
receiving, by an online system . . . information associated with one or more computational resources available at the computing device (Pg. 1-2, Col. 1-1, Para. 5-1, “Our framework differs from existing works in that our local models are compressed from the global model, in order to meet the resource constraints of different devices. In other words, local models are compressed with different compression techniques (e.g., pruning and quantization) to different degrees (e.g., different pruning ratios) . . . Due to the device heterogeneity, different IoT devices need to run different compressed versions of the global model. For example, an IoT hub can afford sophisticated models, whereas an embedded device can only run lightweight models”, where information associated with one or more computational resources, “the resource constraints”, at each computing device, “of different devices”, is used to perform model “compress[ion] with different compression techniques (e.g., pruning and quantization) to different degrees”; see also Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning”, where Fig. 1, reproduced above, depicts performance of “Model Compression” as occurring at the “Server”, thus, since the information associated with one or more computational resources is used to perform model compression, the server must receive this information in order to perform compression, “local models are compressed”, based on the “need[s]” of the “different IoT devices”; see also Pg. 1, Col. 1, Abstract, “Internet of Things (IoT) devices are inherently diverse regarding computation speed and onboard memory. In this paper, we propose an FL framework targeting the heterogeneity of IoT devices . . . [to] facilitate the design of IoT-aware FL”, where the “Server” is an online system when operably connect to “Internet” as part of an “IoT-aware” “system architecture for federated learning”);
compressing, by the online system, a model collection based on [information associated with one or more computational resources available at the computing device] . . . to generate a compressed model collection, the model collection comprising a plurality of machine learning models (Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where a model collection comprising a plurality of models is compressed, “Local models are compressed”, emphasis added to plurally identified models – which demonstrates a collection comprising a plurality of models is compressed, and where Fig. 1, reproduced above, depicts performance of “Model Compression” as occurring at the “Server” for a plurality of “Model Compression” operations acting on the plurality of “Local Model[s]”, which generates a compressed model collection; see also Pg. 2, Col. 2, Para. 6, “models are compressed from the global model using compression techniques such as pruning, quantization, and clustering”; see also Pg. 1-2, Col. 1-1, Para. 5-1, “Our framework differs from existing works in that our local models are compressed from the global model, in order to meet the resource constraints of different devices. In other words, local models are compressed with different compression techniques (e.g., pruning and quantization) to different degrees (e.g., different pruning ratios) . . . Due to the device heterogeneity, different IoT devices need to run different compressed versions of the global model. For example, an IoT hub can afford sophisticated models, whereas an embedded device can only run lightweight models”, where, as discussed above, “local models are compressed . . . to different degrees” based on “the resource constraints of different devices”);
transmitting, by the online system to the computing device, the compressed model collection (Pg. 2, Col. 2, Para. 6, “models are compressed from the global model using compression techniques such as pruning, quantization, and clustering. They receive more data during the model running on the local devices”, where the compressed model collection, “models are compressed from the global model”, are used for “model running on the local devices”; see also Pg. 1, Col. 2, Fig. 1, where Fig. 1, reproduced above, depicts transmission of the “Local Model[s]” after “Model Compression” from the “Server” to “Local Device 1” and “local Device 2”, as shown by the dotted arrows);
receiving, by the online system from the computing device, information associated with an update of at least one machine learning model in the model collection by the computing device (Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model”, where “gradients” of machine learning models in the model collection, “Local models”, are aggregated – which requires the “gradients” to be received, and where a person of ordinary skill in the art would understand model gradients to be information associated with an update to a machine learning model, see Pg. 2, Col. 1, Para. 4, “In our system architecture (Figure 1), the compressed models need to retrain themselves when new data are available and upload the gradients to the server for global model retraining”, where the online system receives the information, “upload the gradients to the server”, and where Fig. 1, reproduced above, depicts training the “Local Model[s]” using “Local data” at the “Local Device[s]” and the receiving of “Gradient[s]” by the “Server” from computing devices “Local Device 1” and “Local Device 2”; see also Pg. 2, Col. 2, Para. 6, “After receiving new data, the local model retrains itself using the new data and uploads the gradient to the server. On the server-side, the global model retrains itself by aggregating the gradients from local models”); and
updating, by the online system, the model collection based on the information received from the computing device (Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the model collection is updated, “retrain the global model, which in turn updates the local models”, based on the information, “Local models[’] gradients are aggregated to”, and where Fig. 1, reproduced above, depicts the retraining of the “Global Model” to update the “Local Model[s]” as occurring at the online system, the “Server”, based on the “Gradient[s]” received by the “Local Device 1” and “Local Device 2”; see also Pg. 2, Col. 2, Para. 6, “After the global model is retrained, the local model is updated by compressing the global model. The whole process repeats whenever new data is available to the local devices”).
Zhang does not explicitly disclose . . . from a computing device, a status report comprising . . . the status report . . . (where, as discussed above, Zhang teaches that the online system receives information associated with one or more computational resources available at the computing device in order to perform device-specific model compression, but Zhang does not specifically describe the source of the information as from the computing device itself or the form of the information as in a status report).
Alternatively, it could be argued that Zhang teaches . . . from a computing device, a status report comprising . . . the status report . . . (Pg. 1-2, Col. 1-1, Para. 5-1, “Our framework differs from existing works in that our local models are compressed from the global model, in order to meet the resource constraints of different devices. In other words, local models are compressed with different compression techniques (e.g., pruning and quantization) to different degrees (e.g., different pruning ratios) . . . Due to the device heterogeneity, different IoT devices need to run different compressed versions of the global model. For example, an IoT hub can afford sophisticated models, whereas an embedded device can only run lightweight models”, where it could reasonably be argued that information conveying “the resource constraints of different devices” is, regardless of form, a status report because it reports on the status of local devices to “run” various types of models, such as “sophisticated models” or “only . . . lightweight models”; and where the status report must be, indirectly or directly, from the computing device – such as transmission from the computing device, estimation or observation from the computing device, or stored in a database containing hardware specifics from the computing device).
However, given that this Office Action relies on the more constrained interpretation of Zhang, Chen teaches . . . [receiving, by an online system] from a computing device, a status report (Fig. 4, where a system, the “Centralized server/controller”, receives a status report, “S406, Measurement Report”, from a computing device, “Distributed client/Base station”; see also Pg. 12, Col. 8, Ln. 4-6, “the AI server may obtain the measurement report of the terminal device through a communication interface between network side devices”, where the system, “the AI server”, is online when operably connected “through a communications interface”, see also Fig. 4-7 for addition details on the online communications)
[comprising information associated with one or more computational resources available at the computing device] . . . (Pg. 18, Col. 20, Ln. 11-15, “The sent measurement report message carries a measurement value of the measurement quantity, and optionally carries an ID of the ML session to which the measurement belongs, and an AI computation processing capacity of the AI distributed client”; see also Pg. 11-12, Col. 6-7, Ln. 67-7, “The AI computation processing capacity of the distributed client may include, for example, a memory size of a central processing unit (CPU), a display memory size of a graphics processing unit (GPU), a CPU/GPU resource utilization rate in a current state, and the like in the AI distributed client. Then, the distributed client is instructed to report one or more supportable ML models”)
[and performing model training processing based on] the status report . . . (Abstract, “performing designated model training processing for network optimization based on whether to perform collaborative training with the client device and measurement data in the received measurement report messages”, where “performing designated model training processing” is based on the status report, “based on whether to perform collaborative training with the client device and measurement data in the received measurement report messages”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the receiving, by an online system, information associated with computational resources available at a computing device and the use of the information, by the online system, to perform model compression of Zhang with the receiving, by an online system from a computing device, a status report comprising information associated with computational resources available at the computing device, and the use of the information, by the online system, to perform model training processing of Chen in order to modify the compression of machine learning models based on the resource capacity of the receiving device (Zhang, Pg. 1-2, Col. 1-1, Para. 5-1, “Our framework differs from existing works in that our local models are compressed from the global model, in order to meet the resource constraints of different devices. In other words, local models are compressed with different compression techniques (e.g., pruning and quantization) to different degrees (e.g., different pruning ratios) . . . Due to the device heterogeneity, different IoT devices need to run different compressed versions of the global model. For example, an IoT hub can afford sophisticated models, whereas an embedded device can only run lightweight models”) such that the resource capacity is based on the current state of the computing device that is directly received from the computing device (Chen, Pg. 11-12, Col. 6-7, Ln. 15-7, “the network optimization method further includes . . . The AI computation processing capacity of the distributed client may include, for example, a memory size of a central processing unit (CPU), a display memory size of a graphics processing unit (GPU), a CPU/GPU resource utilization rate in a current state, and the like in the AI distributed client. Then, the distributed client is instructed to report one or more supportable ML models”, where “distribute client” provides its “AI computation processing capacity” in order to convey which “ML models” are “supportable”, given its “current state”), which will allow for narrowly tailored compression of the machine learning models, based on the current state of the device, in order to retain as much model accuracy as currently possible (Chen, Pg. 11-12, Col. 6-7, Ln. 15-7, “the network optimization method further includes . . . The AI computation processing capacity of the distributed client . . . in a current state, and the like in the AI distributed client. Then, the distributed client is instructed to report one or more supportable ML models”, where “the distributed client” “report[s]” on “one or more supportable ML models”, given its “current state”, which will prevent unnecessary reduces in model accuracy, Zhang, Pg. 3, Col. 1, Para. 6, “the accuracy of a compressed model is worse than its original model”, in instances where the “current state” of the device is computationally better than the assumed capacity of the device, see Zhang, Pg. 3, Col. 1, Para. 2, “Compared to the server, which is usually equipped with a powerful CPU, GPU, and large memory, IoT devices are limited in computational resources” and Zhang, Pg. 1-2, Col. 1-1, Para. 5-1, “Our framework differs from existing works in that our local models are compressed from the global model, in order to meet the resource constraints of different devices. In other words, local models are compressed with different compression techniques (e.g., pruning and quantization) to different degrees (e.g., different pruning ratios) . . . Due to the device heterogeneity, different IoT devices need to run different compressed versions of the global model. For example, an IoT hub can afford sophisticated models, whereas an embedded device can only run lightweight models”, where in the absence of current status reports from the device, the selection of “different compression techniques”, must be based on an assumed capacity that may be higher or lower than the current state).
Regarding Claim 2, Zhang in view of Chen teach the computer-implemented method of claim 1, wherein the online system is in communication with a group of computing devices that includes the computing device (Zhang, Pg. 1, Col. 2, Fig. 1, where, as indicated by the dotted arrows from the online system, the “Server”, to a group of computing devices, the computing device “Local Device 1” and also “Local Device 2”, the online system is in communication with the group of computing devices; see also Zhang, Pg. 2, col. 1, Para. 4, “gradient transmission from a local device to the server and local model
update from the global model”),
and the method further comprises: receiving, by the online system from the group of computing devices, information associated with a set of machine learning models, wherein each machine learning model in the set is included in the model collection or is generated by at least one computing device in the group (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model”, where the online system, the “Server”, receives information associated with the updates of the set machine learning model, “Local model . . . gradients”, which are “aggregated” from the collection of models, comprising the plurality of “Local Model[s]” each of which is included in the set and generated, in part, by the “Local Device[s]”, see also Zhang, Pg. 2, Col. 2, Para. 6, “After receiving new data, the local model retrains itself using the new data and uploads the gradient to the server. On the server-side, the global model retrains itself by aggregating the gradients from local models”);
generating a new machine learning model based on the set of machine learning models; and adding the new machine learning model to the model collection (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the online system, the “Server”, updates the model collection, “retrain the global model, which in turn updates the local models”, which generates a new model with updated parameters that is added to the collection as a replacement for the “global model”, and subsequently the “local models”, based on information received from the computing devices, which in turn is based on the set of machine learning models, “Local model . . . gradients” received from the “Local Device[s]”; see also Zhang, Pg. 2, Col. 2, Para. 6, “After the global model is retrained, the local model is updated by compressing the global model. The whole process repeats whenever new data is available to the local devices”).
Regarding Claim 6, Zhang in view of Chen teach the computer-implemented method of claim 1, wherein the online system is in communication with a group of computing devices that includes the computing device (Zhang, Pg. 1, Col. 2, Fig. 1, where, as indicated by the dotted arrows from the online system, the “Server”, to a group of computing devices, the computing device “Local Device 1” and also “Local Device 2”, the online system is in communication with the group of computing devices; see also Zhang, Pg. 2, col. 1, Para. 4, “gradient transmission from a local device to the server and local model
update from the global model”),
and the method further comprises: receiving, by the online system from the group of computing devices, information associated with one or more machine learning models in the model collection (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model”, where the online system, the “Server”, receives information, “Local model . . . gradients”, associated the one or more machine learning models, the “Local models” that are transmitted “from the global model”, , which are “aggregated” after being transmitted from the group of computing devices, “Local Device[s]”; see also Zhang, Pg. 2, Col. 2, Para. 6, “After receiving new data, the local model retrains itself using the new data and uploads the gradient to the server. On the server-side, the global model retrains itself by aggregating the gradients from local models”),
wherein the one or more machine learning models have been classified by the group computing devices as having worse performance than one or more other machine learning models in the model collection (Zhang, Pg. 1, Col. 2, Fig. 1 and Zhang, Pg. 2, Col. 2, Para. 4, “In our system architecture (Figure 1), the compressed models need to retrain themselves when new data are available and upload the gradients to the server for global model retraining”, where each “Local Device” in the group of computing devices “upload[s] the gradients to the server for global model retraining”, thereby demonstrating that the “Local Device” has classified one or more machine learning models, the “Global Model” and the previous version of the transmitted “Local Model”, as having worse performance on the “Local data” than one or more other machine learning models in the model collection, the “Local Model” with updated parameters corresponding with the “upload[ed]” “gradient”); and
removing at least one of the one or more machine learning models from the model collection (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the “system” uses the information transmitted from the group of devices, “gradients are aggregated”, to remove, by replacement through “retrain[ing]” to “update”, “the global model” and, subsequently, “the local models”).
Regarding Claim 7, Zhang in view of Chen teach the computer-implemented method of claim 6, wherein removing at least one of the one or more machine learning models from the model collection comprises (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the “system” uses the information transmitted from the group of devices, “gradients are aggregated”, to remove, by replacement through “retrain[ing]” to “update”, “the global model” and, subsequently, “the local models”):
selecting a machine learning model from the one or more machine learning models based on a number of computing devices in the group (Zhang, Pg. 2, Col. 2, Para. 6, “After receiving new data, the local model retrains itself using the new data and uploads the gradient to the server. On the server-side, the global model retrains itself by aggregating the gradients from local models. After the global model is retrained, the local model is updated by compressing the global model. The whole process repeats whenever new data is available to the local devices”, where the “the global model is retrained, the local model is updated”, which requires selection of the models for “retrain[ing]” or “update[ing]”, “whenever new data is available to the local devices” to generate “gradient[s]”, thus the selecting is based on whether the number of computing devices in the group with gradients is zero or nonzero)
that has classified the machine learning model as having worse performance than the one or more other machine learning models in the model collection (Zhang, Pg. 1, Col. 2, Fig. 1 and Zhang, Pg. 2, Col. 2, Para. 4, “In our system architecture (Figure 1), the compressed models need to retrain themselves when new data are available and upload the gradients to the server for global model retraining”, where each “Local Device” in the group of computing devices “upload[s] the gradients to the server for global model retraining”, thereby demonstrating that the “Local Device” has classified one or more machine learning models, the “Global Model” and the previous version of the transmitted “Local Model”, as having worse performance on the “Local data” than one or more other machine learning models in the model collection, the “Local Model” with updated parameters corresponding with the “upload[ed]” “gradient”);
and removing the machine learning model from the model collection (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the “system” uses the information transmitted from the group of devices, “gradients are aggregated”, to remove, by replacement through “retrain[ing]” to “update”, “the global model” and, subsequently, “the local models”).
Regarding Claim 10, Zhang in view of Chen teach the computer-implemented method of claim 1, wherein the plurality of machine learning models in the model collection is generated for performing a same machine learning task (Zhang, Pg. 1, Col. 1, Abstract, “we propose an FL framework targeting the heterogeneity of IoT devices. Specifically, local models are compressed from the global model, and the gradients of the compressed local models are used to update the global model”, where the “FL framework” uses a plurality of machine learning models in the model collection, the “global model” and the “local models”, which are generated for performing a same machine learning task, such as “binary classification” using “a 5-layer Multi-Layer Perception (MLP)”, see Zhang, Pg. 3, Col. 1, Para. 8, “We implement a binary classification model using our framework and measure its performance in terms of accuracy, time, and memory overhead. Specifically, we build a 5-layer Multi-Layer Perception (MLP), where each layer has 10 neurons (filters) and uses sigmoid as the activation. To train and evaluate the model, we simulate data samples with 5 features using Gaussian distributions”).
Regarding Claim 11, Zhang teaches one or more non-transitory computer-readable media storing instructions executable to perform operations, the operations comprising: . . . (Pg. 3, Col. 2, Para. 1, “All experiments are conducted in a Lenovo ThinkBook laptop (8 2.4GHz CPU cores and 16GB memory)”, where the “Lenovo ThinkBook laptop”, comprises one or more non-transitory computer-readable media, “16GB memory”, which stores instructions executable to perform operations, “C/C++” code that is executed to perform the “algorithms” of the “framework”, see Pg. 1, Col. 2, Para. 2, “we implement the whole system in C/C++ from scratch. Our framework includes a minimized but holistic model training pipeline such as forwarding and back prorogation of gradients. Our framework can facilitate the research on IoT-aware federated learning. Second, new gradient aggregation algorithms must exploit the gradients from the local compressed models to train the global model”, which the processors, “8 2.4GHz CPU cores”, execute to perform the operations of the “experiments”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Regarding Claim 12, the additional elements of the dependent claim are substantially the same as limitations of Claim 2, therefore it is rejected under the same rationale.
Regarding Claim 13, the additional elements of the dependent claim are substantially the same as limitations of Claim 6, therefore it is rejected under the same rationale.
Regarding Claim 15, the additional elements of the dependent claim are substantially the same as limitations of Claim 10, therefore it is rejected under the same rationale.
Regarding Claim 16, Zhang teaches an apparatus, comprising: a computer processor for executing computer program instructions; and a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations comprising: . . . (Pg. 3, Col. 2, Para. 1, “All experiments are conducted in a Lenovo ThinkBook laptop (8 2.4GHz CPU cores and 16GB memory)”, where the “Lenovo ThinkBook laptop” apparatus comprises one or more non-transitory computer-readable media, “16GB memory”, which stores computer program instructions executable to perform operations, “C/C++” code that is executed to perform the “algorithms” of the “framework”, see Pg. 1, Col. 2, Para. 2, “we implement the whole system in C/C++ from scratch. Our framework includes a minimized but holistic model training pipeline such as forwarding and back prorogation of gradients. Our framework can facilitate the research on IoT-aware federated learning. Second, new gradient aggregation algorithms must exploit the gradients from the local compressed models to train the global model”, which the computer processors, “8 2.4GHz CPU cores”, execute to perform the operations of the “experiments”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Regarding Claim 17, the additional elements of the dependent claim are substantially the same as limitations of Claim 2, therefore it is rejected under the same rationale.
Regarding Claim 18, the additional elements of the dependent claim are substantially the same as limitations of Claim 6, therefore it is rejected under the same rationale.
Regarding Claim 20, the additional elements of the dependent claim are substantially the same as limitations of Claim 10, therefore it is rejected under the same rationale.
Claims 3-4, 8, 14, 19, and 21-25 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Chen and Kelly et al. (hereinafter Kelly) (Pat. Pub. No. US 2022/0114491 A1).
Regarding Claim 3, Zhang in view of Chen teach the computer-implemented method of claim 2, wherein a machine learning model in the set is selected . . . [from] the model collection (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model”, where the online system, the “Server”, transmits the “compressed” models to the compute device, “Local Device”, as “Local models”, which is shown by the dotted arrow from “Global model” to “Local model”, and where the transmitted models must be selected for the “Local Device[s]” prior to transmission).
Zhang in view of Chen do not explicitly disclose . . . by a computing device in the group based on a similarity score determined by the computing device in the group, the similarity score indicating a degree of similarity between the machine learning model in the set and one or more machine learning models in . . . (where a similarity score is not specifically discussed in regard to model selection and the selection is not specifically discussed as being by a computing device in the group).
However, Kelly teaches . . . [wherein a machine learning model in the set is selected] by a computing device in the group based on a similarity score determined by the computing device in the group (Abstract, “Discrete model classes of models are trained on non-anonymous datasets at a centralized server and served to anonymous clients. Clients validate each model against its own localized datasets and retain the most accurate model”, where a computing device, “Client”, in a group of computing devices, “Clients”, selects a machine learning model, “retain the most accurate model”, from a set of machine learning models, “classes of models”, based on an evaluation of accuracy, “the most accurate model”; see also Para. [0109], “after all of the plurality of machine learning models are run by the anonymous client, the machine learning model having the highest accuracy among the plurality of machine learning models based on the local dataset is selected, e.g., a machine learning model that has between 80-95% accuracy”; see also Para. [0031], “a model trained based on local data might not include training data that may be useful for future predictions since the local data might not include data from other clients that have similar conditions to provide a more robust trained model. Machine learning models that are run locally may also provide more accurate predictions “out-of-the-box”, e.g., on initial installation and start-up, than a generic global machine learned model from a cloud-based processing system, since the local machine learning models include client-targeted and specific models that were trained on datasets that may be very similar to another client”, where “accuracy” is a similarity score because, in the context of “run[ning models] locally”, it indicates similarity between training data and local data),
the similarity score indicating a degree of similarity between the machine learning model in the set and one or more machine learning models in [the model collection] (Para. [0053] – [0055], “The parameters of the local machine learning model may then be transmitted or uploaded to the cloud-based computing system periodically . . . The cloud-based computing system 100 . . . aggregate[s] all of the respective received parameters for the machine learning models of the plurality of machine learning models and update the parameters of the machine learning model for the respective model class . . . The cloud-based computing system 100 may then be designed, programmed, or otherwise configured to transmit the updated machine learning models to at least one other client to make predictions”, where the “client” iteratively receives “respective class models” from the “system 100”, such that the “accuracy”-based similarity score, see Para. [0031], “a model trained based on local data might not include training data that may be useful for future predictions since the local data might not include data from other clients that have similar conditions to provide a more robust trained model. Machine learning models that are run locally may also provide more accurate predictions “out-of-the-box”, e.g., on initial installation and start-up, than a generic global machine learned model from a cloud-based processing system, since the local machine learning models include client-targeted and specific models that were trained on datasets that may be very similar to another client”, will indicate a degree of similarity between the machine learning model in the set, “respective class models”, and one or more machine learning models in the collection, “the local machine learning model”; see also Fig. 1, where similarity scores for the set of machine learning models “130”, provided as indicated by the dotted line, indicate the similarity between the model parameters and the “Params” of the locally trained model by “Ex. Cl. 140”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the selection of a machine learning model in the set machine learning models of Zhang in view of Chen with the accuracy-based selection of a machine learning model in the set machine learning models by a computing device in a group, wherein the accuracy score is a similarity score that is used for selection and indicates a degree of similarity between the machine learning model and one or more machine learning models in the model collection of Kelly in order to select a machine learning model with the highest accuracy when applied to the local training data, such as the model closest to a previously used local model, which will increase model training efficiency while preserving privacy (compare Kelly, Para. [0002], “A generic “one-size-fits-all” approach can provide a generic model to all clients, which will be gradually improved with localized, user-specific data over time through transfer learning. Alternatively, client-targeted models can be deployed that were trained on datasets very similar to those generated by the client. These models will provide more accurate predictions “out-of-the-box”, e.g., at installation and initial start-up, but at the expense of potential loss of privacy from the client to the service provider” with Kelly, Para. [0109], “after all of the plurality of machine learning models are run by the anonymous client, the machine learning model having the highest accuracy
among the plurality of machine learning models based on the local dataset is selected, e.g., a machine learning model that has between 80-95% accuracy”, where allowing the computing device, “client”, to “run” and evaluate the “accuracy” of a set of models allows for the benefits of “client-targeted models”, while preserving “privacy”).
Regarding Claim 4, Zhang in view of Chen and Kelly teach the computer-implemented method of claim 2, wherein a machine learning model in the set is selected by a computing device in the group based on an evaluation of an accuracy of the machine learning model in the set (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model”, where the online system, the “Server”, transmits the “compressed” models to the compute device, “Local Device”, as “Local models”, which is shown by the dotted arrow from “Global model” to “Local model”, and where the transmitted models must be selected for the “Local Device[s]” prior to transmission, which, in view of Kelly, is by a computing device in the group and based on an evaluation of accuracy of the models, see Kelly, Abstract, “Discrete model classes of models are trained on non-anonymous datasets at a centralized server and served to anonymous clients. Clients validate each model against its own localized datasets and retain the most accurate model”, where a computing device, “Client”, in a group of computing devices, “Clients”, selects a machine learning model, “retain the most accurate model”, from a set of machine learning models, “classes of models”, based on an evaluation of accuracy, “the most accurate model”; see also Kelly, Para. [0109], “after all of the plurality of machine learning models are run by the anonymous client, the machine learning model having the highest accuracy among the plurality of machine learning models based on the local dataset is selected, e.g., a machine learning model that has between 80-95% accuracy”).
The reasons for obviousness were discussed in regard to the rejection of claim 3 above and remains applicable here.
Regarding Claim 8, Zhang in view of Chen and Kelly teach the computer-implemented method of claim 1, further comprising: receiving, from the computing device, a request for associating with the online system, the request comprising a machine learning model (Zhang, Pg. 2, Col. 2, Para. 6, “After receiving new data, the local model retrains itself using the new data and uploads the gradient to the server. On the server-side, the global model retrains itself by aggregating the gradients from local models. After the global model is retrained, the local model is updated by compressing the global model. The whole process repeats whenever new data is available to the local devices”, where the “process repeats whenever new data is available to the local devices”, thus the “upload[ing of] the gradient[s] to the server” is a request to associate with the online system, “the server”, that is received from the computing devices, “local devices”, which, in view of Chen, comprises “establishing a communication inference” between the online system and the computing device, see Chen, Fig. 6 and Chen, Pg. 27, Col. 22, Ln. 41-51, “FIG. 6 is a schematic flowchart of establishing a communication interface between a server and a network side device . . . At S602, the base station sends a communication interface setup request according to the configured address of the AI centralized server”, and where, in view of Kelly, the “request” comprises a machine learning model, “anonymized parameters”, see Kelly, Fig. 1 and Kelly, Abstract, “Clients improve their model locally through transfer learning on new datasets, and share the updated, anonymized parameters with a centralized computer”); and
generating the model collection by adding the machine learning model in the request to a previous model collection (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the “system” uses the information transmitted from the group of devices, “gradients are aggregated”, to generate the model collection, by replacement through “retrain[ing]” to “update”, “the global model” and, subsequently, “the local models”, which, as discussed above, is done by adding, in part, the model in the request to the previous model collection, see Kelly, Abstract, “Clients improve their model locally through transfer learning on new datasets, and share the updated, anonymized parameters with a centralized computer. The centralized server aggregates and updates model parameters for each respective discrete model class” and Chen, Pg. 27, Col. 22, Ln. 41-51, “FIG. 6 is a schematic flowchart of establishing a communication interface between a server and a network side device . . . At S602, the base station sends a communication interface setup request according to the configured address of the AI centralized server”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the receiving, from the computing device, a request for associating with the online system, in the form of a gradient of Zhang in view of Chen with the receiving, as part of a formalized establishment of a communication interface between an online system and a computing device, a request from the computing device to associate with the online system, in further view of Chen in order to establish a communication interface that meets the capacity and functionality requirements of both the online system and local device (Chen, Pg. 19, Col. 22, Ln. 41-64, “a schematic flowchart of establishing a communication interface between a server and a network side device . . . the communication interface setup request may include: measurement supported by the network side device (e.g., a base station); a measurement reporting mode supported by the network side device; an RAN optimization action supported by the network side device; an AI computing capacity supported by the AI distributed client; and an ML model supported by the AI distributed client”).
Additionally, before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the receiving, from the computing device, a request comprising a gradient, which is used to generate the model collection of Zhang in view of Chen with the transmission of a request comprising a machine learning model, which is used to generate the model collection by adding the machine learning model in the request to the collection of Kelly in order to improve the global model, using a method with increased efficiency and increased privacy protection for the local device (compare Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model”, where malicious actors can utilize “local model . . . gradients” to compromise the privacy of the local data and the information in each “gradient” in localized to a specific subset of model updates, with Kelly, Fig. 1 and Kelly, Abstract, “Clients improve their model locally through transfer learning on new datasets, and share the updated, anonymized parameters with a centralized computer. The centralized server aggregates and updates model parameters for each respective discrete model class”, where the privacy of the client data is further protected through “anonymized parameters”, which do not directly translate to a single iteration of client data, and the “updated” “parameters” can efficiency capture relevant information from multiple iterations of gradient-based learning).
Regarding Claim 14, the additional elements of the dependent claim are substantially the same as limitations of Claim 8, therefore it is rejected under the same rationale.
Regarding Claim 19, the additional elements of the dependent claim are substantially the same as limitations of Claim 8, therefore it is rejected under the same rationale.
Regarding Claim 21, Zhang in view of Chen and Kelly teach an apparatus, comprising: a computer processor for executing computer program instructions; and a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations comprising (Zhang, Pg. 3, Col. 2, Para. 1, “All experiments are conducted in a Lenovo ThinkBook laptop (8 2.4GHz CPU cores and 16GB memory)”, where the “Lenovo ThinkBook laptop” apparatus comprises one or more non-transitory computer-readable media, “16GB memory”, which stores computer program instructions executable to perform operations, “C/C++” code that is executed to perform the “algorithms” of the “framework”, see Zhang, Pg. 1, Col. 2, Para. 2, “we implement the whole system in C/C++ from scratch. Our framework includes a minimized but holistic model training pipeline such as forwarding and back prorogation of gradients. Our framework can facilitate the research on IoT-aware federated learning. Second, new gradient aggregation algorithms must exploit the gradients from the local compressed models to train the global model”, which the computer processors, “8 2.4GHz CPU cores”, execute to perform the operations of the “experiments”):
PNG
media_image1.png
206
404
media_image1.png
Greyscale
Zhang, Figure 1
generating, by a computing device, a first status report comprising information associated with one or more computational resources available at the computing device, transmitting, by the computing device to an online system, the first status report (Zhang, Pg. 1-2, Col. 1-1, Para. 5-1, “Our framework differs from existing works in that our local models are compressed from the global model, in order to meet the resource constraints of different devices. In other words, local models are compressed with different compression techniques (e.g., pruning and quantization) to different degrees (e.g., different pruning ratios) . . . Due to the device heterogeneity, different IoT devices need to run different compressed versions of the global model. For example, an IoT hub can afford sophisticated models, whereas an embedded device can only run lightweight models”, where information associated with one or more computational resources, “the resource constraints”, at each computing device, “of different devices”, is used to perform model “compress[ion] with different compression techniques (e.g., pruning and quantization) to different degrees”; see also Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning”, where Fig. 1, reproduced above, depicts performance of “Model Compression” as occurring at the “Server”, thus, since the information associated with one or more computational resources is used to perform model compression, the server must receive this information in order to perform compression, “local models are compressed”, based on the “need[s]” of the “different IoT devices”; see also Zhang, Pg. 1, Col. 1, Abstract, “Internet of Things (IoT) devices are inherently diverse regarding computation speed and onboard memory. In this paper, we propose an FL framework targeting the heterogeneity of IoT devices . . . [to] facilitate the design of IoT-aware FL”, where the “Server” is an online system when operably connect to “Internet” as part of an “IoT-aware” “system architecture for federated learning”; and which, in view of Chen is generated by the computing device generates a status report on the resource constraints and then transmits it to the online system, see Chen, Fig. 4, where a system, the “Centralized server/controller”, receives a generated status report, “S406, Measurement Report”, from a computing device, “Distributed client/Base station”; see also Chen, Pg. 12, Col. 8, Ln. 4-6, “the AI server may obtain the measurement report of the terminal device through a communication interface between network side devices”, where the system, “the AI server”, is online when operably connected “through a communications interface”, see also Chen, Fig. 7),
receiving, by the computing device from the online system, a compressed model collection (Zhang, Pg. 2, Col. 2, Para. 6, “models are compressed from the global model using compression techniques such as pruning, quantization, and clustering. They receive more data during the model running on the local devices”, where the compressed model collection, “models are compressed from the global model”, are used for “model running on the local devices”; see also Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where Fig. 1, reproduced above, depicts transmission of the “Local Model[s]” after “Model Compression” from the “Server” to “Local Device 1” and “local Device 2”, as shown by the dotted arrows; and where, in view of Kelly, each device receives the collection of models, “Discrete model classes of models”, see Kelly, Abstract, “Discrete model classes of models are trained on non-anonymous datasets at a centralized server and served to anonymous clients. Clients validate each model against its own localized datasets and retain the most accurate model”),
the compressed model collection generated by compressing a model collection, which comprises a plurality of machine learning models, based on the first status report (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where a model collection comprising a plurality of models is compressed, “Local models are compressed”, emphasis added to plurally identified models – which demonstrates a collection comprising a plurality of models is compressed, and where Fig. 1, reproduced above, depicts performance of “Model Compression” as occurring at the “Server” for a plurality of “Model Compression” operations acting on the plurality of “Local Model[s]”, which generates a compressed model collection; see also Zhang, Pg. 2, Col. 2, Para. 6, “models are compressed from the global model using compression techniques such as pruning, quantization, and clustering”; see also Zhang, Pg. 1-2, Col. 1-1, Para. 5-1, “Our framework differs from existing works in that our local models are compressed from the global model, in order to meet the resource constraints of different devices. In other words, local models are compressed with different compression techniques (e.g., pruning and quantization) to different degrees (e.g., different pruning ratios) . . . Due to the device heterogeneity, different IoT devices need to run different compressed versions of the global model. For example, an IoT hub can afford sophisticated models, whereas an embedded device can only run lightweight models”, where, as discussed above, “local models are compressed . . . to different degrees” based on “the resource constraints of different devices”, which, in view of Chen, is contained within the status report, see Chen, Fig. 4, where a system, the “Centralized server/controller”, receives a status report, “S406, Measurement Report”, from a computing device, “Distributed client/Base station”; see also Chen, Pg. 12, Col. 8, Ln. 4-6, “the AI server may obtain the measurement report of the terminal device through a communication interface between network side devices”, where the system, “the AI server”, is online when operably connected “through a communications interface”),
updating, by the computing device, a machine learning model in the model collection (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the model collection is updated, “retrain the global model, which in turn updates the local models”, based on the information, “Local models[’] gradients are aggregated to”, and where Fig. 1, reproduced above, depicts the “Local Device[s]” updating the “Local Model[s]” in the model collection using “Local Data”; see also Zhang, Pg. 2, Col. 2, Para. 2, “At each round, the global model broadcasts the current model to the local clients on devices for the local training. After this step, local models are fine-tuned using local data and aggregated to update the global model”),
generating, by the computing device, a second status report comprising information associated with the machine learning model in the model collection, and transmitting, by the computing device to an online system, the second status report (Zhang, Pg. 1, Col. 2, Fig. 1, where a “Server”, which is an online system when operably connect to “Internet” as part of an “IoT-aware FL” system, see Zhang, Pg. 1, Col. 1, Abstract, “Internet of Things (IoT) devices are inherently diverse regarding computation speed and onboard memory. In this paper, we propose an FL framework targeting the heterogeneity of IoT devices . . . [to] facilitate the design of IoT-aware FL”, performs “Model Compression”; see also Zhang, Pg. 1-2, Col. 1-1, Para. 5-1, “Our framework differs from existing works in that our local models are compressed from the global model, in order to meet the resource constraints of different devices. In other words, local models are compressed with different compression techniques (e.g., pruning and quantization) to different degrees (e.g., different pruning ratios) . . . Due to the device heterogeneity, different IoT devices need to run different compressed versions of the global model. For example, an IoT hub can afford sophisticated models, whereas an embedded device can only run lightweight models”, where the online system, the “Server”, must receive information associated with one or more computational resources available, “the resource constraints”, at each computing device, “of different devices”, in order to perform model “compress[ion] with different compression techniques (e.g., pruning and quantization) to different degrees” based in the “need[s]” of the “different IoT devices”, which must be reported to the “Server” in as data to convey the status of the “different devices”, which, in view of Chen is in a status report generated by the computing device and then transmitted to the online system by the computing device, see Chen, Fig. 4, where a system, the “Centralized server/controller”, receives a generated status report, “S406, Measurement Report”, from a computing device, “Distributed client/Base station”; see also Chen, Pg. 12, Col. 8, Ln. 4-6, “the AI server may obtain the measurement report of the terminal device through a communication interface between network side devices”, where the system, “the AI server”, is online when operably connected “through a communications interface”, see also Chen, Fig. 7; see also Zhang, Pg. 2, col. 2, Para. 6, “After the global model is retrained, the local model is updated by compressing the global model. The whole process repeats whenever new data is available to the local devices”, where, during “repeat[ed]” “process[es]” the status report will be a second status report and will be associated with associated with the machine learning model in the model collection because it will provide information on the device used to generate the model).
The reasons for obviousness were discussed in regard to the rejection of claim 1, for the combination with Chen, and in regard the rejection of claim 3, for the combination with Kelly, and remain applicable here.
Regarding Claim 22, Zhang in view of Chen and Kelly teach the apparatus of claim 21, wherein updating the machine learning model in the model collection comprises: training the machine learning model by using data available at the computing device (Zhang, Pg. 1, Col. 2, Fig. 1, where the “Local Device[s]” update the “Local Model[s]” in the model collection using “Local Data”; see also Zhang, Pg. 2, Col. 2, Para. 2, “At each round, the global model broadcasts the current model to the local clients on devices for the local training. After this step, local models are fine-tuned using local data and aggregated to update the global model”, where the machine learning models, “local models”, are trained using data available at the computing device, “fine-tuned using local data”).
Regarding Claim 23, Zhang in view of Chen and Kelly teach the apparatus of claim 21, wherein the operations further comprise: evaluating, by the computing device, performances of the plurality of machine learning models (Kelly, Abstract, “Discrete model classes of models are trained on non-anonymous datasets at a centralized server and served to anonymous clients. Clients validate each model against its own localized datasets and retain the most accurate model”);
identifying, by the computing device, one or more machine learning models from the plurality of machine learning models based on the performances (Kelly, Para. [0109], “after all of the plurality of machine learning models are run by the anonymous client, the machine learning model having the highest accuracy among the plurality of machine learning models based on the local dataset is selected, e.g., a machine learning model that has between 80-95% accuracy”); and
including, by the computing device, information associated with the one or more machine learning models in the second status report (Zhang, Pg. 1, Col. 2, Fig. 1, where a “Server”, which is an online system when operably connect to “Internet” as part of an “IoT-aware FL” system, see Zhang, Pg. 1, Col. 1, Abstract, “Internet of Things (IoT) devices are inherently diverse regarding computation speed and onboard memory. In this paper, we propose an FL framework targeting the heterogeneity of IoT devices . . . [to] facilitate the design of IoT-aware FL”, performs “Model Compression”; see also Zhang, Pg. 1-2, Col. 1-1, Para. 5-1, “Our framework differs from existing works in that our local models are compressed from the global model, in order to meet the resource constraints of different devices. In other words, local models are compressed with different compression techniques (e.g., pruning and quantization) to different degrees (e.g., different pruning ratios) . . . Due to the device heterogeneity, different IoT devices need to run different compressed versions of the global model. For example, an IoT hub can afford sophisticated models, whereas an embedded device can only run lightweight models”, where the online system, the “Server”, must receive information associated with one or more computational resources available, “the resource constraints”, at each computing device, “of different devices”, in order to perform model “compress[ion] with different compression techniques (e.g., pruning and quantization) to different degrees” based in the “need[s]” of the “different IoT devices”, which must be reported to the “Server” as data to convey the status of the “different devices”, which is within the broadest reasonable interpretation of a status report, which, in view of Chen is generated by the computing device and then transmitted to the online system by the computing device, see Chen, Fig. 4, where a system, the “Centralized server/controller”, receives a generated status report, “S406, Measurement Report”, from a computing device, “Distributed client/Base station”; see also Chen, Pg. 12, Col. 8, Ln. 4-6, “the AI server may obtain the measurement report of the terminal device through a communication interface between network side devices”, where the system, “the AI server”, is online when operably connected “through a communications interface”, see also Chen, Fig. 7; see also Zhang, Pg. 2, col. 2, Para. 6, “After the global model is retrained, the local model is updated by compressing the global model. The whole process repeats whenever new data is available to the local devices”, where, during “repeat[ed]” “process[es]” the status report will be a second status report and will be associated with associated with the machine learning model in the model collection because it will provide information on the device used to generate the model).
The reasons for obviousness were discussed in regard to the rejection of claim 3 above and remains applicable here.
Regarding Claim 24, Zhang in view of Chen and Kelly teach the apparatus of claim 23, wherein evaluating performances of the plurality of machine learning models comprises: for each respective machine learning model, determining a similarity score of the respective machine learning model (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model”, where the online system, the “Server”, transmits the “compressed” models to the compute device, “Local Device”, as “Local models”, which is shown by the dotted arrow from “Global model” to “Local model”, and where the transmitted models must be selected for the “Local Device[s]” prior to transmission, which, in view of Kelly, is by a computing device in the group and based on an evaluation of accuracy of the models, see Kelly, Abstract, “Discrete model classes of models are trained on non-anonymous datasets at a centralized server and served to anonymous clients. Clients validate each model against its own localized datasets and retain the most accurate model”, where a computing device, “Client”, in a group of computing devices, “Clients”, selects a machine learning model, “retain the most accurate model”, from a set of machine learning models, “classes of models”, based on an evaluation of accuracy, “the most accurate model”; see also Kelly, Para. [0109], “after all of the plurality of machine learning models are run by the anonymous client, the machine learning model having the highest accuracy among the plurality of machine learning models based on the local dataset is selected, e.g., a machine learning model that has between 80-95% accuracy”; see also Kelly, Para. [0109], “after all of the plurality of machine learning models are run by the anonymous client, the machine learning model having the highest accuracy among the plurality of machine learning
models based on the local dataset is selected, e.g., a machine learning model that has between 80-95% accuracy”; see also Kelly, Para. [0031], “a model trained based on local data might not include training data that may be useful for future predictions since the local data might not include data from other clients that have similar conditions to provide a more robust trained model. Machine learning models that are run locally may also provide more accurate predictions “out-of-the-box”, e.g., on initial installation and start-up, than a generic global machine learned model from a cloud-based processing system, since the local machine learning models include client-targeted and specific models that were trained on datasets that may be very similar to another client”, where “accuracy” is a similarity score because, in the context of “run[ning models] locally”, it indicates similarity between training data and local data),
the similarity score indicating a degree of similarity between the respective machine learning model and one or more other machine learning models in the model collection (Kelly, Para. [0053] – [0055], “The parameters of the local machine learning model may then be transmitted or uploaded to the cloud-based computing system periodically . . . The cloud-based computing system 100 . . . aggregate[s] all of the respective received parameters for the machine learning models of the plurality of machine learning models and update the parameters of the machine learning model for the respective model class . . . The cloud-based computing system 100 may then be designed, programmed, or otherwise configured to transmit the updated machine learning models to at least one other client to make predictions”, where the “client” iteratively receives “respective class models” from the “system 100”, such that the “accuracy”-based similarity score, see Kelly, Para. [0031], “a model trained based on local data might not include training data that may be useful for future predictions since the local data might not include data from other clients that have similar conditions to provide a more robust trained model. Machine learning models that are run locally may also provide more accurate predictions “out-of-the-box”, e.g., on initial installation and start-up, than a generic global machine learned model from a cloud-based processing system, since the local machine learning models include client-targeted and specific models that were trained on datasets that may be very similar to another client”, will indicate a degree of similarity between the machine learning model in the set, “respective class models”, and one or more machine learning models in the collection, “the local machine learning model”; see also Kelly, Fig. 1, where similarity scores for the set of machine learning models “130”, provided as indicated by the dotted line, indicate the similarity between the model parameters and the “Params” of the locally trained model by “Ex. Cl. 140”).
The reasons for obviousness were discussed in regard to the rejection of claim 3 above and remains applicable here.
Regarding Claim 25, Zhang in view of Chen and Kelly teach the apparatus of claim 23, wherein evaluating performances of the plurality of machine learning models comprises: for each respective machine learning model, determining an accuracy of the respective machine learning model (Kelly, Abstract, “Discrete model classes of models are trained on non-anonymous datasets at a centralized server and served to anonymous clients. Clients validate each model against its own localized datasets and retain the most accurate model”; see also Kelly, Para. [0109], “after all of the plurality of machine learning models are run by the anonymous client, the machine learning model having the highest accuracy among the plurality of machine learning models based on the local dataset is selected, e.g., a machine learning model that has between 80-95% accuracy”).
The reasons for obviousness were discussed in regard to the rejection of claim 3 above and remains applicable here.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Chen and Martin et al. (hereinafter Martin) (“Scalable XML Collaborative Editing with Undo”).
Regarding Claim 5, Zhang in view of Chen teach the computer-implemented method of claim 2, wherein adding the new machine learning model to the model collection comprises: identifying a machine learning model in the model collection . . . associated with the machine learning model; and replacing the machine learning model with the new machine learning model (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the online system, the “Server”, updates the model collection, “retrain the global model, which in turn updates the local models”, which generates a new model with updated parameters that is added to the collection as a replacement for the “global model”, and subsequently the “local models”, based on information received from the computing devices, which in turn is based on the set of machine learning models, “Local model . . . gradients” received from the “Local Device[s]”; see also Zhang, Pg. 2, Col. 2, Para. 6, “After the global model is retrained, the local model is updated by compressing the global model. The whole process repeats whenever new data is available to the local devices”, where “retrain[ing]” “the global model” and “update[ing]” “local model[s]” requires identification the models using identifiers associated with the models, particularly in the context of “algorithm[s]” that need to differentiate between “models that are compressed using pruning and quantization at the same time”, see Zhang, Pg. 5, Col. 1, Para. 2, “we plan to leverage our framework to design algorithms that can aggregate gradients of compressed models to train the global model . . . one can design a gradient aggregation algorithm for different pruning ratios, while another one can design an algorithm for models that are compressed using pruning and quantization at the same time”).
Zhang in view of Chen do not explicitly disclose . . . based on a timestamp. . . (where identification of models is not specifically discussed as being based on a timestamp).
However, Martin teaches . . . [data identifying] based on a timestamp . . . (Pg. 507, Abstract, “In this paper, we present a CRDT to edit XML data. Compared to existing approaches for XML collaborative editing, our approach is more scalable and handles all the XML editing aspects : elements, contents, attributes and undo”; see also Pg. 509, Para. 5, “To allow Add and Del operations to commute, we use a unique timestamp identifier”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the identifying of a machine learning model in the model collection of Zhang in view of Chen with the identifying of data based on a timestamp of Martin in order to utilize a scalable model updating approach (Martin, Pg. 507, Abstract, “In this paper, we present a CRDT to edit XML data. Compared to existing approaches for XML collaborative editing, our approach is more scalable and handles all the XML editing aspects : elements, contents, attributes and undo”), which allows for temporal differentiation of iteratively updated models (Zhang, Pg. 2, Col. 2, Para. 6, “After the global model is retrained, the local model is updated by compressing the global model. The whole process repeats whenever new data is available to the local devices”).
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Chen, Kelly, and Martin.
Regarding Claim 9, Zhang in view of Chen, Kelly, and Martin teach the computer-implemented method of claim 8, wherein the previous model collection comprises a plurality of previous machine learning models, and generating the model collection comprises (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the “system” uses the information transmitted from the group of devices, “gradients are aggregated”, to generate the model collection, by replacement of the previous model collection through “retrain[ing]” to “update”, “the global model” and, subsequently, “the local models”,):
identifying a previous machine learning model in the previous model collection based on timestamps associated with the plurality of previous machine learning models; and replacing the previous machine learning model with the machine learning model in the request (Zhang, Pg. 1, Col. 2, Fig. 1, “Our system architecture for federated learning. Local models are compressed from the global model, and their gradients are aggregated to retrain the global model, which in turn updates the local models”, where the online system, the “Server”, updates the model collection, “retrain the global model, which in turn updates the local models”, which generates a new model with updated parameters that is added to the collection as a replacement for the “global model”, and subsequently the “local models”, based on information received from the computing devices, which in turn is based on the set of machine learning models, “Local model . . . gradients” received from the “Local Device[s]”; see also Zhang, Pg. 2, Col. 2, Para. 6, “After the global model is retrained, the local model is updated by compressing the global model. The whole process repeats whenever new data is available to the local devices”, where “retrain[ing]” “the global model” and “update[ing]” “local model[s]” requires identification the models using identifiers associated with the models, particularly in the context of “algorithm[s]” that need to differentiate between “models that are compressed using pruning and quantization at the same time”, see Zhang, Pg. 5, Col. 1, Para. 2, “we plan to leverage our framework to design algorithms that can aggregate gradients of compressed models to train the global model . . . one can design a gradient aggregation algorithm for different pruning ratios, while another one can design an algorithm for models that are compressed using pruning and quantization at the same time”, which, in view of Martin, is based on a timestamp identifier, see Martin, Pg. 507, Abstract, “In this paper, we present a CRDT to edit XML data. Compared to existing approaches for XML collaborative editing, our approach is more scalable and handles all the XML editing aspects : elements, contents, attributes and undo”; see also Martin, Pg. 509, Para. 5, “To allow Add and Del operations to commute, we use a unique timestamp identifier”).
The reasons for obviousness were discussed in regard to the rejection of claim 5 above and remains applicable here.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Bhuyan et al. (“Multi-Model Federated Learning”) discloses a federated learning system where a system maintains a multi-model machine learning collection (see Pg. 1, Col. 1, Abstract).
Meyer et al. (Pat. No. US 12,265,888 B1) disclosures an alternative mechanism for device-provided status reports in a federated learning framework (see Fig. 8).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW BRYCE GOLAN whose telephone number is (571)272-5159. The examiner can normally be reached Monday through Friday, 8:00 AM to 5:00 PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW BRYCE GOLAN/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123