DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-11 are presented for examination.
Claim Objections
Claims 2 are objected to because of the following informalities,
Claim 2 [line 5]: “used for generating the plurality local models, respectively” should be “used for generating the plurality of local models, respectively.”
Appropriate corrections are required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 3-6 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
With respect to claims 3-6, it is unclear what the limitation “the display unit” [line 2] refers to. Claims 3-6 each depended on claim 1. However, claim 1 only recited “display”, and claim 2, which also depended on claim 1, only recited “the display” as well. Either claims 3-6 or claims 1 and 2 should be fixed to make the limitation consistent. For the purposes of examination, Examiner will interpret claims 3-6 as “the display.”
With respect to claim 3,
The limitation “own classification item classifying own learning data” [line 3] is confusing. The claim does not provide sufficient clarity as to the meaning of “own” in the recited “own classification item” and “own learning data”, whether “own” refers to the claimed node, the processing circuitry, or another entity. Further, it is unclear whether the “own classification item” itself performs a classification of the “own learning data” or merely indicates a classification of the “own learning data.” For the purposes of examination, Examiner would interpret the limitation as “a classification item classifying a learning data used for …”
The limitation “own classification item classifying own learning data” [line 3] is also confusing in the way that “classification item” is “classifying” something. Claim 2 recited “a plurality of classification items indicating classifications of the learning data.” This creates uncertainty about what “classification item” actually does. The limitation “classification item classifying learning data” means the item appears to perform the classification. Meanwhile, the limitation “classification item indicating a classification of learning data” means the item appears to represent/display information about a classification. The dependent claims 2 and 3 depended on claim 1, but they actually recited in different meanings and caused confusion.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5, 7-11 are rejected under 35 U.S.C. 103 as being unpatentable over Takasaki et al (US 20220343219 A1) hereafter Takasaki, and further in view of Zhang et al (US 20240311645 A1) hereafter Zhang.
With respect to claim 1, teaches a node comprising: processing circuitry (cloud computing environment is a service that includes a network of interconnected nodes, wherein one or more cloud computing nodes with which local computing devices are used by cloud consumers. Nodes may communicate with each other [par. 0050, 0072, 0073]) configured to:
display on a display, local model identification information identifying another local model generated by another node (local model identification information may be known as client/user IDs. Figure 1 shows multiple groups of local devices, wherein each local device is tagged with a number, for example, Local Device 1-5. The method includes receiving, by each local device, local models from other local devices, such that the each local device has all local models in one group [par. 0004-0008 and FIG. 1]) and
receive selection of said another local model (a local device is selected as a leader of a group. This local device is responsible for collecting validation results from other local devices. Each local device receives local model from other local devices [par. 0004-0008, 0028-0034]); and
transmit selection information indicating the selection of said another local model to another device (each local device receives local model from other local devices. Randomly selecting one or more local models whose accuracies do not exceed a predetermined threshold and sending to the server weight parameters of the selected local models [par. 0004-0008, 0028-0034 and FIG. 2A]).
However, Takasaki does not explicitly disclose a classification item that classifies learning data used to generate said another local model.
In the same field of endeavor, Zhang teaches a classification item that classifies learning data used to generate said another local model (federated learning is a machine learning (ML) technique that enables clients to participate in learning a model related to a task. The local model receives data sample from the local dataset and processes data sample through a sequence of neural network layers to generate a prediction result. Normalization of the feature vector using the normalization layer may help to control the divergence of the feature norm among multiple clients. If a local model is designed to perform a class prediction task (a classification task), the final layer may be a classification layer that compares the normalized feature with different class vectors to generate a predicted class label as the prediction output [par. 0051, 0062-0066]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have incorporated the concept of using feature normalization in federated learning to implement a local model by a client/user to extract feature vector from input data as suggested by Zhang into the concept of parallel collaborative machine learning that group local models for the local devices as suggested by Takasaki because both of these systems addressing the process of using federated learning to compare local models of local devices to transmit the result to the global server. Doing so would be desirable because the concept of Takasaki would be more efficient by including a classification item (a classification task) to train a local model related to a task using large amounts of local data, and the feature extraction subnetwork may be dependent on a particular task to be performed by a local model (Zhang, [par. 0003, 0063]).
With respect to claim 2, the combination of Takasaki and Zhang teaches wherein the processing circuitry is further configured to display on the display, a plurality of local model identification information indicating a plurality of local models generated by a plurality of nodes, respectively (Takasaki, the system includes a plurality of groups, and each group includes a plurality of local devices hosting respective local models. For example, a local device may receive a local model from another local device, and there are 5 local devices tagged with numbers from 1 to 5 [par. 0023, 0024, 0031-0034]), and a plurality of classification items indicating classifications of the learning data used for generating the plurality local models, respectively (Zhang, particular neural network layers of the feature extraction subnetwork may be dependent on the classification task to be performed by the local model. For example, if the local model is designed to perform an image processing task, then the neural network layers may include one or more convolutional layers [par. 0061-0064]).
With respect to claim 3, the combination of Takasaki and Zhang teaches wherein the processing circuitry is configured to display on the display unit, own classification item classifying own learning data used for executing learning processing for generating own local model (Zhang, each client uses received global parameters to update its own local model. The client then applies the local model to its own local dataset to compute an update for the local model. The global parameters may be used to execute the client’s own local model to generate predictions [par. 0054, 0057]).
With respect to claim 4, the combination of Takasaki and Zhang teaches wherein the processing circuitry is configured to further display on the display unit, a distribution of the learning data associated with the classification item (Zhang, data heterogeneity means statistical distribution of data is different between different local datasets. A type of data heterogeneity is referred to as label shift, which may occur when different local datasets have different class distribution. Each local dataset has a respective unique label distribution [par. 0004, 0058, 0065]).
With respect to claim 5, the combination of Takasaki and Zhang teaches wherein the processing circuitry is configured to display on the display unit, local model evaluation information indicating an evaluation of said another local model (Takasaki, any local models sorted in a group are exchanged before integration into a global model. Existing local data is inputted to exchange models to evaluate their outputs. Only local models with secured reliability resulting from the evaluation are treated as subjects taken into a global model [par. 0019-0020]).
With respect to claim 7, the combination of Takasaki and Zhang teaches wherein the processing circuitry is configured to receive information indicating said another local model from said another device (Takasaki, each local device receives local model from other local devices. Randomly selecting one or more local models whose accuracies do not exceed a predetermined threshold and sending to the server weight parameters of the selected local models [par. 0004-0008, 0028-0034 and FIG. 2A]).
With respect to claim 8, the combination of Takasaki and Zhang teaches wherein the processing circuitry is configured to execute learning processing using own learning data on said another local model to generate own local model (Takasaki, each of the local devices include a model validator which is responsible for using its own local data to validate accuracies of all the models in the group and a global model. For example, a model validator in each of other devices in the group validates the five local models and the global model using its own local data [par. 0027, 0028]).
With respect to claim 9, the combination of Takasaki and Zhang teaches wherein the processing circuitry is configured to:
receive selection of a global model updated based on a plurality of local models generated by a plurality of nodes (Takasaki, the method for parallel cross validation in collaborative ML may send to the server the weight parameters of selected local models, and the server updates the global model. The central server updates the global model through averaging the weight parameters uploaded from respective selected groups [par. 0004-0008, 0043]); and
transmit selection information indicating that the global model has been selected to said another device (Takasaki, server sends a global model to the respective local devices, receives weight parameters of selected local models, and receives validation results from local devices. Server also updates the global model using weight parameters of selected local models [par. 0004-0008, 0023]).
With respect to claim 10, it is an information processing method claim that is corresponding to the processing circuitry of claim 1. Therefore, it is rejected for the same reason as claimed in claim 1 above.
With respect to claim 11, it is an information processing system claim that is corresponding to the processing circuitry of claim 1. Therefore, it is rejected for the same reason as claimed in claim 1 above.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Takasaki et al (US 20220343219 A1) hereafter Takasaki, in view of Zhang et al (US 20240311645 A1) hereafter Zhang, and further in view of Chakraborty et al (US 20230316090 A1) hereafter Chakraborty.
With respect to claim 6, the combination of Takasaki and Zhang teaches all limitations as claimed in claim 5 above.
However, the combination of Takasaki and Zhang does not explicitly teach wherein the processing circuitry is configured to display on the display unit, local model evaluation information for each classification item.
In the same field of endeavor, Chakraborty teaches wherein the processing circuitry is configured to display on the display unit, local model evaluation information for each classification item (ML models being trained using federated learning techniques that satisfy various inference performance characteristics including classification accuracy. The local evaluation metrics are computed by the federated learning clients using validation data on the same device where the local model was trained [par. 0023]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have incorporated the concept of performing federated learning to send model update data to a server as suggested by Chakraborty into the combination of Takasaki and Zhang because all of these systems addressing the process of using federated learning to generate and update the local model to be sent to the global server. Doing so would be desirable because the combination of Takasaki and Zhang would be more efficient by including the classification accuracy and computing the local evaluation metrics by federated learning clients using validation data on the device where the local model was trained and generated (Chakraborty, [par. 0023]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Spyridopoulos et al (US 20220156368 A1) disclosed a method for detecting an attack on a distributed artificial intelligence deployment comprising a plurality of worker devices. Each of the plurality of worker devices comprises a local machine learning model. Each local machine learning model comprises a plurality of layers. The method comprises calculating a first inference from first input data using a first machine learning model comprising layers of the plurality of layers of one or more of the local machine learning models and calculating additional inferences from the first input data using one or more additional machine learning models. Each of the additional machine learning models comprises at least one of the layers used in the first machine learning model and at least one layer from the pluralities of layers of the one or more local machine learning models that is not used by the first machine learning model. The method further comprises calculating differences between the first inference and each of the one or more additional inferences.
Satheesh Kumar et al (US 20240378457 A1) disclosed a method for distributed machine learning (ML) at a central computing device is provided. The method includes: providing a global ML model to a plurality of local computing devices, wherein the global ML model includes a plurality of parameters; receiving, from each local computing device in a subset of the plurality of local computing devices, a local ML model updated based on the global ML model, wherein the local ML model includes weights with values corresponding to one or more of the plurality of parameters; constructing, for each weight value in each of the received local ML models, a probability distribution for each of the plurality of parameters with corresponding received weight values; sampling, using the constructed probability distribution for each weight value for each of the plurality of parameters, for all of the plurality of local computing devices to generate representative values for each weight; and updating the global ML model by averaging the representative values for each weight for each of the plurality of parameters.
Dai et al (US 20240054350 A1) disclosed a central system may store a neural network model which has a body of a number of layers, and a classification layer comprising class prototypes which classifies the latent representations output by the body of the model. The central system may initialize the class prototypes so that they are uniformly distributed in the representation space. The model and class prototypes may be broadcast to a number of client systems, which update the body of the model locally while keeping the class prototypes fixed. The clients may return information to the central system including updated local model parameters, and a local representation of the classes based on the latent representation of items in the local training data. Based on the information from the clients, the neural network model may be updated. This process may be repeated iteratively.
Subramanya et al (US 20230351245 A1) disclosed an apparatus configured to obtain reliability values for each user equipment in a group of user equipments, obtain, for each user equipment in the group, a reliability value for a training data set stored in the user equipment, each user equipment storing a distinct training data set, and direct a subset of the group of user equipments to separately perform a machine learning training process in the user equipments in the subset, wherein the apparatus is configured to select the subset based on the reliability values for the user equipments and the reliability values for the training data sets.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Quoc Phung whose telephone number is (703) 756 1330. The examiner can normally be reached on Monday through Friday from 9am to 5pm PT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached on 571-272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Q.L.P./Examiner, Art Unit 2143
/JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143