DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Status of Claims
The present application is being examined under the claims filed on 9/26/2025.
Claims 1-22 are pending.
Claims 1-22 are rejected.
Drawings
The drawings are considered acceptable for examination purposes.
Specification
The objection to the abstract of the disclosure and the title, has been withdrawn as necessitated by the amendment.
Claim Interpretation
Claim 16 has been amended, and no longer invokes 35 USC 112 (f).
Claim Rejections - 35 USC § 112
The rejection of claim 14 under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite, has been withdrawn as necessitated by the amendment.
Claim Rejections - 35 USC § 101
The rejection of claims 1-22 under 35 U.S.C. 101 because the claimed invention is directed toward an abstract idea without significantly more, have been withdrawn as necessitated by the amendment to describe a specific manner to create an inference model.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-22 are rejected under 35 U.S.C. 103.
Claims 1-7, 11-13, and 16-22
Claims 1-7, 11-13, and 16-22 are rejected under 35 U.S.C. 103 as being unpatentable over Otsuka (“Information Processing Device, System, and Information Processing Method”—WO 2019022052 A1, and English translation) in view of Shen et al. (“AUROR: Defending Against Poisoning Attacks in Collaborative Deep Learning Systems”; hereinafter “Shen”).
Regarding claim 1
Otsuka teaches:
A machine learning system comprising: a plurality of client terminals; and an integration server (paragraph [0016], “the system 10 includes a server device 100 and client devices 300a, 300b, 300c,....”);
the integration server comprises a first processor and a non-transitory first computer-readable medium storing a trained master model (paragraph [0022], “Referring to FIG. 2, the server apparatus 100 includes a storage 110 […] and a parameter update processing unit 140”; parameter update processing unit is considered to be a processor; paragraph [0023], “In the server apparatus 100, the storage 110 functions as a model holding unit and holds a learning model 111”);
each of the plurality of client terminals comprises a second processor (paragraph [0012], “a processor of the client device calculates a parameter update amount”);
the second processor is configured to: execute machine learning of a learning model using, as learning data, data stored in a data storage apparatus of a medical institution (paragraph [0015], “Specifically, the client device 300a is installed at the location S1, the client device 300b at the location S2, and the client device 300c at the location S3. The places S1, S2, and S3 are places to hold actual data usable as training data for the learning model, more specifically hospitals and establishments, for example”; paragraph [0021], “It should be noted that the client device 300 acquires actual data via a medium (internal network 301, internal transmission path, and removable medium 303) that is independent from the external network 200 connected to the server device 100”; paragraph [0025], “In addition, in the client apparatus 300, the update amount calculation unit 330 executes the training of the learning model 111 received by the model reception unit 310 using the real data acquired by the data acquisition unit 320”; client devices train models using data stored locally at locations such as hospitals, which are a type of medical institution);
transmit a learning result of the learning model to the integration server (paragraph [0025], “in the client apparatus 300 […] the update amount calculation unit 330 calculates the update amount of the parameter P of the learning model 111 based on the result of the training. The update amount transmission unit 340 includes a function of a communication device that transmits data via the external network 200 and transmits the update amount calculated by the update amount calculation unit 330 to the server device 100”; a learning value on the client device determined from machine learning model training is sent to the server device);
the first processor is configured to: synchronize the learning model of each client terminal with the master model before training of the learning model is performed on each of the plurality of client terminals (paragraph [0011], “Also, the server device includes […] a model providing unit for providing the learning model to the client device via the first medium”; sending the model to client devices is considered a synchronization of client models to server model);
receive each of the learning results from the plurality of client terminals (paragraph [0023], “The update amount receiving unit 130 includes a function of a communication device that receives data via the external network 200, and receives an update amount described later from the client device 300”; update amounts are considered learning results); and
create master model candidates the client clusters by integrating the learning results [for the client clusters] (paragraph [0023], “The parameter update processing unit 140 includes a function of a processor that updates the data of the storage 110 and updates at least a part of the parameter P based on the update amount received by the update amount reception unit 130”; update amount results in modification of at least one machine learning parameter, which is considered integration; Examiner notes that for this and subsequent limitations, the portion within the square brackets is not being taught by this particular prior art reference).
Further, Otsuka does not teach create a plurality of client clusters comprising a first client cluster and a second client cluster by dividing the plurality of client terminals into a plurality of groups, wherein the first client cluster includes a first plurality of client terminals and the second client cluster includes a second plurality of client terminals that are not overlapping with the first plurality of client terminals; wherein a first master model candidate of the master model candidates is created from the first client cluster and a second master model candidate of the master model candidates is created from the second client cluster; detect whether any of the master model candidates has an inference accuracy lower than an accuracy threshold value by evaluating a corresponding inference accuracy of each of the master model candidates; extract a first client terminal as an accuracy deterioration cause from the client cluster in response to having detected that the first master model candidate having a first inference accuracy lower than the accuracy threshold value; extract a second client terminal as the accuracy deterioration cause from the second client cluster in response to having detected that the second master model candidate having a second inference accuracy lower than the accuracy threshold value.
perform another iteration of training the first master model candidate by excluding the first client terminal and another iteration of training the second master model candidate by excluding the second client terminal;
and create an inference model from the master model candidates in response to an inference accuracy of each of the master model candidates have exceeded the accuracy threshold value.
However, Shen teaches:
(Shen, pg. 510-511, “Defense Solution” section, Models Mp and MA “A global model trained using AUROR MA is robust and effective if the attack success rate SRMA and the accuracy drop AD (MB →MA) of MA are small enough to be acceptable for practical purposes”)-- create a plurality of client clusters comprising a first client cluster and a second client cluster by dividing the plurality of client terminals into a plurality of groups, wherein the first client cluster includes a first plurality of client terminals and the second client cluster includes a second plurality of client terminals that are not overlapping with the first plurality of client terminals; wherein a first master model candidate of the master model candidates is created from the first client cluster and a second master model candidate of the master model candidates is created from the second client cluster; detect the master model candidate having an inference accuracy lower than an accuracy threshold value by evaluating the inference accuracy of each of the master model candidates created ;
Shen teaches (pg. 514, section 5, 1-3rd paragraph, table 2 “AUROR filters the malicious users before creating the final model…The second step in AUROR is to identify the malicious users based on the indicative features […] The users that appear in suspicious clusters for more than τ = 50% of the total indicative features are confirmed as malicious users”; such clients are identified as contributing to a lower accuracy score)-- extract a first client terminal as an accuracy deterioration cause [from the client cluster used] from the client cluster in response to having detected that the first master model candidate having a first inference accuracy lower than the accuracy threshold value; extract a second client terminal as the accuracy deterioration cause from the second client cluster in response to having detected that the second master model candidate having a second inference accuracy lower than the accuracy threshold value.
Additionally, Shen teaches dividing users into different clusters. Every cluster where the number of users is smaller than n/2 are labeled as suspicious. The users that appear in suspicious cluster more than T=50% of the total indicative features are confirmed as malicious users. Then, the malicious users are excluded and the global model is trained without the malicious users (pg. 514, section 5, 4-5th paragraph)-- perform another iteration of training the first master model candidate by excluding the first client terminal and another iteration of training the second master model candidate by excluding the second client terminal;
Further, Shen teaches filtering out or excluding input values from malicious user computers, generating and training a global model on the remaining data (pg.514, sect.5)-- and create an inference model from the master model candidates in response to an inference accuracy of each of the master model candidates have exceeded the accuracy threshold value.
Otsuka and Shen are considered analogous to the claimed invention since they are making improvements to the federated learning process. It would have been obvious to a person having ordinary skill in the art (hereinafter “PHOSITA”), before the effective filing date of the claimed invention, to incorporate the teachings of Shen into those of Otsuka. One would be motivated to do so in order to improve the security and efficacy of models trained using distributed learning techniques (Shen, Abstract, paragraph 3).
Regarding claim 2
The combination of Otsuka and Shen does not teach, but Shen teaches dividing users into different clusters. Every cluster where the number of users is smaller than n/2 are labeled as suspicious. The users that appear in suspicious cluster more than T=50% of the total indicative features are confirmed as malicious users. Then, the malicious users are excluded and the global model is trained without the malicious users (pg. 514, section 5, 4-5th paragraph)--
wherein the first processor is configured to exclude the first client terminal as the accuracy deterioration cause from any subsequent learning after perform the another iteration of training the first master model candidate. It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Shen into those of Otsuka and Ghosh. One would be motivated to do so in order to improve the security and efficacy of models trained using distributed learning techniques (Shen, Abstract, paragraph 3).
Regarding claim 3
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, and Shen teaches wherein the first computer-readable medium stores information of the client terminal as the accuracy deterioration cause (pg. 514, section 5, third paragraph, “The second step in AUROR is to identify the malicious users based on the indicative features […] The users that appear in suspicious clusters for more than τ = 50% of the total indicative features are confirmed as malicious users”; Algorithm identifies malicious users; based on the information, these users do not participate in further learning in subsequent steps as detailed prior). It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Shen into those of Otsuka and Ghosh. One would be motivated to do so in order to improve the security and efficacy of models trained using distributed learning techniques (Shen, Abstract, paragraph 3).
Regarding claim 4
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, and Otsuka teaches wherein each of the plurality of client terminals is a terminal provided in a medical institution network of different medical institutions (paragraph [0017], “On the other hand, the client devices 300a, 300b, 300c,... Are installed at different places. Specifically, the client device 300a the location S1, the client device 300b is the location S2, the client device 300c to a location S3, are installed respectively. The places S1, S2 and S3 are places that hold actual data that can be used as training data for a learning model, specifically, for example, a hospital or a business place”).
Regarding claim 5
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, and Otsuka teaches wherein the integration server is provided in a medical institution network or outside the medical institution network (paragraph [0016], “The server apparatus 100 is connected to the client apparatuses 300 a, 300 b, 300 c,... Via the external network 200. Here, the external network 200 includes, for example, the Internet”; server device may exist in medical institution network or somewhere outside the network).
Regarding claim 6
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, and Otsuka teaches wherein the learning result transmitted from the client terminal to the integration server includes a weight parameter of the trained learning model (paragraph [0032], “the parameter update processing section 140 updates the update transmitted from each client apparatus 300 The [update] amount ΔP may be weighted”; weight parameter of a model is not specific and may be considered a weighted value).
Regarding claim 7
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, and Shen teaches wherein the data used as the learning data includes at least one type of data among a two-dimensional image, a three-dimensional image, a moving image, time-series data, and document data (pg. 512, section 4.1, first paragraph, “We use the MNIST dataset of handwritten digits [24], a popular benchmark for training and testing deep learning models”; MNIST dataset containing two-dimensional images of digits used as learning data). It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Shen into those of Otsuka and Ghosh. One would be motivated to do so in order to improve the security and efficacy of models trained using distributed learning techniques (Shen, Abstract, paragraph 3).
Regarding claim 11
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, and Ghosh teaches wherein the number of the client terminals included in each of the plurality of client clusters is the same (pg. 9, section 6.1, first paragraph, “Then ⌊(1−𝛼)𝑚⌋ machines are uniformly assigned to the K clusters, and ⌈𝛼𝑚⌉ machines are considered adversarial machines”; uniform assignment to clusters implies that each cluster receives the same number of machines where no machine is a part of more than one cluster). It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Ghosh into those of Otsuka and Shen. The motivation to do so involves increasing the robustness and performance of models trained using distributed learning techniques (Ghosh, Abstract).
Regarding claim 12
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, and Ghosh teaches wherein the first computer-readable medium stores information indicating a correspondence relationship as to which client cluster among the plurality of client clusters each of the plurality of master model candidates created is based on (pg. 4, section 3.2, first paragraph, “All m compute nodes send local ERMs [empirical risk minimizers], w(i), for i ⋲ [m]1 to the center machine, and
the center machine runs a clustering algorithm on these data points to find K clusters C1, . . . , CK”; center node runs and stores the results of a clustering algorithm, indicating that the center node stores information about which cluster each machine belongs to). It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Ghosh into those of Otsuka and Shen. The motivation to do so involves increasing the robustness and performance of models trained using distributed learning techniques (Ghosh, Abstract).
Regarding claim 13
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, and Shen teaches wherein the first processor is configured to determine whether or not the inference accuracy of the master model candidate is lower than the accuracy threshold value based on a comparison between an instantaneous value of the inference accuracy of each of the master model candidates and the accuracy threshold value, or based on a comparison between a statistical value of the inference accuracy in a learning iteration of each of the master model candidates and the accuracy threshold value (Shen, pg. 515, “Evaluating the Final Model” section, second paragraph, “We measure the accuracy drop of the final global model as compared to the benign model and study the improvement over the poisoned model”; accuracy of benign model is seen as the threshold accuracy). It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Shen into those of Otsuka and Ghosh. One would be motivated to do so in order to improve the security and efficacy of models trained using distributed learning techniques (Shen, Abstract, paragraph 3).
Regarding claim 16
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, and Otsuka teaches further (Otsuka, paragraph [0022], “the server apparatus 100 includes a storage 110”) while Shen teaches wherein the non-transitory first computer-readable medium is further configured to store verification data and the first processor is configured to evaluate the inference accuracy of the master model candidate using the verification data (Shen, pg. 515, “Evaluating the Final Model” section, second paragraph, “We measure the accuracy drop of the final global model as compared to the benign model and study the improvement over the poisoned model”; benign model accuracy of Shen reference is seen as verification data to evaluate global model against; such data necessarily must be present in a storage device). It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Shen into those of Otsuka and Ghosh. One would be motivated to do so in order to improve the security and efficacy of models trained using distributed learning techniques (Shen, Abstract, paragraph 3).
Regarding claim 17
Otsuka teaches:
A machine learning method using a plurality of client terminals and an integration server (paragraph [0016], “the system 10 includes a server device 100 and client devices 300 a, 300 b, 300 c,....”), the method comprising:
synchronizing a learning model of each client terminal with a trained master model stored in the integration server before training of the learning model is performed on each of the plurality of client terminals (paragraph [0011], “Also, the server device includes […] a model providing unit for providing the learning model to the client device via the first medium”; sending the model to client devices is considered a synchronization of client models to server model);
executing machine learning of the learning model using, as learning data, data stored in a data storage apparatus of each of medical institutions different from each other by each of the plurality of client terminals (paragraph [0021], “It should be noted that the client device 300 acquires actual data via a medium (internal network 301, internal transmission path, and removable medium 303) that is independent from the external network 200 connected to the server device 100”; paragraph [0025], “In addition, in the client apparatus 300, the update amount calculation unit 330 executes the training of the learning model 111 received by the model reception unit 310 using the real data acquired by the data acquisition unit 320”);
transmitting a learning result of the learning model to the integration server from each of the plurality of client terminals (paragraph [0025], “in the client apparatus 300… the update amount calculation unit 330 calculates the update amount of the parameter P of the learning model 111 based on the result of the training. The update amount transmission unit 340 includes a function of a communication device that transmits data via the external network 200 and transmits the update amount calculated by the update amount calculation unit 330 to the
model training is sent to the server device);
receiving each of the learning results from the plurality of client terminals by the integration server (paragraph [0023], “The update amount receiving unit 130 includes a function of a communication device that receives data via the external network 200, and receives an update amount described later from the client device 300”); and
creating master model candidates for the clients by the integration server, by integrating the learning results [for each of the client clusters] (paragraph [0023], “The parameter update processing unit 140 includes a function of a processor that updates the data of the storage 110 and updates at least a part of the parameter P based on the update amount received by the update amount reception unit 130”; update amount results in modification of at least one machine learning parameter, which is considered integration).
Otsuka does not teach creating master model candidates for the client clusters; creating a plurality of client clusters comprising a first client cluster and a second client cluster by the integration server, by dividing the plurality of client terminals into a plurality of groups, wherein the first client cluster includes a first plurality of client terminals and the second client cluster includes a second plurality of client terminals that are not overlapping with the first plurality of client terminals;
wherein a first master model candidate of the master model candidates is created from the first client cluster and a second master model candidate of the master model candidates is created from the second client cluster; detecting whether any of the master model candidates having has an inference accuracy lower than an accuracy threshold value by the integration server, by evaluating a corresponding inference accuracy of each of the master model candidates; extracting a first client terminal as an accuracy deterioration cause from the client cluster in response to having detected that the firstmaster model candidate having a first inference accuracy lower than the accuracy threshold value by the integration server; extracting a second client terminal as the accuracy deterioration cause from the second client cluster in response to having detected that the second master model candidate having a second inference accuracy lower than the accuracy threshold value; performing another iteration of training the first master model candidate by excluding the first client terminal and another iteration of training the second master model candidate by excluding the second client terminal; and creating an inference model from the master model candidates in response to an inference accuracy of each of the master model candidates have exceeded the accuracy threshold value.
However, Shen teaches:
(Shen, pg. 510-511, “Defense Solution” section, Models Mp and MA “A global model trained using AUROR MA is robust and effective if the attack success rate SRMA and the accuracy drop AD (MB →MA) of MA are small enough to be acceptable for practical purposes”)-- creating a plurality of client clusters comprising a first client cluster and a second client cluster by the integration server, by dividing the plurality of client terminals into a plurality of groups, wherein the first client cluster includes a first plurality of client terminals and the second client cluster includes a second plurality of client terminals that are not overlapping with the first plurality of client terminals; wherein a first master model candidate of the master model candidates is created from the first client cluster and a second master model candidate of the master model candidates is created from the second client cluster; detecting whether any of the master model candidate has an inference accuracy lower than an accuracy threshold value by evaluating a corresponding inference accuracy of each of the master model candidates;
(Shen, pg. 510, “Defense Solution” section, “A global model trained using AUROR MA is robust and effective if the attack MA and the accuracy drop AD (MB →MA) of MA are small enough to be acceptable for practical purposes”); and
Shen teaches (pg. 514, section 5, 1-3rd paragraph, table 2 “AUROR filters the malicious users before creating the final model…The second step in AUROR is to identify the malicious users based on the indicative features […] The users that appear in suspicious clusters for more than τ = 50% of the total indicative features are confirmed as malicious users”; such clients are identified as contributing to a lower accuracy score)-- extracting a first client terminal as an accuracy deterioration cause from the client cluster in response to having detected that the first master model candidate having a first inference accuracy lower than the accuracy threshold value by the integration server; extracting a second client terminal as the accuracy deterioration cause from the second client cluster in response to having detected that the second master model candidate having a second inference accuracy lower than the accuracy threshold value.
Additionally, Shen teaches dividing users into different clusters. Every cluster where the number of users is smaller than n/2 are labeled as suspicious. The users that appear in suspicious cluster more than T=50% of the total indicative features are confirmed as malicious users. Then, the malicious users are excluded and the global model is trained without the malicious users (pg. 514, section 5, 4-5th paragraph)--performing another iteration of training the first master model candidate by excluding the first client terminal and another iteration of training the second master model candidate by excluding the second client terminal;
Further, Shen teaches filtering out or excluding input values from malicious user computers, generating and training a global model on the remaining data (pg.514, sect.5)-- and creating an inference model from the master model candidates in response to an inference accuracy of each of the master model candidates have exceeded the accuracy threshold value.
Otsuka and Shen are considered analogous to the claimed invention since they are making improvements to the federated learning process. It would have been obvious to a person having ordinary skill in the art (hereinafter “PHOSITA”), before the effective filing date of the claimed invention, to incorporate the teachings of Shen into those of Otsuka. One would be motivated to do so in order to improve the security and efficacy of models trained using distributed learning techniques (Shen, Abstract, paragraph 3).
Claim 18 is directed towards an integration server for performing the limitations found in claim 1, and is likewise rejected.
Regarding claim 19
The combination of Otsuka and Shen teaches An information processing apparatus that is used as one of the plurality of client terminals connected to the integration server according to claim 18 via a communication line. Otsuka teaches the apparatus comprising:
a second processor (paragraph [0012], “a processor of the client device calculates a parameter update amount”);
a second computer-readable medium as a non-transitory tangible medium in which a second program to be executed by the second processor is recorded (paragraph [0044], “The memory 903 includes, for example, a ROM (Read Only Memory) and a RAM (Random Access Memory). The ROM stores, for example, a program for the processor 901”; part of information processing apparatus which can be on a client device);
wherein the second processor is configured to, according to an instruction of the second program, execute machine learning of a learning model by setting, as the learning model in an initial state before learning is started, a learning model synchronized with the master model stored in the integration server (paragraph [0024], “in the client device 300, the model receiving unit 310 includes a function of a communication device that receives data via the external network 200, and receives the learning model 111 transmitted from the server device 100”);
using, as learning data, data stored in a data storage apparatus of a medical institution (paragraph [0015], “Specifically, the client device 300a is installed at the location S1, the client device 300b at the location S2, and the client device 300c at the location S3. The places S1, S2 and S3 are places to hold actual data usable as training data for the learning model, more specifically hospitals and establishments, for example”; paragraph [0021], “It should be noted that the client device 300 acquires actual data via a medium (internal network 301, internal transmission path, and removable medium 303) that is independent from the external network
300, the update amount calculation unit 330 executes the training of the learning model 111 received by the model reception unit 310 using the real data acquired by the data acquisition unit 320”; client devices train models using data stored locally at locations such as hospitals, which are a type of medical institution); and
transmit a learning result of the learning model to the integration server (paragraph [0025], “in the client apparatus 300… the update amount calculation unit 330 calculates the update amount of the parameter P of the learning model 111 based on the result of the training. The update amount transmission unit 340 includes a function of a communication device that transmits data via the external network 200 and transmits the update amount calculated by the update amount calculation unit 330 to the server device 100”; a learning value on the client device determined from machine learning model training is sent to the server device).
Regarding claim 20
The combination of Otsuka and Shen teaches A non-transitory computer readable medium storing a program causing a computer to function as one of the plurality of client terminals connected to the integration server according to claim 18 via a communication line (Otsuka, paragraph [0044], “The memory 903 includes, for example, a ROM (Read Only Memory) and a RAM (Random Access Memory). The ROM stores, for example, a program for the processor 901”; part of information processing apparatus which can be on a client device). Otsuka teaches the program causing the computer to realize:
a function of executing machine learning of a learning model by setting, as the learning model in an initial state before learning is started, a learning model synchronized with the master model stored in the integration server (paragraph [0024], “in the client device 300, the model external network 200, and receives the learning model 111 transmitted from the server device 100”);
using, as learning data, data stored in a data storage apparatus of a medical institution (paragraph [0015], “Specifically, the client device 300a is installed at the location S1, the client device 300b at the location S2, and the client device 300c at the location S3. The places S1, S2 and S3 are places to hold actual data usable as training data for the learning model, more specifically hospitals and establishments, for example”; paragraph [0021], “It should be noted that the client device 300 acquires actual data via a medium (internal network 301, internal transmission path, and removable medium 303) that is independent from the external network 200 connected to the server device 100”; paragraph [0025], “In addition, in the client apparatus 300, the update amount calculation unit 330 executes the training of the learning model 111 received by the model reception unit 310 using the real data acquired by the data acquisition unit 320”; client devices train models using data stored locally at locations such as hospitals, which are a type of medical institution); and
a function of transmitting a learning result of the learning model to the integration server (paragraph [0025], “in the client apparatus 300… the update amount calculation unit 330 calculates the update amount of the parameter P of the learning model 111 based on the result of the training. The update amount transmission unit 340 includes a function of a communication device that transmits data via the external network 200 and transmits the update amount calculated by the update amount calculation unit 330 to the server device 100”; a learning value on the client device determined from machine learning model training is sent to the server device).
Claim 21 is directed towards a non-transitory computer readable medium for storing the limitations found in claim 1, and is likewise rejected.
Regarding claim 22
Otsuka teaches:
An inference model creation method for creating an inference model by performing machine learning using a plurality of client terminals and an integration server, the method comprising: synchronizing a learning model of each client terminal with a trained master model stored in the integration server before training of the learning model is performed on each of the plurality of client terminals (paragraph [0011], “Also, the server device includes […] a model providing unit for providing the learning model to the client device via the first medium”; sending the model to client devices is considered a synchronization of client models to server model);
executing machine learning of the learning model using, as learning data, data stored in a data storage apparatus of each of medical institutions different from each other via each of the plurality of client terminals (paragraph [0021], “It should be noted that the client device 300 acquires actual data via a medium (internal network 301, internal transmission path, and removable medium 303) that is independent from the external network 200 connected to the server device 100”; paragraph [0025], “In addition, in the client apparatus 300, the update amount calculation unit 330 executes the training of the learning model 111 received by the model reception unit 310 using the real data acquired by the data acquisition unit 320”);
transmitting a learning result of the learning model to the integration server from each of the plurality of client terminals (paragraph [0025], “in the client apparatus 300… the update amount calculation unit 330 calculates the update amount of the parameter P of the learning model 111 based on the result of the training. The update amount transmission unit 340 includes a function of a communication device that transmits data via the external network 200 and transmits the update amount calculated by the update amount calculation unit 330 to the server device 100”; a learning value on the client device determined from machine learning model training is sent to the server device);
receiving each of the learning results from the plurality of client terminals by the integration server (paragraph [0023], “The update amount receiving unit 130 includes a function of a communication device that receives data via the external network 200, and receives an update amount described later from the client device 300”); and
creating master model candidates for each of the client clusters by the integration server, by integrating the learning results [for each of the client clusters] (paragraph [0023], “The parameter update processing unit 140 includes a function of a processor that updates the data of the storage 110 and updates at least a part of the parameter P based on the update amount received by the update amount reception unit 130”; update amount results in modification of at least one machine learning parameter, which is considered integration).
Otsuka does not teach detecting the master model candidate having an inference accuracy lower than an accuracy threshold value by the integration server, by evaluating the inference accuracy of each of the master model candidates created for each of the client clusters or extracting a client terminal as an accuracy deterioration cause from the client cluster used for creation of the master model candidate having the inference accuracy lower than the accuracy threshold value by the integration server.
However, Shen teaches:
detecting the master model candidate having an inference accuracy lower than an accuracy threshold value by the integration server, by evaluating the inference accuracy of each of the master model candidates created [for each of the client clusters] (Shen, pg. 510, “Defense Solution” section, “A global model trained using AUROR MA is robust and effective if the attack success rate SRMA and the accuracy drop AD (MB →MA) of MA are small enough to be acceptable for practical purposes”);
extracting a client terminal as an accuracy deterioration cause [from the client cluster used] for creation of the master model candidate having the inference accuracy lower than the accuracy threshold value by the integration server (pg. 514, section 5, third paragraph, “The second step in AUROR is to identify the malicious users based on the indicative features […] The users that appear in suspicious clusters for more than τ = 50% of the total indicative features are confirmed as malicious users”); and
creating the inference model having an inference accuracy higher than the inference accuracy of the master model based on the master model candidate having an inference accuracy equal to or higher than the accuracy threshold value by the integration server (pg. 514, section 5, fourth paragraph, “The server excludes the input values from the malicious users identified in the previous step and trains the global model on the remaining masked features”).
It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Shen into those of Otsuka. One would be motivated to do so in order to improve the security and efficacy of models trained using distributed learning techniques (Shen, Abstract, paragraph 3).
The combination of Otsuka and Shen does not teach creating a plurality of client clusters by the integration server, by dividing the plurality of client terminals into a plurality of groups. However, Ghosh teaches this limitation (pg. 3, section 2, first paragraph, “Out of the non-Byzantine compute nodes, we assume that there are K different data distributions, D1, . . . ,DK, and that the (1−α)m machines are partitioned into K clusters, C1, . . . , CK”). Additionally, Ghosh further teaches the limitations for each of the client clusters and for the client cluster used, where the teachings of Otsuka and Shen that describe creating master model candidates, evaluating the inference accuracy of each of the master model candidates, and extracting a client terminal as an accuracy deterioration cause can be substituted by clusters upon which those same functions are performed. It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Ghosh into those of Otsuka and Shen. The motivation to do so involves increasing the robustness and performance of models trained using distributed learning techniques (Ghosh, Abstract).
Claims 8, 9, and 15
Claims 8, 9, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Otsuka and Shen as applied to claim 1 above, and further in view of McMahan et al. (“Federated Learning of Deep Networks using Model Averaging”; hereinafter “McMahan”).
Regarding claim 8
The combination of Otsuka and Shen teaches The machine learning system according to claim 1 but does not teach wherein each model of the learning model, the master model, and the master model candidate is configured by using a neural network. However, McMahan teaches this limitation (pg. 5, section 3, first paragraph, “Our initial study [of the FederatedAveraging algorithm] includes three model families on two datasets. The first two are for the MNIST digit recognition task (LeCun et al., 1998): […] 2) A CNN with two 5x5 convolution layers (the first with 32 channels, the second with 64, each followed with 2x2 max pooling), a fully connected layer with 512 units and ReLu activation, and a final softmax output layer (1,663,370 total parameters)”; CNN (a type of neural network) used in federated learning scheme and can be considered to be “configured” during model training and aggregation).
McMahan is considered analogous to the claimed invention since they make improvements on federated learning techniques. It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the methodologies of McMahan into those of Otsuka and Shen. One would be motivated to do so to implement the popular FedAveraging algorithm of distributed learning (McMahan, Abstract, second paragraph; Conclusion section).
Regarding claim 9
The combination of Otsuka and Shen teaches The machine learning system according to claim 1 but does not teach wherein the data used as the learning data includes a two-dimensional image, a three-dimensional image, or a moving image. However, McMahan teaches this limitation (pg. 5, section 3, first paragraph, “We study two ways of partitioning the MNIST data over clients: IID, where the data is shuffled, and then partitioned into 100 clients each receiving 600 examples”; MNIST dataset containing two-dimensional images of handwritten digits used as learning data). It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the methodologies of McMahan into those of Otsuka and Shen. One would be motivated to do so to implement the popular FedAveraging algorithm of distributed learning (McMahan, Abstract, second paragraph; Conclusion section).
Regarding claim 15
The combination of Otsuka and Shen teaches The machine learning system according to claim 1, where Otsuka teaches wherein the integration server further comprises a display device (Otsuka, paragraph [0041], “The information processing apparatus 900 shown in FIG. 5 functions, for example, as the server apparatus 100”; paragraph [0042], “The information processing apparatus 900 includes […] an output device 907”; paragraph [0046], “The output device 907 may include, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display”; Server apparatus of the prior art contains an output device which can be a form of display device) and the display device is configured to display the inference accuracy [in each learning iteration of each of the master model candidates created for each of the client clusters] (Otsuka, paragraph [0046], “The output device 907 outputs the result obtained by the processing of the information processing device 900 as a video such as text or image, sound such as voice or sound, vibration, or the like”).
Otsuka does not teach displaying the inference accuracy in each learning iteration of each of the master model candidates created. However, McMahan teaches this limitation (McMahan, pg. 6, at least fig. 2 shows a graph of test accuracy vs. number of rounds of communication (analogous to learning iterations) which constitutes visual output that can be shown on a display device). It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the methodologies of McMahan into those of Otsuka and Shen. One would be motivated to do so to implement the popular FedAveraging algorithm of distributed learning (McMahan, Abstract, second paragraph; Conclusion section).
Otsuka and McMahan do not teach performing this operation for each client cluster. However, Ghosh further teaches the limitation for each of the client clusters where the teachings of Otsuka and McMahan that describe the remaining portion of the limitation can be substituted by clusters upon which those same functions are performed. It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the teachings of Ghosh into those of Otsuka, Shen, and McMahan. The motivation to do so involves increasing the robustness and performance of models trained using distributed learning techniques (Ghosh, Abstract).
Claim 10
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Otsuka and Shen as applied to claim 1 above, and further in view of Goldberg et al. (“A Primer on Neural Network Models for Natural Language Processing”; hereinafter Goldberg). The combination of Otsuka and Shen teaches The machine learning system according to claim 1 but does not teach wherein the data used as the learning data includes time-series data or document data, and each model of the learning model, the master model, and the master model candidate is configured by using a recursive neural network. However, Goldberg teaches this limitation (pg. 404, third paragraph, “A recursive neural network (RecNN) is a function that takes as input a parse tree over an n-word sentence x1, . . . , xn”; words and sentences can be considered document data, and such a neural network can be used in place of those described in Otsuka’s paragraph [0023]).
Goldberg is considered analogous to the claimed invention since they are in the same field of machine learning. It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the methodologies of Goldberg into those of Otsuka and Shen. Otsuka teaches federated learning using a variety of machine learning models as described in paragraph [0023], and Goldberg discloses a recursive neural network that uses a particular input. This is seen as applying a known technique (recursive neural network) to a known product (federated learning algorithm) ready for improvement to yield predictable results.
Claim 14
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Otsuka and Shen as applied to claim 1 above, and further in view of Bagdasaryan et al. (“How to Backdoor Federated Learning”; hereinafter “Bagdasaryan”).
The combination of Otsuka, and Shen, teaches The machine learning system according to claim 1 but does not teach wherein, in a case where the master model candidate having the inference accuracy lower than the accuracy threshold value is detected, the first processor is configured to notify the integration server of information of the master model candidate related to the detection. However, Bagdasaryan teaches this limitation (pg. 7, “Accuracy Auditing” section, “Another plausible anomaly detection technique is to reject local models whose accuracy on the main task is abnormally low”).
Bagdasaryan is considered analogous to the claimed invention since they improve upon federated learning techniques. It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the methodologies of Goldberg into those of Otsuka, and Shen,. One would be motivated to make this change to improve the robustness and security of distributed learning systems (Bagdasaryan, pg. 10, section 7, third paragraph).
Response to Arguments
Applicant's arguments filed 9/6/2025 have been fully considered but they are not persuasive. Applicant states that “The main issue is that none of the three references used to reject independent claims 1, 17, 18, 21, and 22 disclose the concept of the intermediate level of master model candidates that is between the inference model to be finalized and the local models that is used locally. Also, while the Ghosh reference was cited to disclose the concept of clustering, the Ghosh reference does not contain any details that would suggest the implementation of cluster is the same as the claimed invention. The Ghosh does not disclose or suggest "perform another iteration of training the first master model candidate by excluding the first client terminal and another iteration of training the second master model candidate by excluding the second client terminal". The prior art would not teach excluding two terminals from two different clusters, since the prior art at best suggest excluding one terminal in one cluster. Therefore, based on the amendments and the reasons provided, the curently cited
references, alone or in combination, does not render obvious independent claims 1, 17, 18, 21, and 22. All other claim are also novel and non-obvious due to claim dependencies.”. The examiner disagrees, because as indicated above Shen teaches dividing users into different clusters. Every cluster where the number of users is smaller than n/2 are labeled as suspicious. The users that appear in suspicious cluster more than T=50% of the total indicative features are confirmed as malicious users. Then, the malicious users are excluded and the global model is trained without the malicious users (pg. 514, section 5, 4-5th paragraph).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CESAR PAULA whose telephone number is (571)272-4128. The examiner can normally be reached Monday - Friday, 6.30am- 4:30 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Wiley can be reached at (571)272-3923. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145