DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The term “more specific” in claims 5-7 is a relative term which renders the claim indefinite. The term “more specific” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (mental process) without significantly more.
Claim 1:
Regarding claim 1, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites
“A system, comprising: a first computing device configured to gather memory usage data and device characteristic data associated with a first plurality of memory devices monitored by the first computing device …”, and a system is one of the four statutory categories of invention.
In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components:
and predict aging of the first plurality of memory devices and the second plurality of memory devices… (mental process, a person can mentally evaluate and predict aging of first and second memory devices, see MPEP 2106.04(a)(2)(III)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
A system, comprising: a first computing device configured to: (this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
gather memory usage data and device characteristic data associated with a first plurality of memory devices monitored by the first computing device; (In step 2A, prong 2, this recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)),
and train a first machine learning model based on the gathered memory usage data and the device characteristic data associated with the first plurality of memory devices; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
a second computing device configured to: … (this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
gather memory usage data and device characteristic data associated with a second plurality of memory devices monitored by the second computing device; (In step 2A, prong 2, this recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)),
and train a second machine learning model based on the gathered memory usage data and the device characteristic data associated with the second plurality of memory devices; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
and a local federated server in communication with the first computing device and the second computing device and configured to: … (this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
aggregate the first machine learning model and the second machine learning model into a third machine learning model; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
a global federated server in communication with the local federated server and configured to: … (this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
aggregate the third machine learning model with a fourth machine learning model comprising a plurality of aggregated machine learning models into a fifth machine learning model; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
… based on the fifth machine learning model, (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional elements ii, v, viii, and x recite using generic computer components as a tool. Additional elements iv, vii, ix, xi, and xii recite mere instructions to apply the judicial exception using generic computer components, which are not indicative of significantly more. The additional elements iii and vi recite mere data gathering, and are considered insignificant extra-solution activities. In step 2B, these insignificant extra-solution activities are well understood routine and conventional activities, which include receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)),
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 2:
Regarding claim 2, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 2 recites the following additional element:
The system of claim 1, wherein the first plurality of memory devices comprises a first categorized group of memory devices having first device characteristics different than a second categorized group of memory devices having second device characteristics of the second plurality of memory devices. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 3:
Regarding claim 3, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 3 recites the following abstract ideas:
to predict aging of the first plurality of memory devices … (This recites a mental process, a person can mentally evaluate and predict aging for the first plurality of memory devices, see MPEP 2106.04(a)(2)(III)),
to predict aging of the second plurality of memory devices … (This recites a mental process, a person can mentally evaluate and predict aging for the second plurality of memory devices, see MPEP 2106.04(a)(2)(III)),
Further, claim 3 recites the following additional elements:
The system of claim 1, wherein the first computing device is configured … (In step 2A, prong 2, and step 2B, this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
… based on the first machine learning model, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
and the second computing device is configured … (In step 2A, prong 2 and step 2B, this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
… based on the second machine learning model. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 4:
Regarding claim 4, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1.
Further, claim 4 recites the following abstract idea:
predict aging of the first plurality of memory devices and the second plurality of memory devices, (This recites a mental process, a person can mentally evaluate and predict aging for the first and second plurality of memory devices, see MPEP 2106.04(a)(2)(III)),
Further, claim 4 recites the following additional element:
The system of claim 1, wherein the local federated server is configured to … based on the third machine learning model. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 5:
Regarding claim 5, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 5 recites the following additional element:
The system of claim 1, wherein the first machine learning model provides a more specific aging prediction of the first plurality of memory devices as compared to the third machine learning model; (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
and wherein the second machine learning model provides a more specific aging prediction of the second plurality of memory devices as compared to the third machine learning model.(In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 6:
Regarding claim 6, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 6 recites the following additional element:
The system of claim 1, wherein the first machine learning model provides a more specific aging prediction of the first plurality of memory devices; (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
and wherein the second machine learning model provides a more specific aging prediction of the second plurality of memory devices as compared to the fifth machine learning model. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 7:
Regarding claim 7, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 7 recites the following additional element:
The system of claim 1, wherein the third machine learning model provides a more specific aging prediction of the first plurality of memory devices and the second plurality of memory devices as compared to the fifth machine learning model. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 8:
Regarding claim 8, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites
“A system, comprising: a first plurality of memory devices grouped together based on memory device characteristics of the first plurality of memory devices …”, and a system is one of the four statutory categories of invention.
In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components:
determine a first mean time to failure (MTTF) for each one of the first plurality of memory devices …; (mental process, a person can mentally evaluate and determine a MTTF for each of the first memory devices, see MPEP 2106.04(a)(2)(III)),
determine a second MTTF for each one of the second plurality of memory devices …; (mental process, a person can mentally evaluate and determine a MTTF for each of the second memory devices, see MPEP 2106.04(a)(2)(III)),
and determine a respective MTTF for each of the first plurality of memory devices and the second plurality of memory devices …(mental process, a person can mentally evaluate and determine a respective MTTF for each of the first memory devices and second memory devices, see MPEP 2106.04(a)(2)(III)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
A system, comprising: (this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
a first plurality of memory devices grouped together based on memory device characteristics of the first plurality of memory devices; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
a second plurality of memory devices, grouped together based on memory device characteristics of the second plurality of memory devices; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
a first computing device… (this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
configured to: gather memory usage data and the memory device characteristics from the first plurality of memory devices; (In step 2A, prong 2, this recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)),
train a first local machine learning model based on the gathered memory usage data and the memory device characteristics for the first plurality of memory devices; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
based on the first local machine learning model, (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
a second computing device configured to: … (this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
gather memory usage data and the memory device characteristics from the second plurality of memory devices; (In step 2A, prong 2, this recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)),
train a second local machine learning model based on the gathered memory usage data and the memory device characteristics for the second plurality of memory devices; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
based on the second local machine learning model, (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
a local federated server configured to: … (this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
aggregate the first and the second local machine learning models into a first MTTF machine learning model; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
and a global federated server configured to: … (this recites a generic computer component being used as a tool – see MPEP 2106.05(f)),
aggregate weights derived from the first generic machine learning model with weights derived from a second generic machine learning model into a global MTTF machine learning model; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
… based on the first local machine learning model, the second local machine learning model, the first generic machine learning model, the second generic machine learning model, the global MTTF machine learning model, or any combination thereof. (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional elements iv, vii, xi, xv, and xvii recite using generic computer components as a tool. Additional elements v, vi, ix, x, xiii, xiv, xvi, xviii, and xix recite mere instructions to apply the judicial exception using generic computer components, which are not indicative of significantly more. The additional elements viii and xii recite mere data gathering, and are considered insignificant extra-solution activities. In step 2B, these insignificant extra-solution activities are well understood routine and conventional activities, which include receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)),
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 9:
Regarding claim 9, it is dependent upon claim 8, and thereby incorporates the limitations of, and corresponding analysis applied to claim 8. Further, claim 9 recites the following abstract idea:
to: determine one of the first plurality of memory devices has reached an MTTF; (This recites a mental process, a person can mentally evaluate and determine if first memory devices reached an MTTF, see MPEP 2106.04(a)(2)(III)),
Further, claim 9 recites the following additional elements:
The system of claim 8, wherein the global federated server, the local federated server, the first computing device, or a combination thereof is configured … (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)),
and based on determination, provide an alert to a host device of the plurality of first memory devices. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 10:
Regarding claim 10, it is dependent upon claim 8, and thereby incorporates the limitations of, and corresponding analysis applied to claim 8. Further, claim 10 recites the following additional element:
The system of claim 8, wherein the local federated server is configured to update the first generic machine learning model in response to a change in the first local machine learning model, the second machine learning model, or both. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 11:
Regarding claim 11, it is dependent upon claim 8, and thereby incorporates the limitations of, and corresponding analysis applied to claim 8. Further, claim 11 recites the following additional element:
The medium of claim 8, global federated server is configured to update the global machine learning model in response to a change in the first local machine learning model, the second local machine learning model, the first generic machine learning model, the second generic machine learning model, or any combination thereof. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 12:
Regarding claim 12, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites
“A method, comprising: deploying a first mean time to failure (MTTF) machine learning model specific to a plurality of first memory devices based on characteristics of each one of the plurality of first memory devices … “ , and a method is one of the four statutory categories of invention.
In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components:
and predicting a respective MTTF for each of the plurality of first memory devices and the plurality of second memory devices … (mental process, a person can mentally evaluate and predict a MTTF for each of the first and second memory devices, see MPEP 2106.04(a)(2)(III))
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
A method, comprising: deploying a first mean time to failure (MTTF) machine learning model specific to a plurality of first memory devices based on characteristics of each one of the plurality of first memory devices … (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
deploying a second MTTF machine learning model specific to a plurality of second memory devices based on characteristics of each one of the plurality of second memory devices; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
aggregating output data received from the first MTTF machine learning model and the second MTTF machine learning model into a third MTTF machine learning model; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
aggregating output data received from the third MTTF machine learning model and a fourth MTTF machine learning model into a fifth MTTF machine learning model; (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
… based on output data of the fifth MTTF machine learning model, the output data received from the third machine learning model, the output data of the first MTTF machine learning model, the output data of the second MTTF machine learning model, or any combination thereof. (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)),
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional elements ii, iii, iv, v, and vi recite mere instructions to apply the judicial exception using generic computer components, which are not indicative of significantly more.
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 13:
Regarding claim 13, it is dependent upon claim 12, and thereby incorporates the limitations of, and corresponding analysis applied to claim 12. Further, claim 13 recites the following additional element:
The method of claim 12, further comprising providing the predicted MTTF to a host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 14:
Regarding claim 14, it is dependent upon claim 12, and thereby incorporates the limitations of, and corresponding analysis applied to claim 12. Further, claim 14 recites the following additional element:
The method of claim 13, further comprising providing an updated predicted MTTF to the host device each time one of the first, the second, the third, the fourth, or the fifth MMTF machine learning models is updated. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 15:
Regarding claim 15, it is dependent upon claim 12, and thereby incorporates the limitations of, and corresponding analysis applied to claim 12. Further, claim 15 recites the following additional element:
The method of claim 12, further comprising aggregating output data from a sixth MTTF machine learning model and a seventh MTTF machine learning model into the fourth MTTF machine learning model, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
wherein the sixth MTTF machine learning model is specific to a plurality of third memory devices, and the seventh MTTF machine learning model is specific to a plurality of fourth memory devices. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 16:
Regarding claim 16, it is dependent upon claim 12, and thereby incorporates the limitations of, and corresponding analysis applied to claim 12. Further, claim 16 recites the following additional element:
The method of claim 12, further comprising deploying the first MTTF machine learning model, the second MTTF machine learning model, or both, on memory devices of a second host device different than a first host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 17:
Regarding claim 17, it is dependent upon claim 12, and thereby incorporates the limitations of, and corresponding analysis applied to claim 12. Further, claim 17 recites the following additional element:
The method of claim 12, further comprising deploying the third MTTF machine learning model, the fourth MTTF machine learning model, or both on memory devices of a second host device different than a first host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 18:
Regarding claim 18, it is dependent upon claim 12, and thereby incorporates the limitations of, and corresponding analysis applied to claim 12. Further, claim 18 recites the following additional element:
The method of claim 12, further comprising deploying the fifth MTTF machine learning model on memory devices of a second host device different than a first host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 19:
Regarding claim 19, it is dependent upon claim 12, and thereby incorporates the limitations of, and corresponding analysis applied to claim 12. Further, claim 19 recites the following additional element:
The method of claim 12, further comprising updating the fifth MTTF machine learning model in response to a change in the third MTTF machine learning model, the fourth MTTF machine learning model, or both. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 20:
Regarding claim 20, it is dependent upon claim 12, and thereby incorporates the limitations of, and corresponding analysis applied to claim 12. Further, claim 20 recites the following additional element:
The method of claim 12, further comprising updating the third MTTF machine learning model in response to a change in the first MTTF machine learning model, the second MTTF machine learning model, or both. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
Claim 1 is rejected under 35 U.S.C. 103 over Liu, J. et al., Pub. No. CN115456194A, published on December 9, 2022, (hereafter, Liu J.), in view of Abdulrahman S., et al in "A Survey on Federated Learning: The Journey From Centralized to Distributed On-Site Learning and Beyond," published on April 1, 2021, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9220780 , (hereafter, Abdulrahman), further in view of Bemani, A. et al. in “Aggregation Strategy on Federated Machine Learning Algorithm for Collaborative Predictive Maintenance”, published on August 19, 2022, available on https://www.mdpi.com/1424-8220/22/16/6252, (hereafter, Bemani).
Claim 1:
Regarding claim 1, Liu J. teaches “1. A system, comprising: a first computing device configured to:…” See Liu J. in paragraphs [n0006- n0009] describe “[n0006] Send the first parameters of the first global model to multiple edge devices. The first global model is the initial global model. [n0007] Receive the second parameter of the second global model returned by the first edge device among multiple edge devices. The second global model is a global model obtained by the first edge device after training the first global model based on the local dataset,” Here, Liu J. describes using multiple edge devices, which shows more than one computing device.
Further, Liu J. teaches “ a second computing device configured to:”
See Liu J. in paragraphs [n0006- n0009] describe “[n0006] Send the first parameters of the first global model to multiple edge devices. The first global model is the initial global model. [n0007] Receive the second parameter of the second global model returned by the first edge device among multiple edge devices. The second global model is a global model obtained by the first edge device after training the first global model based on the local dataset. [n0008] When a third global model is obtained by aggregating at least one second global model, the third parameters of the third global model are sent to the second edge devices among multiple edge devices. ... [n0009] According to a second aspect of this disclosure, a model training control method based on asynchronous federated learning, applied to a second edge device,” Here, Liu J. describes using multiple edge devices, which shows more than one computing device.
Further, Liu J. teaches “aggregate the first machine learning model and the second machine learning model into a third machine learning model;”
See Liu J. describe in paragraph [n0017] “Upon receiving the second parameters of the second global model returned by the first edge device among multiple edge devices, the base station determines the third global model, which is the latest global model relative to the first global model.” Here, Liu J. describes using a first model and second global model, and later combine or aggregate them to determine a third global model, showing this as a part of an iterative aggregation step. The base station combines the returned parameters with other updates to form the latest, improved global model.
Further, Liu J. teaches “aggregate the third machine learning model with a fourth machine learning model comprising a plurality of aggregated machine learning models into a fifth machine learning model;”
See Liu J. in [n0011-n0012] mention “in response to receiving the third parameters of the third global model sent by the base station, the fourth global model is determined. The third global model is the latest global model determined by the base station. [n0012] The third and fourth global models are aggregated to obtain the fifth global model;” Here, Liu J. mentions taking a third and fourth model and aggregate them into a fifth model.
However, Liu J. did not teach “gather memory usage data and device characteristic data associated with a first plurality of memory devices monitored by the first computing device” or “and train a first machine learning model based on the gathered memory usage data and the device characteristic data associated with the first plurality of memory devices;” or “gather memory usage data and device characteristic data associated with a second plurality of memory devices monitored by the second computing device;” or “and train a second machine learning model based on the gathered memory usage data and the device characteristic data associated with the second plurality of memory devices;” or “and a local federated server in communication with the first computing device and the second computing device and configured to:” or “and predict aging of the first plurality of memory devices and the second plurality of memory devices based on the fifth machine learning model,” or “a global federated server in communication with the local federated server and configured to”
In an analogous method, Abdulrahman teaches “gather memory usage data … associated with a first plurality of memory devices monitored by the first computing device;”
See Abdulrahman in page 5490 in SECTION VIII. Resource Management describe “FL is applied in dynamic environments, in which the clients have constrained resource devices and are communicating through bandwidth-constrained networks where some devices can share the same link. Therefore, many contributions have been focusing on resource management to take the best decision related to the selected clients, learning hyperparameters, number and duration of training rounds, and aggregation strategies. In this context, various optimization problems have been defined and solved assuming the availability/predictability of subsets of the following metrics (Fig. 8).
• Clients Reliability: Resources (CPU, energy), location traceability (GPS coordinates), local training time, and quality of updated parameters (accuracy, loss). Some related work assume the availability of the “actual” values of these metrics at each learning round while others adopt various approaches to “predict” these values.
• Network Link Quality: Uplink/downlink bandwidth either already available or possibly allocated.” Here from Abdulrahman in page 5490, the examiner construes the clients’ resource data such as CPU or energy usage, or location information as well as local training time are construed memory usage data. Here, Abdulrahman describes memory usage data associated with devices such as resource usage with CPU or energy.
Further, see Abdulrahman in page 5480 in section IV. FL Technical Challenges and Research Fields: New Classification describe "3) Massively Distributed Data: The participants in FL can form multiple millions of clients, ranging from mobile phones to IoT devices, organizations/institutions, vehicles, and many more. It has been reported in [3] that the number of the participants is expected to be larger than the average number of samples per participant." Here, Abdulrahman mentions that this method can be applied to multiple clients or devices, and includes a first, second, or other subsequent devices.
Further, Abdulrahman teaches “gather memory usage data … associated with a second plurality of memory devices monitored by the second computing device;”
See Abdulrahman in page 5490 in SECTION VIII. Resource Management describe “FL is applied in dynamic environments, in which the clients have constrained resource devices and are communicating through bandwidth-constrained networks where some devices can share the same link. Therefore, many contributions have been focusing on resource management to take the best decision related to the selected clients, learning hyperparameters, number and duration of training rounds, and aggregation strategies. In this context, various optimization problems have been defined and solved assuming the availability/predictability of subsets of the following metrics (Fig. 8).
• Clients Reliability: Resources (CPU, energy), location traceability (GPS coordinates), local training time, and quality of updated parameters (accuracy, loss). Some related work assume the availability of the “actual” values of these metrics at each learning round while others adopt various approaches to “predict” these values.
• Network Link Quality: Uplink/downlink bandwidth either already available or possibly allocated.” Here from Abdulrahman in page 5490, the examiner construes the clients’ resource data such as CPU or energy usage, or location information as well as local training time are construed memory usage data. Here, Abdulrahman describes memory usage data associated with devices such as resource usage with CPU or energy.
Further, see Abdulrahman in page 5480 in section IV. FL Technical Challenges and Research Fields: New Classification describe "3) Massively Distributed Data: The participants in FL can form multiple millions of clients, ranging from mobile phones to IoT devices, organizations/institutions, vehicles, and many more. It has been reported in [3] that the number of the participants is expected to be larger than the average number of samples per participant." Here, Abdulrahman mentions that this method can be applied to multiple clients or devices, and includes a first, second, or other subsequent devices.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Liu J. and incorporate into the teachings of Abdulrahman because both references teach generating a model that predicts aging for each device based on device characteristics and from memory usage attributes.
One of ordinary skill in the art would be motivated to do so because this would achieve a “goal of FL is to train and generate high-quality global model through the learning rounds. With high-dimensional data distributed on-devices, the following approaches have been proposed to efficiently federate client-provided model updates,” (see Abdulrahman in page 5482, section C. Optimization and Aggregation Algorithms, last paragraph on page).
Further, Abdulrahman teaches “and train a first machine learning model based on the gathered memory usage data …. associated with the first plurality of memory devices” and “and train a second learning model based on the gathered memory usage data …. associated with the second plurality of memory devices”
See Abdulrahman in page 5490 in SECTION VIII. Resource Management describe “FL is applied in dynamic environments, in which the clients have constrained resource devices and are communicating through bandwidth-constrained networks where some devices can share the same link. Therefore, many contributions have been focusing on resource management to take the best decision related to the selected clients, learning hyperparameters, number and duration of training rounds, and aggregation strategies. In this context, various optimization problems have been defined and solved assuming the availability/predictability of subsets of the following metrics (Fig. 8).” Abdulrahman mentions multiple clients, which relate to multiple devices.
See Abdulrahman in page 4729, in algorithm 1,
PNG
media_image1.png
839
722
media_image1.png
Greyscale
where Abdulrahman describes training each device using memory usage attributes such as CPU, memory or energy, etc.
However, Liu J. in view of Abdulrahman did not teach “gather … device characteristic data associated with a first plurality of memory devices monitored by the first computing device;” or “and train a first machine learning model based on the gathered … device characteristic data associated with the first plurality of memory devices;” or “and train a second machine learning model based on the gathered … device characteristic data associated with the second plurality of memory devices;” or “and a local federated server in communication with the first computing device and the second computing device and configured to:” or “and predict aging of the first plurality of memory devices and the second plurality of memory devices based on the fifth machine learning model,” or “a global federated server in communication with the local federated server and configured to…”
In an analogous art, Bemani teaches “gather … device characteristic data associated with a first plurality of memory devices monitored by the first computing device;”
See Bemani in page 12, section 5, Structure of Performance Evaluation, section 5.1. Distributing CMAPSS for a Collaborative PM describe "Each edge device can train the global model (FedSVM or FedLSTM) by using its local time-series data." Here, Bemani describes each device using its device characteristics data such as a device’s local time series data to train a model.
Further, see Bemani in page 7, section 3.1 Network model, describe "We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement." Here, Bemani mentions collecting information or data about measurements of devices of each asset or per device.
Further, see Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set N = {1, . . . , N}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani here describes multiple devices, indicated by the set N from 1 through N, where N can be any number, or relate to collecting data as well as respective features for first computing device or subsequent devices.
Further, Bemani teaches “and train a first machine learning model based on the gathered memory usage data and the device characteristic data associated with the first plurality of memory devices;”
See Bemani in page 12, section 5, Structure of Performance Evaluation, section 5.1. Distributing CMAPSS for a Collaborative PM describe "Each edge device can train the global model (FedSVM or FedLSTM) by using its local time-series data." Here, Bemani describes each device using its device characteristics data such as its local time series data to train a model, and this applies for each subsequent device such as the second, third device, etc.
See Bemani describe on page 9 for more details that each edge node trains the FedLSTM or FedSVM locally before aggregation in algorithm 1, based on device attributes or features of the data of first device group.
PNG
media_image2.png
973
1080
media_image2.png
Greyscale
Also, see also Bemani in page 7, section 3.1 Network model, describe "We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement." Here, Bemani mentions collecting information or data about measurements of devices of each asset or per device.
Further, see Bemani in page 3 , paragraph 3, mention “FL consists of two steps, local training and global aggregation. In local training, the edge device downloads the model from the fog and computes an updated model using its local data. The fog server then collects these updated models mainly by averaging. FL can be used between fog and cloud also.” Bemani mentions the training is based on the gathered data per device.
Further, Bemani teaches “gather … device characteristic data associated with a second plurality of memory devices monitored by the second computing device;”
See Bemani in page 12, section 5, Structure of Performance Evaluation, section 5.1. Distributing CMAPSS for a Collaborative PM describe "Each edge device can train the global model (FedSVM or FedLSTM) by using its local time-series data." Here, Bemani describes each device using its device characteristics data such as its local time series data, that it collected to train a model, and this applies for each subsequent device such as the second, third device, etc.
Also, see Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set N = {1, . . . , N}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani here describes multiple devices, indicated by the set N from 1 through N, where N can be 2, or relate to collecting data as well as respective features for a second computing system or device.
Further, see Bemani in page 7, section 3.1 Network model, describe "We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement." Here, Bemani mentions collecting information or data about measurements of devices of each asset or per device.
Further, see Bemani teach “ and train a second machine learning model based on the gathered … the device characteristic data associated with the second plurality of memory devices;”
See Bemani in page 12, section 5, Structure of Performance Evaluation, section 5.1. Distributing CMAPSS for a Collaborative PM describe "Each edge device can train the global model (FedSVM or FedLSTM) by using its local time-series data." Here, Bemani describes each device using its device characteristics data such as its local time series data to train a model, and this applies for each subsequent device such as the second, third device, etc.
Further, see Bemani in page 6, section 2.1. Machine Learning at the Edge, Fog, and Cloud Levels, describe "Different strategies for model aggregation and hierarchical FL are essential issues in a distributed system at the edge, fog, and cloud levels. Lumin Liu et al. proposed the hierarchical FL based on FedAvg for distributed systems [28]. They tried to formulate an optimization problem based on the number of aggregations at the fog and cloud levels compared to the number of iterations at the edge level. Based on their proposed architecture, this model can be trained faster and achieve better communication efficiency.” Here, Bemani describe training at each round or iteration, which relates to training not only a first model based on collected data, but also subsequent models such as a second model, third model, etc.
Also, see Bemani in page 7, section 3.1 Network model, describe "We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement." Here, Bemani mentions collecting information or data about measurements of devices of each asset or per device.
Further, see Bemani in page 3 , paragraph 3, mention “FL consists of two steps, local training and global aggregation. In local training, the edge device downloads the model from the fog and computes an updated model using its local data. The fog server then collects these updated models mainly by averaging. FL can be used between fog and cloud also.” Bemani mentions the training is based on the gathered data per device.
Further, see Bemani teach “vi. and a local federated server in communication with the first computing device and the second computing device and configured to:”
See Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set N = {1, . . . , N}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices. The edge devices could be some similar assets in a factory site, and communicate with the fog server for an FL task via a wireless link. We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement. This dataset is denoted as Di = ξi,l Di l=1 , which ξ represents the lth training sample at edge device i,∀i ∈ N. The whole dataset in each factory is the union of its edge devices datasets D = i∈N Di. We consider two machine learning models (SVM and LSTM) over this wireless network between fog server and edge devices in different factory sites.” See Bemani describe fog server as a local server, and this fog server communicates among N distributed edge devices.
Further, Bemani teaches “and predict aging of the first plurality of memory devices and the second plurality of memory devices based on the fifth machine learning model,”
See Bemani in page 7, section 3.1 Network model, describe "We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement." Here, Bemani mentions collecting information or data about measurements of devices of each asset or per device.
Also, see Bemani in page 12, section 5, Structure of Performance Evaluation, section 5.1. Distributing CMAPSS for a Collaborative PM mention “Each time series consists of 21 sensor observations, three operating settings, a trajectory-id, and a cycle count, which the RUL of an engine is estimated from it in terms of the number of operation cycles before the engine runs to failure. The goal is to predict the RUL based on the time series data by the model at the cloud level, which was trained by the federated aggregation from the fog servers and, subsequently, edge devices.” Bemani mentions a metric called remaining useful life or RUL, which measures an estimated amount of time or usage a hardware system, component, or asset is expected to operate before it becomes unusable or requires either repair or replacement. This measures aging of a system since aging requires time elapsed before a system needs to get repaired. This metric gets applied to all edge devices for FedLSTM or federated learning.
Further, see Bemani in page 6, section 2.1, third paragraph mention “Different strategies for model aggregation and hierarchical FL are essential issues in a distributed system at the edge, fog, and cloud levels. Lumin Liu et al. proposed the hierarchical FL based on FedAvg for distributed systems [28]. They tried to formulate an optimization problem based on the number of aggregations at the fog and cloud levels compared to the number of iterations at the edge level. Based on their proposed architecture, this model can be trained faster and achieve better communication efficiency. Another study proposed a hierarchical FL to minimize training loss and latency.” Here, Bemani mentions using hierarchical federated learning methods with a number of aggregations and iterations at the edge level or on devices. Bemani shows that the fifth machine learning model is a result of a second round of aggregation from previous models, and each model can be used to predict RUL or aging for models from edge devices mentioned in page 12.
Further, Bemani teaches “a global federated server in communication with the local federated server and configured to …”
See Bemani in page 6, section 2.1 mention “Different strategies for model aggregation and hierarchical FL are essential issues in a distributed system at the edge, fog, and cloud levels. Lumin Liu et al. proposed the hierarchical FL based on FedAvg for distributed systems [28]. They tried to formulate an optimization problem based on the number of aggregations at the fog and cloud levels compared to the number of iterations at the edge level. Based on their proposed architecture, this model can be trained faster and achieve better communication efficiency. Another study proposed a hierarchical FL to minimize training loss and latency by formulating an optimization problem of edge aggregation interval control and time allocation [29]." Here, Bemani mention that the distributed system is communicating at the edge or device, fog or local server, and cloud or global server levels, indicating all three parts talk with one another to train models.
Also, see Bemani in page 8, section 3.2 Federated Learning Process mention “Broadcasting from cloud server: Cloud server broadcasts the global model parameter wt cloud to all fog servers through a wireless link in the tth round." Here, Bemani shows that a cloud server sends global model parameters to all local servers.
Further, see Bemani in page 15, in section 6. Experimental Results and Discussion describe “All the edge devices, fog servers, and cloud servers work as virtual workers inside the python script and collaborate based on the proposed algorithms and two communication topologies.” Here, Bemani mentions using a cloud server (i.e. global federated server since this cloud server sends global model parameters) that collaborates and communicates with edge devices and fog servers, where fog servers are local federated servers. Examiner construes cloud level server to mean a global federated server that talks to a local server.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Liu J. and Abdulrahman, and incorporate with the teachings of Bemani by using the teachings of Liu J. and Abdulrahman of creating model for each device from device characteristics, with Bemani’s teaching of predicting aging.
One of ordinary skill in the art would be motivated to do so because by integrating Bemani’s framework into the methods of Liu J. and Abdulrahman, one with ordinary skill in the art would achieve the goal of providing “By comparison of the FedLSTM results with other prediction methods in Table 5, it can be confirmed that the performance of the proposed FedLSTM in PM applications has comparable efficiency to the conventional centralized approaches in terms of prediction accuracy,” (see page 20, section 6.2, FedLSTM results, first paragraph in Bemani).
Claim 2 is rejected under 35 U.S.C. 103 over Liu J. in view of Abdulrahman, further in view of Bemani , further in view of Yang, W. et al in “Optimizing Federated Learning With Deep Reinforcement Learning for Digital Twin Empowered Industrial IoT,” published on July 4, 2022, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9815106 , (hereafter, Yang).
Claim 2:
However, Liu J. in view of Abdulrahman, further in view of Bemani did not teach “The system of claim 1, wherein the first plurality of memory devices comprises a first categorized group of memory devices having first device characteristics different than a second categorized group of memory devices having second device characteristics of the second plurality of memory devices,”
In an analogous field, Yang teaches “The system of claim 1, wherein the first plurality of memory devices comprises a first categorized group of memory devices having first device characteristics different than a second categorized group of memory devices having second device characteristics of the second plurality of memory devices”
See Yang in page 1889 describe “C. DRL-Based Asynchronous Federated Learning
Devices in the IIoT application scenario are highly heterogeneous, and the slowest device will limit the training speed of the synchronous learning solution, causing the so-called straggler effect. Therefore, we propose an asynchronous FL framework to address this issue. The main idea is to select the optimally participating devices, classify the devices with different utility values by cluster, and configure the corresponding aggregator for each cluster to realize asynchronous learning. In this case, each cluster can be trained at different local aggregation frequencies. In addition, we adopt Algorithm 1, based on the actor-critic DRL framework, to select devices that participate in asynchronous FL. Our proposed asynchronous FL framework mainly includes the following four steps.
Device selection: For the sake of improving the convergence rate and the model accuracy, a device with higher utility is selected to participate in FL within the given communication time. In the beginning, the server initializes the FL process through broadcasting the global model and the initialization parameter wini. The server then selects the optimal subset of devices Ni∈N through the DRL-based algorithm.” Here, Yang shows that there are different groups of devices, which have different utility values (i.e. different device characteristics). Examiner construes memory devices to be any IoT device stated from the specification in [0018] “The memory device 150 and host 103 can be a satellite, a communications tower, a personal laptop computer, a desktop computer, a digital camera, a mobile telephone, a memory card reader, an Internet-of-Things (IOT) enabled device, an automobile, among various other types of systems. For clarity, the system 101 has been simplified to focus on features with particular relevance to the present disclosure. ”.
Also see Yang in pages 1890-1891 in section V. Experiments mention “we evaluate the proposed scheme’s performance on different numbers of training devices. To verify the developed device selection scheme’s performance, we established three training device groups, 30, 50, and 70, with 10 inefficient devices in each group. Then, the proposed device selection algorithm is evaluated on the three groups of training devices and compared with the first group of devices without using the device selection algorithm. Figs. 6 and 7 show the prediction accuracy and the loss of the training model, respectively. The experimental results demonstrate that the scheme has excellent convergence and accuracy. As the number of devices involved increases from 30 to 50 to 70, the model’s accuracy increases slightly. This is because a larger number of utility devices involved in the training results in a higher quality model.” Here, Yang shows three different groups of devices, which are considered at least 2 groups of devices used for federated learning.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Liu J., Abdulrahman, and Bemani, and incorporate with the teachings of Yang by using the teachings of Liu J., Abdulrahman, and Bemani of devices and each training a model to predict aging, with Yang’s teaching of two different groups of memory devices each with different device characteristics.
One of ordinary skill in the art would be motivated to do so because by integrating Yang’s framework into the methods of Liu J., Abdulrahman, and Bemani, one with ordinary skill in the art would achieve the goal of providing a “ proposed asynchronous framework with the device selection and clustering mechanism eliminates the straggler effect, effectively avoids inefficient devices and even malicious attacks, and improves the convergence rate and learning quality. Although the DRL-based method requires massive samples for training, it can improve the training effect and maintain the practical significance through addressing the problems like resource consumption and sample distribution,” (see Yang in page 1890, in section C. DRL-Based Asynchronous Federated Learning, fifth paragraph after fourth item).
Claims 3 and 4 are rejected under 35 U.S.C. 103 over Liu J., in view of Abdulrahman, further in view of Bemani, and further in view of Lo, S. et al. in “A systematic literature review on federated machine learning: From a software engineering perspective,” published on 25 May 2021 , available at https://dl.acm.org/doi/pdf/10.1145/3450288, (hereafter, Lo).
Claim 3:
Regarding claim 3, Liu J., in view of Abdulrahman, further in view of Bemani, teach the limitations in claim 1.
Further, Bemani teaches “3. The system of claim 1, wherein the first computing device is configured to predict aging of the first plurality of memory devices based on the first machine learning model, and the second computing device is configured to predict aging of the second plurality of memory devices based on the second machine learning model.”
See Bemani in page 7, section 3.1 describe “The edge devices could be some similar assets in a factory site, and communicate with the fog server for an FL task via a wireless link. We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement. This dataset is denoted as 𝒟𝑖={𝝃𝑖,𝑙}𝐷𝑖𝑙=1, which 𝝃 represents the 𝑙th training sample at edge device 𝑖,∀𝑖∈𝒩. The whole dataset in each factory is the union of its edge devices datasets 𝒟=⋃𝑖∈𝒩𝒟𝑖. We consider two machine learning models (SVM and LSTM) over this wireless network between fog server and edge devices in different factory sites. The fog server and edge devices collaboratively build shared SVM and LSTM models for predicting the labels. These shared models are trained by exchanging model parameter information while keeping all the data locally at the edge devices. The global shared model will be made by aggregating all the model parameters from fog servers.” Here, Bemani mentions that each factory site has their own group of edge devices, and each of those device groups use their own respective models.
Also, see Bemani in page 4, in section 1.2 describe “FedLSTM utilizes a federated long short term memory (LSTM) model in each edge device to predict the absolute values of an asset’s RUL. This method is useful for learning from sequence data in each edge device. By applying the moving average strategy, the number of consecutive blocks in FedLSTM is reduced compared to without it, which significantly affects the training time of the model at the fog level.” Bemani mentions a model is used for each device to predict RUL.
Also, see Bemani in page 12, section 5, Structure of Performance Evaluation, section 5.1. Distributing CMAPSS for a Collaborative PM mention “Each time series consists of 21 sensor observations, three operating settings, a trajectory-id, and a cycle count, which the RUL of an engine is estimated from it in terms of the number of operation cycles before the engine runs to failure. The goal is to predict the RUL based on the time series data by the model at the cloud level, which was trained by the federated aggregation from the fog servers and, subsequently, edge devices.” Bemani mentions a metric called remaining useful life or RUL, which measures an estimated amount of time or usage a hardware system, component, or asset is expected to operate before it becomes unusable or requires either repair or replacement. This measures aging of a system since aging requires time elapsed before a system needs to get repaired. This metric gets applied to all edge devices for FedLSTM or federated learning and is synonymous with predicting aging.
Further, see Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set N = {1, . . . , N}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani here describes multiple devices, indicated by the set N from 1 through N, where N can be any number, or relate to collecting data as well as respective features for first computing device or subsequent devices.
Claim 4:
Regarding claim 4, Liu J., in view of Abdulrahman, further in view of Bemani, teach the limitations in claim 1.
Further, Bemani teaches “The system of claim 1, wherein the local federated server is configured to predict aging …”
Also, see Bemani in page 4, in section 1.2 describe “FedLSTM utilizes a federated long short term memory (LSTM) model in each edge device to predict the absolute values of an asset’s RUL. This method is useful for learning from sequence data in each edge device. By applying the moving average strategy, the number of consecutive blocks in FedLSTM is reduced compared to without it, which significantly affects the training time of the model at the fog level.” Bemani mentions a model is used for each device to predict RUL.
Also, see Bemani in page 12, section 5, Structure of Performance Evaluation, section 5.1. Distributing CMAPSS for a Collaborative PM mention “Each time series consists of 21 sensor observations, three operating settings, a trajectory-id, and a cycle count, which the RUL of an engine is estimated from it in terms of the number of operation cycles before the engine runs to failure. The goal is to predict the RUL based on the time series data by the model at the cloud level, which was trained by the federated aggregation from the fog servers and, subsequently, edge devices.” Bemani mentions a metric called remaining useful life or RUL, which measures an estimated amount of time or usage a hardware system, component, or asset is expected to operate before it becomes unusable or requires either repair or replacement. This measures aging of a system since aging requires time elapsed before a system needs to get repaired. This metric gets applied to all edge devices for FedLSTM or federated learning and is synonymous with predicting aging. Further, see Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set N = {1, . . . , N}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani here describes multiple devices, indicated by the set N from 1 through N, where N can be any number, or relate to collecting data as well as respective features for first computing device or subsequent devices. Fog server is construed to mean a local federated server, which is connected to the system to predict RUL or predict aging.
However, Liu J., in view of Abdulrahman, further in view of Bemani, did not teach “4. The system of claim 1, wherein the local federated server is configured to predict aging of the first plurality of memory devices and the second plurality of memory devices based on the third machine learning model.”
In an analogous art, Lo teaches “4. The system of claim 1, wherein the local federated server is configured to predict aging of the first plurality of memory devices and the second plurality of memory devices based on the third machine learning model.”
See Lo in page 95:21, section 3.8.1 Central server, mention “After global model initialisation, the central servers broadcast the global model that includes the model parameters or gradients to the participating client devices. The global model can be broadcasted to all the participating client devices every round [5, 65, 142, 206], or only to specific client devices, either randomly [32, 88, 114, 159, 182] or through selection based on the model training performance [131, 179] and the resources availability [38, 168]. Similarly, the trained local models are also collected from either all the participating client devices [167, 168, 198]or only from selected client devices [18, 111]. The collection of models can either be in an asynchronous [29, 64, 99, 174], or synchronous manner [42, 48]. Finally, the central server performs model aggregations when it receives all or a specific amount of updates, followed by the redistribution of the updated global model to the client devices.” Here, Lo mentions that after the central server perform aggregation to a global model (i.e. third machine learning model), this information is then redistributed to the client devices. The global model is viewed as a combined or aggregated model form the local models, and examiner interprets this aggregated model to be part of the third model, which is combined from the first and second models mentioned from the specification abstract “The local federated server can aggregate the first machine learning model and the second machine learning model into a third machine learning model.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Liu J., Abdulrahman, and Bemani, and incorporate with the teachings of Lo by using the teachings of Liu J., Abdulrahman, and Bemani of devices and each training a model to predict aging, with Lo’s teaching of using two different models for each of the two devices, then aggregate them to a third model.
One of ordinary skill in the art would be motivated to do so because by integrating Yang’s framework into the methods of Liu J., Abdulrahman, and Bemani, one with ordinary skill in the art would achieve the goal of providing “the proposed approaches use control algorithms [168], reinforcement learning [118, 193], and edge computing methods [121] to optimize the resource usage and improve the system efficiency,” (see Lo in page 95:18, last paragraph).
Claims 5 and 6 are rejected under 35 U.S.C. 103 over Liu J., in view of Abdulrahman, further in view of Bemani, and further in view of Mills, J., et al. in “Multi-task federated learning for personalised deep neural networks in edge computing,” published on July 21, 2021, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9492755 , (hereafter, Mills).
Claim 5:
Regarding claim 5, Liu J., in view of Abdulrahman, further in view of Bemani, teach the limitations in claim 1. However, Liu J., in view of Abdulrahman, further in view of Bemani, did not teach “5. The system of claim 1, wherein the first machine learning model provides a more specific aging prediction of the first plurality of memory devices as compared to the third machine learning model; and wherein the second machine learning model provides a more specific aging prediction of the second plurality of memory devices as compared to the third machine learning model.”
In an analogous field, Mills teaches “5. The system of claim 1, wherein the first machine learning model provides a more specific aging prediction of the first plurality of memory devices as compared to the third machine learning model; and wherein the second machine learning model provides a more specific aging prediction of the second plurality of memory devices as compared to the third machine learning model.”
See Mills in page 5 in section 3.2 Effect of BN Patches on Inference mention "To understand the impact that BN-patch layers have on UA, we consider the change in internal DNN activations over a client’s local test-set immediately before and immediately after the FL aggregation step. As illustrated in Fig. 3 (a), UA typically drops after the aggregation step in iterative FL. This is because the model has been tuned on the local training set for several epochs, and suddenly has its model weights replaced by the Federated weights, which are unlikely to have better test performance than the preaggregation model. This idea is further examined in [32] and showed later in our experimental section." Here, Mills shows that the user model accuracy (UA), which is considered a metric used for the prediction of local client devices or local models. Since Mills mentions the model accuracy drops after aggregation, and the pre-aggregated local model performance (i.e. includes a first model) has better performance than the post-aggregation model (i.e. third model). The term ‘more specific’ is a relative term. Since more specific is interpreted to mean a metric or value that provides better performance in the pre-aggregated model or performs well compared to post- aggregated model, Mills shows that the model accuracy is better for the pre-aggregated model (which in this case applies to a second machine learning model) than the post-aggregated model, which is the third model. See Mills in section 3. Multi-Task Federated Learning (MTFL) mention for more details
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Liu J., Abdulrahman, and Bemani, and incorporate with the teachings of Mills by using the teachings of Liu J., Abdulrahman, and Bemani of devices and each training a model to predict aging, with the teaching of Mills of providing a more specific prediction by a local model than a global model.
One of ordinary skill in the art would be motivated to do so because by integrating the framework of Mills into the methods of Liu J., Abdulrahman, and Bemani, one with ordinary skill in the art would achieve the goal of providing “Experiments using MNIST and CIFAR10 demonstrate that MTFL is able to significantly reduce the number of rounds required to reach a target UA, by up to 5× when using existing FL optimisation strategies, and with a further 3× improvement when using FedAvg-Adam. We compare MTFL to competing personalised FL algorithms, showing that it is able to achieve the best UA for MNIST and CIFAR10 in all considered scenarios,” (see Mills in page 1, abstract).
Claim 6:
Regarding claim 6, Liu J., in view of Abdulrahman, further in view of Bemani, teach the limitations in claim 1. However, Liu J., in view of Abdulrahman, further in view of Bemani, did not teach “6. The system of claim 1, wherein the first machine learning model provides a more specific aging prediction of the first plurality of memory devices; and wherein the second machine learning model provides a more specific aging prediction of the second plurality of memory devices as compared to the fifth machine learning model.”
Further, Mills teaches “6. The system of claim 1, wherein the first machine learning model provides a more specific aging prediction of the first plurality of memory devices; and
wherein the second machine learning model provides a more specific aging prediction of the second plurality of memory devices as compared to the fifth machine learning model.”
See Mills from page 5 in section 3.2 Effect of BN Patches on Inference mention “To understand the impact that BN-patch layers have on UA, we consider the change in internal DNN activations over a client’s local test-set immediately before and immediately after the FL aggregation step. As illustrated in Fig. 3 (a), UA typically drops after the aggregation step in iterative FL. This is because the model has been tuned on the local training set for several epochs, and suddenly has its model weights replaced by the Federated weights, which are unlikely to have better test performance than the preaggregation model. This idea is further examined in [32] and showed later in our experimental section.” Also see Mills in page 1 in abstract mention “MTFL benefits UA and convergence speed by allowing users to train models personalised to their own data.” Here, Mills shows that the user model accuracy (UA), which is considered a metric used for the prediction of local client devices or local models. Since Mills mentions the model accuracy drops after aggregation, and the pre-aggregated model performance (i.e. includes a second model from the training of more than one model from page 1, abstract) has better performance than the post-aggregation model (i.e. fifth model or a model that is aggregated twice). The term ‘more specific’ is a relative term. Since more specific is interpreted to mean a metric or value that provides better performance in the pre-aggregated model or performs well compared to post- aggregated model, Mills shows that the model accuracy is better for the pre-aggregated model (which in this case applies to a second machine learning model) than the post-aggregated model, which is the fifth model.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Liu J., Abdulrahman, and Bemani, and incorporate with the teachings of Mills by using the teachings of Liu J., Abdulrahman, and Bemani of devices and each training a model to predict aging, with the teaching of Mills of providing a more specific prediction by a local model than a global model.
One of ordinary skill in the art would be motivated to do so because by integrating the framework of Mills into the methods of Liu J., Abdulrahman, and Bemani, one with ordinary skill in the art would achieve the goal of providing “Experiments using MNIST and CIFAR10 demonstrate that MTFL is able to significantly reduce the number of rounds required to reach a target UA, by up to 5× when using existing FL optimisation strategies, and with a further 3× improvement when using FedAvg-Adam. We compare MTFL to competing personalised FL algorithms, showing that it is able to achieve the best UA for MNIST and CIFAR10 in all considered scenarios,” (see Mills in page 1, abstract).
Claim 7 is rejected under 35 U.S.C. 103 over Liu J., in view of Abdulrahman, further in view of Bemani, and further in view of Wang Y. in “Performance Enhancement Schemes and Effective Incentives for Federated Learning,” published on November 16, 2021, available at https://ruor.uottawa.ca/items/fb7bb89a-0b36-4116-99c0-cd03e0c6517e , (hereafter, Wang).
Claim 7:
Regarding claim 7, Liu J., in view of Abdulrahman, further in view of Bemani, teach the limitations in claim 1. However, Liu J., in view of Abdulrahman, further in view of Bemani, did not teach “7. The system of claim 1, wherein the third machine learning model provides a more specific aging prediction of the first plurality of memory devices and the second plurality of memory devices as compared to the fifth machine learning model.”
In an analogous art, Wang teaches “7. The system of claim 1, wherein the third machine learning model provides a more specific aging prediction of the first plurality of memory devices and the second plurality of memory devices as compared to the fifth machine learning model.”
See Wang in page 27, section 3.3, describe “ By utilizing a temporary global model we asses[s] the possible aggregated performance of all local models for a specific iteration. " Here, Wang mentions an iterative aggregation of models, which relates to a first round of aggregation (i.e. creates a third model), and a second round of aggregation creates a fifth machine learning model. Examiner construes third model to be a model created from a first round of aggregation, and examiner construes this third model to be more local than the fifth model (created from second round of aggregation).
Later, see Wang in page 48, section 5.2.1 Reputation score, mention "(2) Comparison of Ai with the temporary global model’s accuracy Agtemp that is generated from the trained local models’ aggregation in this specific iteration is denoted as RT i . A positive contribution means that the local models outperform the temporary global model. (3) Comparison of Ai with the global model accuracy Agold of the last iteration is denoted as RP i . This would normally have a positive score because the global model from the last iteration improves with further training. The purpose of the reputation score is to choose reliable local models to participate in the updating process.” Here, Wang shows that one of the scores show that local models (such as third model) perform better than the temporary global or aggregated model, and is measured by model accuracy. Here, the third model is viewed as more local than the fifth model, since the third model only required one round of model aggregation as specified by the abstract in the specification state "The local federated server can aggregate the first machine learning model and the second machine learning model into a third machine learning model. The global federated server can aggregate the third machine learning model with a fourth machine learning model comprising a plurality of aggregated machine learning models into a fifth machine learning model". Note, the examiner construes more specific to mean having a performance metric such as accuracy, precision, or other similar measure that provides an improved model performance, such as high accuracy, since more is a relative term, having a better accuracy of local models outperform the temporary global model, also relates to more specific.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Liu J., Abdulrahman, and Bemani, and incorporate with Wang’s teachings by using the teachings of Liu J., Abdulrahman, and Bemani of devices and each training a model to predict aging, with the teaching of Wang of providing a more specific prediction by a third model than a fifth model.
One of ordinary skill in the art would be motivated to do so because by integrating Wang’s framework into the methods of Liu J., Abdulrahman, and Bemani, one with ordinary skill in the art would achieve “experimental results show that the proposed reputation-aware FL scheme can achieve improvements in test accuracy varying between 1.73% to 9.30% under different data sets,” (see Wang in page 5, section 1.2 Contributions).
Claims 8, 9, and 10 are rejected under 35 U.S.C. 103 over Shamshiri, A. et al. in “ML-based aging monitoring and lifetime prediction of IoT devices with cost-effective embedded tags for edge and cloud operability”, published on September 28, 2021 , available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9551269&tag=1 , (hereafter, Shamshiri), in view of Bemani, and further in view of Abdulrahman.
Claim 8:
Regarding claim 8, Shamshiri teaches “A system, comprising: a first plurality of memory devices grouped together based on memory device characteristics of the first plurality of memory devices;”
See Shamshiri in page 7433, Introduction, third paragraph of section describe “Therefore, it is necessary to have a system that automatically and intelligently notifies owners on the status and predicts potential failure. Moreover, this prediction should not create too many overheads (e.g., area) in circuits and their limited power, nor cause system performance degradation due to computation costs.” Here, Shamshiri describes a system that predicts potential failure of devices.
Further, see Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. These groups relate to memory device characteristics associated with a first group of memory devices.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to an initial or first machine learning model.
Further, see Shamshiri in page 7438, section B. mention “There is no soft limit on the available storage for cloud-oriented databases (e.g., distributed, federated, and clustered).” Here, Shamshiri mentions that the method involves various types of learning such as distributed, federated learning.
Further, Shamshiri teaches “ a second plurality of memory devices, grouped together based on memory device characteristics of the second plurality of memory devices;”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. These groups relate to memory device characteristics associated with two groups of memory devices, and relate to a first and second group of devices.
Further, Shamshiri teaches “determine a first mean time to failure (MTTF) for each one of the first plurality of memory devices based on the first local machine learning model;”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices.
See Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to an initial or first machine learning model.
Further, Shamshiri teaches “determine a second MTTF for each one of the second plurality of memory devices based on the second local machine learning model;”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices.
See Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting a MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to a machine learning model for the second group of devices mentioned from pages 7434-7435.
Further, Shamshiri teaches “… into a first MTTF machine learning model;”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices for their respective models.
Also, see Shamshiri in page 7433, in abstract describe “The proposed framework is implemented in three distinct types of cloud machine learning (ML), edge ML, and a combination of proportion estimation and Zscore.” Here, Shamshiri mentions implementing or deploying the ML models.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to an initial or first machine learning model that predicts MTTF.
See Shamshiri in algorithm 1 in page 7441 for more details.
PNG
media_image3.png
767
546
media_image3.png
Greyscale
Further, Shamshiri teaches “and determine a respective MTTF for each of the first plurality of memory devices and the second plurality of memory devices based on the first local machine learning model, the second local machine learning model, the first generic machine learning model, the second generic machine learning model, the global MTTF machine learning model, or any combination thereof.”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices.
See Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to an initial or first machine learning model.
Further, see Shamshiri in page 7438, section B. mention “There is no soft limit on the available storage for cloud-oriented databases (e.g., distributed, federated, and clustered).” Here, Shamshiri mentions that the method involves various types of learning such as distributed, federated learning.
Further, Shamshiri teaches “… into a global MTTF machine learning model”
See Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting a MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to a machine learning model for the second group of devices mentioned from pages 7434-7435.
However, Shamshiri did not teach “a first computing device configured to: gather memory usage data and the memory device characteristics from the first plurality of memory devices;” or “train a first local machine learning model based on the gathered memory usage data and the memory device characteristics for the first plurality of memory devices;” or “a second computing device configured to: gather memory usage data and the memory device characteristics from the second plurality of memory devices”, or “train a second local machine learning model based on the gathered memory usage data and the memory device characteristics for the second plurality of memory devices” or “a local federated server configured to: aggregate the first and the second local machine learning models into a first … machine learning model; and a global federated server configured to: aggregate weights derived from the first generic machine learning model with weights derived from a second generic machine learning model into a global … machine learning model;”
In an analogous art, Bemani teaches “ a first computing device configured to: …”
See Bemani in page 7, section 3.1 Network model, describe "We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement … we consider two machine learning models (SVM and LSTM) over this wireless network between fog server and edge devices in different factory sites.” When Bemani mentions devices in different factory sites, Bemani refers to using different computing devices. Here, Bemani also mentions collecting information or data about measurements of devices of each asset or per device.
Further, Bemani teaches “train a first local machine learning model based on the gathered … device characteristics for the first plurality of memory devices;”
See Bemani in page 12, section 5, Structure of Performance Evaluation, section 5.1. Distributing CMAPSS for a Collaborative PM describe "Each edge device can train the global model (FedSVM or FedLSTM) by using its local time-series data." Here, Bemani describes each device using its device characteristics data such as its local time series data, that it collected to train a model, and this applies for each subsequent device such as the second, third device, etc.
Further, see Bemani in page 6, section 2.1 mention “Different strategies for model aggregation and hierarchical FL are essential issues in a distributed system at the edge, fog, and cloud levels. Lumin Liu et al. proposed the hierarchical FL based on FedAvg for distributed systems [28]. They tried to formulate an optimization problem based on the number of aggregations at the fog and cloud levels compared to the number of iterations at the edge level. Based on their proposed architecture, this model can be trained faster and achieve better communication efficiency. Another study proposed a hierarchical FL to minimize training loss and latency by formulating an optimization problem of edge aggregation interval control and time allocation [29]." Here, Bemani mention that the distributed system is communicating at the edge or device, fog or local server, and cloud or global server levels, indicating all three parts talk with one another to train models.
Further, see Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set N = {1, . . . , N}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani here describes multiple devices, indicated by the set N from 1 through N, where N can be 1, or relate to collecting data as well as respective features for a first device. Bemani here also describes multiple devices, indicated by the set N from 1 through N, where N can be 2, 3 or subsequent numbers, and also relate to collecting data as well as respective features for subsequent devices.
Further, see Bemani in page 21, section 6.3 Model aggregation analysis, describe “Therefore we have considered the following two distribution cases and distributed the training data regarding labels 1 and 7 between the edge devices.
Edge-iid: The training data from labels 1 and 7 are identically distributed between the ten edge devices.
Edge-non-iid: The training data of label 1 are distributed among edge numbers 1 to 5 under one fog server, and training data of label 7 are distributed among edge numbers 6 to 10 under another fog server.” Bemani mentions that training is done from the data the edge devices collected from page 7, where the data includes device characteristics.
Also, see Bemani in page 7, section 3.1 Network model, describe "We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement." Here, Bemani mentions collecting information or data about measurements of devices of each asset or per device.
Further, see Bemani in page 3 , paragraph 3, mention “FL consists of two steps, local training and global aggregation. In local training, the edge device downloads the model from the fog and computes an updated model using its local data. The fog server then collects these updated models mainly by averaging. FL can be used between fog and cloud also.” Bemani mentions the training is based on the gathered data per device. Local training means training each model for each device.
Further, Bemani teaches “a second computing device configured to:…”
See Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set N = {1, . . . , N}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani here describes multiple devices, indicated by the set N from 1 through N, where N can be 2, or relate to collecting data as well as respective features for a second device. Bemani here also describes multiple devices, indicated by the set N from 1 through N, where N can be 2, or relate to collecting data as well as respective features for a second device.
Further, Bemani teaches “train a second local machine learning model based on the gathered … device characteristics for the second plurality of memory devices;”
See Bemani in page 12, section 5, Structure of Performance Evaluation, section 5.1. Distributing CMAPSS for a Collaborative PM describe "Each edge device can train the global model (FedSVM or FedLSTM) by using its local time-series data." Here, Bemani describes each device using its device characteristics data such as its local time series data, that it collected to train a model, and this applies for each subsequent device such as the second, third device, etc.
See Bemani in page 6, section 2.1 mention “Different strategies for model aggregation and hierarchical FL are essential issues in a distributed system at the edge, fog, and cloud levels. Lumin Liu et al. proposed the hierarchical FL based on FedAvg for distributed systems [28]. They tried to formulate an optimization problem based on the number of aggregations at the fog and cloud levels compared to the number of iterations at the edge level. Based on their proposed architecture, this model can be trained faster and achieve better communication efficiency. Another study proposed a hierarchical FL to minimize training loss and latency by formulating an optimization problem of edge aggregation interval control and time allocation [29]." Here, Bemani mention that the distributed system is communicating at the edge or device, fog or local server, and cloud or global server levels, indicating all three parts talk with one another to train models.
Further, see Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set N = {1, . . . , N}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani here describes multiple devices, indicated by the set N from 1 through N, where N can be 2, or relate to collecting data as well as respective features for a second device. Bemani here also describes multiple devices, indicated by the set N from 1 through N, where N can be 2, or relate to collecting data as well as respective features for a second device.
Further, see Bemani in page 21, section 6.3 Model aggregation analysis, describe “Therefore we have considered the following two distribution cases and distributed the training data regarding labels 1 and 7 between the edge devices.
Edge-iid: The training data from labels 1 and 7 are identically distributed between the ten edge devices.
Edge-non-iid: The training data of label 1 are distributed among edge numbers 1 to 5 under one fog server, and training data of label 7 are distributed among edge numbers 6 to 10 under another fog server.” Bemani mentions that training is done from the data the edge devices collected from page 7, where the data includes device characteristics.
Further, Bemani teaches “a local federated server configured to: …”
See Bemani in page 6, section 2.1. Machine Learning at the Edge, Fog, and Cloud Levels, describe "Different strategies for model aggregation and hierarchical FL are essential issues in a distributed system at the edge, fog, and cloud levels. Lumin Liu et al. proposed the hierarchical FL based on FedAvg for distributed systems [28]. They tried to formulate an optimization problem based on the number of aggregations at the fog and cloud levels compared to the number of iterations at the edge level. Based on their proposed architecture, this model can be trained faster and achieve better communication efficiency. Another study proposed a hierarchical FL to minimize training loss and latency by formulating an optimization problem of edge aggregation interval control and time allocation [29]." Here, Bemani mention that the distributed system is communicating at the edge or device, fog or local server, and cloud or global server levels, indicating all three parts talk with one another to train model for distributed federated learning.
Also, see Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set 𝒩={1,…,𝑁}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani mentions here that the fog server provide communication and computation services to its edge devices. The fog server is viewed here as a local server since the fog server connects to its local devices to get information from those devices.
Further, Bemani teaches “aggregate the first and the second local machine learning models into a first … machine learning model;”
See Bemani in page 6, section 2.1, Machine Learning at the Edge, Fog, and Cloud Levels, describe "Different strategies for model aggregation and hierarchical FL are essential issues in a distributed system at the edge, fog, and cloud levels. Lumin Liu et al. proposed the hierarchical FL based on FedAvg for distributed systems [28]. They tried to formulate an optimization problem based on the number of aggregations at the fog and cloud levels compared to the number of iterations at the edge level. Based on their proposed architecture, this model can be trained faster and achieve better communication efficiency. Another study proposed a hierarchical FL to minimize training loss and latency by formulating an optimization problem of edge aggregation interval control and time allocation [29]." Here, Bemani mention that the distributed system is communicating at the edge or device, fog or local server, and cloud or global server levels, indicating all three parts talk with one another to train model for distributed federated learning. Bemani mentions model aggregation, which is also done iteratively as part of hierarchical federated learning, which relates to aggregate the first and the second local models into a combined model (i.e. first combined machine learning model).
Further, Bemani teaches “and a global federated server configured to: …”
See Bemani in page 6, section 2.1. Machine Learning at the Edge, Fog, and Cloud Levels, describe "Different strategies for model aggregation and hierarchical FL are essential issues in a distributed system at the edge, fog, and cloud levels. Lumin Liu et al. proposed the hierarchical FL based on FedAvg for distributed systems [28]. They tried to formulate an optimization problem based on the number of aggregations at the fog and cloud levels compared to the number of iterations at the edge level. Based on their proposed architecture, this model can be trained faster and achieve better communication efficiency. Another study proposed a hierarchical FL to minimize training loss and latency by formulating an optimization problem of edge aggregation interval control and time allocation [29]." Here, Bemani mention that the distributed system is communicating at the edge or device, fog or local server, and cloud or global server levels (relate to global federated server), indicating all three parts talk with one another to train model for distributed federated learning.
Also, see Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set 𝒩={1,…,𝑁}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani mentions here that the cloud server (viewed as a global server) talks with the fog server (local server). Since the cloud server involves a virtual server that operates in a cloud computing environment and can be accessed on demand by multiple users, this can be accessed globally and is considered a global server.
Further, Bemani teaches “aggregate weights derived from the first generic machine learning model with weights derived from a second generic machine learning model into a global MTTF machine learning model;”
See Bemani in pages 6-7 in section 2.2 DML for collaborative PM, describe “One principle in predictive maintenance is anomaly detection, and some research has been done to deploy FL algorithms, mainly in this field. For instance, the authors in [34] introduced a novel FL algorithm for the LSTM framework and evaluated it on anomaly detection of sensors behavior in smart building applications. They proposed an FSLSTM network which consists of a local LSTM model that runs on sensor edge devices and a global model on fog that aggregates the weights, updates the parameters, and distributes them between the sensor edge devices. Their results show that this method converges twice as fast as the centralized LSTM model in the training phase. Similar work on anomaly detection is [35] in which authors proposed a novel communication efficient FL algorithm for sensing time-series data in distributed anomaly detection applications. They offered an attention mechanism-based convolutional neural network long short-term memory (AMCNN-LSTM) model to detect anomalies accurately. This model captures the most important features with the use of a CNN and then passes data to an LSTM, which predicts the future time-series data.” Here, Bemani shows the process of aggregating weights from a first LSTM model (i.e. first generic machine learning model) with weights from a second model on fog (i.e. second generic machine learning model) into a global model.
Further, see Bemani in pages 8-9, describe the steps “The training process at the 𝑡th communication round is described as follows:
Broadcasting from cloud server: Cloud server broadcasts the global model parameter 𝑤𝑡𝑐𝑙𝑜𝑢𝑑 to all fog servers through a wireless link in the 𝑡th round.
Fog updating phase: All fog servers update their model parameters with the received parameters from the cloud.
Broadcasting from local fog server: Fog servers broadcast the updated model parameters 𝑤𝑡𝑓𝑜𝑔,𝑗 to all edge devices located at factory site number j through a wireless link.
Edge device updating phase: After receiving the fog level model parameter, each edge device 𝑖∈𝒩 in factory 𝑗∈ℳ trains its local model by applying E epochs of a kind of optimization algorithms such as SGD and Adam. The iteration for SGD becomes
𝒘𝑡+1𝑖,𝑗=𝒘𝑡𝑖,𝑗−𝜂∇𝑓𝑗𝑖(𝒘𝑡𝑖,𝑗), (3)
where 𝜂 is the learning rate and ∇𝑓𝑗𝑖(𝒘𝑡𝑖,𝑗) is the stochastic gradient of local loss function in edge device i in factory j. After E epochs, edge device i uploads its update model parameter 𝒘𝑡+1𝑖,𝑗 to the connected fog server j.
Aggregating on fog: After E iterations on the edge devices in each factory, once receiving all the local model parameters, the fog server aggregates them and obtains an updated model, which is known as a synchronous method.
𝒘𝑡+1𝑓𝑜𝑔,𝑗=∑𝑖=1𝑁𝑗𝐷𝑗𝑖𝐷𝑗𝒘𝑡+1𝑖,𝑗, (4)
where 𝐷𝑗 is the whole data sample at factory site j. Another aggregation method is that when an edge device updates its model parameter in fog, the server in fog immediately creates an intermediate form of that agent’s parameters with its parameters and returns it to the agent for the next iteration. In this method, defined as asynchronous aggregation, the servers in the fog do not need to wait until they receive all the agents’ parameters.
Aggregating on cloud: The model parameter aggregation on the cloud happens once in a while. The number of communication rounds between cloud and fogs is much lower than the number of communication rounds between edge devices and fogs (𝑇𝐺<<𝑇𝑗𝑙). Therefore, when the cloud requests an update, a simple averaging with different weights (𝐴𝑗) depending on the size of the factory is performed on all fog parameters” Here, Bemani shows the process of aggregating model parameters.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reference of Shamshiri, and incorporate with the teachings of Bemani by using the teachings of Shamshiri of creating a model per device to predict MTTF for each device, with Bemani’s teaching of prediction of models based on device characteristics data, and model aggregation methods.
One of ordinary skill in the art would be motivated to do so because by integrating Bemani’s framework into the methods of Shamshiri, one with ordinary skill in the art would achieve the goal of providing “By comparison of the FedLSTM results with other prediction methods in Table 5, it can be confirmed that the performance of the proposed FedLSTM in PM applications has comparable efficiency to the conventional centralized approaches in terms of prediction accuracy,” (see page 20, section 6.2, FedLSTM results, first paragraph in Bemani).
However, Shamshiri in view of Bemani did not teach “a first computing device configured to: gather memory usage data and the memory … characteristics from the first plurality of memory devices;” or “a second computing device configured to: gather memory usage data and the memory … characteristics from the second plurality of memory devices”,
In an analogous art, Abdulrahman teaches “gather memory usage data and the memory … characteristics from the first plurality of memory devices;
See Abdulrahman in page 5490 in SECTION VIII. Resource Management describe “FL is applied in dynamic environments, in which the clients have constrained resource devices and are communicating through bandwidth-constrained networks where some devices can share the same link. Therefore, many contributions have been focusing on resource management to take the best decision related to the selected clients, learning hyperparameters, number and duration of training rounds, and aggregation strategies. In this context, various optimization problems have been defined and solved assuming the availability/predictability of subsets of the following metrics (Fig. 8).
• Clients Reliability: Resources (CPU, energy), location traceability (GPS coordinates), local training time, and quality of updated parameters (accuracy, loss). Some related work assume the availability of the “actual” values of these metrics at each learning round while others adopt various approaches to “predict” these values.
• Network Link Quality: Uplink/downlink bandwidth either already available or possibly allocated.” Here from Abdulrahman in page 5490, the examiner construes the clients’ resource data such as CPU or energy usage, or location information as well as local training time are construed memory usage data and related memory characteristics related to the device. Here, Abdulrahman describes memory usage data associated with the memory characteristics of the devices such as resource usage with CPU or energy, and can be associated with a first device.
Further, Abdulrahman teaches “ gather memory usage data and the memory … characteristics from the second plurality of memory devices;”
See Abdulrahman in page 5490 in SECTION VIII. Resource Management describe “FL is applied in dynamic environments, in which the clients have constrained resource devices and are communicating through bandwidth-constrained networks where some devices can share the same link. Therefore, many contributions have been focusing on resource management to take the best decision related to the selected clients, learning hyperparameters, number and duration of training rounds, and aggregation strategies. In this context, various optimization problems have been defined and solved assuming the availability/predictability of subsets of the following metrics (Fig. 8).
• Clients Reliability: Resources (CPU, energy), location traceability (GPS coordinates), local training time, and quality of updated parameters (accuracy, loss). Some related work assume the availability of the “actual” values of these metrics at each learning round while others adopt various approaches to “predict” these values.
• Network Link Quality: Uplink/downlink bandwidth either already available or possibly allocated.” Here from Abdulrahman in page 5490, the examiner construes the clients’ resource data such as CPU or energy usage, or location information as well as local training time are construed memory usage data and related memory characteristics related to the device. Here, Abdulrahman describes memory usage data associated with devices such as resource usage with CPU or energy. Abdulrahman mentions multiple clients, which relate to multiple devices.
Further, see Abdulrahman in page 5480 in section IV. FL Technical Challenges and Research Fields: New Classification describe "3) Massively Distributed Data: The participants in FL can form multiple millions of clients, ranging from mobile phones to IoT devices, organizations/institutions, vehicles, and many more. It has been reported in [3] that the number of the participants is expected to be larger than the average number of samples per participant." Here, Abdulrahman mentions that this method can be applied to multiple clients or devices, and includes a first, second, or other subsequent devices. This corresponds to data associated with a second device.
Further, Abdulrahman teaches “and train a first machine learning model based on the gathered memory usage data and the memory ….characteristics for the first plurality of memory devices” and “and train a second learning model based on the gathered memory usage data and the memory ….characteristics for the second plurality of memory devices”
See Abdulrahman in page 5490 in SECTION VIII. Resource Management describe “FL is applied in dynamic environments, in which the clients have constrained resource devices and are communicating through bandwidth-constrained networks where some devices can share the same link. Therefore, many contributions have been focusing on resource management to take the best decision related to the selected clients, learning hyperparameters, number and duration of training rounds, and aggregation strategies. In this context, various optimization problems have been defined and solved assuming the availability/predictability of subsets of the following metrics (Fig. 8).” Abdulrahman mentions multiple clients, which relate to multiple devices.
See Abdulrahman in page 4729, in algorithm 1, where Abdulrahman describes training each device using memory usage attributes such as CPU, memory or energy, etc.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Shamshiri and Bemani, and incorporate into the teachings of Abdulrahman because the references teach generating a model that predicts a respective MTTF for each device based on memory usage attributes.
One of ordinary skill in the art would be motivated to do so because this would achieve a “goal of FL is to train and generate high-quality global model through the learning rounds. With high-dimensional data distributed on-devices, the following approaches have been proposed to efficiently federate client-provided model updates,” (see Abdulrahman in page 5482, section C. Optimization and Aggregation Algorithms, last paragraph on page).
Claim 9:
Regarding claim 9, Shamshiri in view of Bemani, further in view of Abdulrahman, teach the limitations of claim 8.
Further, Shamshiri teaches “9. The system of claim 8, wherein the global federated server, the local federated server, the first computing device, or a combination thereof is configured to: determine one of the first plurality of memory devices has reached an MTTF;”
See Shamshiri in page 7433, Introduction, paragraph 3, describe "it is necessary to have a system that automatically and intelligently notifies owners on the status and predicts potential failure. Moreover, this prediction should not create too many overheads (e.g., area) in circuits and their limited power, nor cause system performance degradation due to computation costs." Further, see Shamshiri in page Here, Shamshiri mentions that once a system detects degradation (with MTTF), then this notifies owners prior to potential failure of that system.
Further, see Shamshiri in page 7434, section II. Aging Monitoring Tags note "The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices. " Shamshiri here uses MTTF as a metric, to determine potential outage of a group of device.
Further, Shamshiri teaches “and based on determination, provide an alert to a host device of the plurality of first memory devices”
See Shamshiri in page 7439, section IV. Machine Learning-Based Anomaly Detection And Aging Estimation, part A. Failure Prediction and Aging Monitoring Recent Works, third paragraph of section describe "In intrachip detection techniques, each device indicates failures internally and dispatches alarm. It is a conventional method for designing integrated circuits,... The main advantage of this technique is its automatic detection of a failed device." Also, see Shamshiri in page 7439, section B. Anomaly Detection Methods mention “ Devices in IoT systems are usually center oriented. That means all devices are only connected to the central cloud.” Here, Shamshiri describes that each devices dispatches an alarm, which relates to providing an alert to a central cloud (i.e. global federated server). Shamshiri describes an alarm to indicate an alert provided by the device to signal that the device is not working. The examiner construes 'the global federated server, the local federated server, the first computing device, or a combination' to mean either one of these systems, and not all of them are used for running the method the claim requires.
Further, see Shamshiri in page 7434, section II. Aging Monitoring Tags note "The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices. " Shamshiri here uses MTTF as a metric, to determine potential outage of a group of device.
Later, see Shamshiri in page 7434, describe "Our proposed framework continuously monitors the situation of each device and provides an estimation of the current device state. As new datum flows in, the system will recalculate the estimated remaining useful life. This solution can be implemented on popular platforms and connect IoT devices through local networks." Shamshiri mentions that these devices are connected through local networks and are constantly monitored to predict current device state to update to local networks. See Shamshiri in page 7443 for more information.
Claim 10:
Regarding claim 10, Shamshiri in view of Bemani, further in view of Abdulrahman, teach the limitations of claim 8.
Further, Bemani teaches “10. The system of claim 8, wherein the local federated server is configured to update the first generic machine learning model in response to a change in the first local machine learning model, the second machine learning model, or both.”
See Bemani in page 3, Introduction section, third paragraph on page, describe “In local training, the edge device downloads the model from the fog and computes an updated model using its local data. The fog server then collects these updated models mainly by averaging. FL can be used between fog and cloud also. In this case, the models’ parameters will typically be aggregated in the cloud by federating method and then distributed between the fog servers.” Here, Bemani shows that the fog server (i.e. local federated server) collects any updates from the local models (which includes any update to a first local machine model), and then updates from those models are distributed to the other fog servers in the cloud. The examiner construes the language in response to a change in the first model, second model, or both, to mean updates to either of these models, and not all of these models.
Claim 11 is rejected under 35 U.S.C. 103 over Shamshiri in view of Bemani, further in view of Abdulrahman, and further in view of Deng Y. et al., "FAIR: Quality-Aware Federated Learning with Precise User Incentive and Model Aggregation," published on July 26, 2021, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9488743 , (hereafter, Deng).
Claim 11:
Regarding claim 11, Shamshiri in view of Bemani, further in view of Abdulrahman, teach the limitations of claim 8.
However, Shamshiri in view of Bemani, further in view of Abdulrahman, did not teach “11. The medium of claim 8, global federated server is configured to update the global machine learning model in response to a change in the first local machine learning model, the second local machine learning model, the first generic machine learning model, the second generic machine learning model, or any combination thereof.”
In an analogous art, Deng teaches “11. The medium of claim 8, global federated server is configured to update the global machine learning model in response to a change in the first local machine learning model, the second local machine learning model, the first generic machine learning model, the second generic machine learning model, or any combination thereof.”
See Deng in page 4, section IV. Design of FAIR, part A. Estimating Learning Quality , part 1) Learning Quality Quantification , describe “In federated learning, both the volume and quality of the training data can affect the learning quality significantly. The quantification of the learning quality should adequately reflect how useful that the local model updates can contribute to the global model. One plausible approach is to adopt each node’s local model accuracy tested on a global dataset as the learning quality. However, in this approach, the test on each local model is required in each iteration, which can inflict significant overhead. Different from the accuracy measurement, the loss value is calculated in training with no additional overhead. Therefore, we leverage the loss reduction in each iteration to quantify the training data quality. Specifically, suppose iteration t starts at time ts and ends at time te. At time te, the received local model updates are aggregated to update the global models, and the next iteration starts." Further, see Deng in page 1, Introduction section, middle of first paragraph, describe, " Specifically, federated learning is a distributed learning framework, where all nodes independently train the global model based on local data and only model updates are committed to the cloud server for aggregation. In this way, distributed model updates can be aggregated to improve the global model quality in a privacy-preserving manner." Here, Deng shows that once local models are updated, this also updates the global model, and this information is sent to a cloud server (i.e. global federated server) that receives the updates and then make changes to the global models.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Shamshiri, Bemani, and Abdulrahman, and incorporate with Deng’s teachings by using the teachings of Shamshiri, Bemani, and Abdulrahman of devices and each training a model to predict MTTF, with the teaching of Deng of providing a global server to update the global machine learning model from changes to local models.
One of ordinary skill in the art would be motivated to do so because by integrating Deng’s framework into the methods of Shamshiri, Bemani, and Abdulrahman, one with ordinary skill in the art would achieve “In this way, distributed model updates can be aggregated to improve the global model quality in a privacy-preserving manner,” (see Deng in page 1, Introduction section, paragraph 1).
Claims 12 and 14 are rejected under 35 U.S.C. 103 over Shamshiri, in view of Liu L. et al. in "Client-Edge-Cloud Hierarchical Federated Learning,", published on July 27, 2020 , available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9148862, (hereafter, Liu L.), and further in view of Bemani.
Claim 12:
Regarding claim 12, Shamshiri teaches “A method, comprising: deploying a first mean time to failure (MTTF) machine learning model specific to a plurality of first memory devices based on characteristics of each one of the plurality of first memory devices;”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices.
Also, see Shamshiri in page 7433, in abstract describe “The proposed framework is implemented in three distinct types of cloud machine learning (ML), edge ML, and a combination of proportion estimation and Zscore.” Here, Shamshiri mentions implementing or deploying the ML models.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to an initial or first machine learning model.
Further, see Shamshiri in page 7438, section B. mention “There is no soft limit on the available storage for cloud-oriented databases (e.g., distributed, federated, and clustered).” Here, Shamshiri mentions that the method involves various types of learning such as distributed, federated learning.
Further, Shamshiri teaches “deploying a second MTTF machine learning model specific to a plurality of second memory devices based on characteristics of each one of the plurality of second memory devices;”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices.
Also, see Shamshiri in page 7433, in abstract describe “The proposed framework is implemented in three distinct types of cloud machine learning (ML), edge ML, and a combination of proportion estimation and Zscore.” Here, Shamshiri mentions implementing or deploying the ML models.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to an initial or first machine learning model.
Further, see Shamshiri in page 7438, section B. mention “There is no soft limit on the available storage for cloud-oriented databases (e.g., distributed, federated, and clustered).” Here, Shamshiri mentions that the method involves various types of learning such as distributed, federated learning.
Further, Shamshiri teaches “and predicting a respective MTTF for each of the plurality of first memory devices and the plurality of second memory devices based on output data of the fifth MTTF machine learning model, the output data received from the third machine learning model, the output data of the first MTTF machine learning model, the output data of the second MTTF machine learning model, or any combination thereof.”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated... In this article, we used Vth variation as a criterion for estimating the lifetime of devices since all aging phenomena influence Vth . It is possible to evaluate the aging effects on integrated circuits by both pre- and post-profiling. Preprofiling is a theoretical evaluation of aging effects on transistors’ fundamental parameters such as Vth , as expressed as follows:
Vth∝⟨L,k,P,V,T,t⟩(3)
where L and k are the length and substrate factor of the transistor, respectively. As well, P,V , and T are the process, voltage, and temperature parameters, and t is the time. Whereas post-profiling evaluates the aging effects on specific circuit outputs according to
Vout=k1Vth+k2 (4) where k1 and k2 are constant coefficients. If a sensor obtains the threshold voltage as an output voltage of a node in designing an aging monitoring embedded tag, the results of pre- and post-profiling have the same electrical properties.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. The voltage output relates to the data output information.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to an initial or first machine learning model.
Further, see Shamshiri in page 7438, section B. mention “There is no soft limit on the available storage for cloud-oriented databases (e.g., distributed, federated, and clustered).” Here, Shamshiri mentions that the method involves various types of learning such as distributed, federated learning.
Regarding the limitations “first MTTF machine learning model and the second MTTF machine learning model” and “third MTTF machine learning model” , and “fourth MTTF …model” or “fifth MTTF … model”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated... In this article, we used Vth variation as a criterion for estimating the lifetime of devices since all aging phenomena influence Vth . It is possible to evaluate the aging effects on integrated circuits by both pre- and post-profiling. Preprofiling is a theoretical evaluation of aging effects on transistors’ fundamental parameters … If a sensor obtains the threshold voltage as an output voltage of a node in designing an aging monitoring embedded tag, the results of pre- and post-profiling have the same electrical properties.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. The voltage output relates to the data output information.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to a machine learning model that predicts MTTF.
However, Shamshiri did not teach “ aggregating output data received from the first …machine learning model and the second … machine learning model into a third … machine learning model;” or “aggregating output data received from the third … machine learning model and a fourth … machine learning model into a fifth … machine learning model;”
In an analogous method, Liu L. teaches “aggregating output data received from the first … machine learning model and the second … machine learning model into a third … machine learning model;”
See Liu L. in pages 2-3, in section II. Federated learning systems, part C. Client-Edge-Cloud Hierarchical FL describe “In FAVG, the model aggregation step can be interpreted as a way to exchange information among the clients. Thus, aggregation at the cloud parameter server can incorporate many clients, but the [communication] cost is high. On the other hand, aggregation at the edge parameter server only incorporates a small number of clients with much cheaper [communication] cost. To combine their advantages, we consider a hierarchical FL system, which has one cloud server, L edge servers indexed by ℓ, with disjoint client sets {Cℓ}Lℓ=1, and N clients indexed by i and ℓ, with distributed datasets {Dℓi}Ni=1. Denote Dℓ as the aggregated dataset under edge ℓ. Each edge server aggregates models from its clients.
Algorithm 1: Hierarchical Federated Averaging (HierFAVG)
PNG
media_image4.png
879
357
media_image4.png
Greyscale
PNG
media_image5.png
41
183
media_image5.png
Greyscale
With this new architecture, we extend the FAVG to a HierFAVG algorithm. The key steps of the HierFAVG algorithm proceed as follows. After every κ1 local updates on each client, each edge server aggregates its clients’ models. Then after every κ2 edge model aggregations, the cloud server aggregates all the edge servers’ models, which means that the communication with the cloud happens every κ1κ2 local updates.” Here, Liu L. describes that “After every κ1 local updates on each client, each edge server aggregates its clients’ models” showing that any local model from a first or second client device can be combined to form an aggregated model (i.e. third model). Algorithm 1 and figure 3 illustrates the hierarchical steps of multiple iteration of aggregations of local models into one global model, repeated in many rounds.
Further, see Liu L. teach “ aggregating output data received from the third … machine learning model and a fourth … machine learning model into a fifth …machine learning model;”
See Liu L. in pages 2-3, in section II. Federated learning systems, part C. Client-Edge-Cloud Hierarchical FL describe “In FAVG, the model aggregation step can be interpreted as a way to exchange information among the clients. Thus, aggregation at the cloud parameter server can incorporate many clients, but the [communication] cost is high. On the other hand, aggregation at the edge parameter server only incorporates a small number of clients with much cheaper [communication] cost. To combine their advantages, we consider a hierarchical FL system, which has one cloud server, L edge servers indexed by ℓ, with disjoint client sets {Cℓ}Lℓ=1, and N clients indexed by i and ℓ, with distributed datasets {Dℓi}Ni=1. Denote Dℓ as the aggregated dataset under edge ℓ. Each edge server aggregates models from its clients.
Algorithm 1: Hierarchical Federated Averaging (HierFAVG)
PNG
media_image6.png
734
594
media_image6.png
Greyscale
PNG
media_image5.png
41
183
media_image5.png
Greyscale
With this new architecture, we extend the FAVG to a HierFAVG algorithm. The key steps of the HierFAVG algorithm proceed as follows. After every κ1 local updates on each client, each edge server aggregates its clients’ models. Then after every κ2 edge model aggregations, the cloud server aggregates all the edge servers’ models, which means that the communication with the cloud happens every κ1κ2 local updates.” Here, Liu L. describes that “After every κ1 local updates on each client, each edge server aggregates its clients’ models”. With the iterative nature of this method, Liu L. also shows that any local model from a third or fourth client device can be combined to form an aggregated model (i.e. fifth model). Algorithm 1 and figure 3 illustrates the hierarchical steps of multiple iteration of aggregations of local models into one global model, repeated in many rounds.
For more information, see Liu L. in page 2, first paragraph describes “First, by extending the FAVG algorithm to the hierarchical setting, will the new algorithm still converge? Given the two levels of model aggregation (one at the edge, one at the cloud), how often should the models be aggregated at each level? … In this paper, we address these key questions. First, a rigorous proof is provided to show the convergence of the training algorithm. Through convergence analysis, some [qualitative] guidelines on picking the aggregation frequencies at two levels are also given.” Liu L. describes two levels of aggregation in federated learning.
Further, see Liu L. teach “ aggregating output data received from the third …machine learning model and a fourth … machine learning model into a fifth …machine learning model;”
See Liu L. in pages 2-3, in section II. Federated learning systems, part C. Client-Edge-Cloud Hierarchical FL describe “In FAVG, the model aggregation step can be interpreted as a way to exchange information among the clients. Thus, aggregation at the cloud parameter server can incorporate many clients, but the [communication] cost is high. On the other hand, aggregation at the edge parameter server only incorporates a small number of clients with much cheaper [communication] cost. To combine their advantages, we consider a hierarchical FL system, which has one cloud server, L edge servers indexed by ℓ, with disjoint client sets {Cℓ}Lℓ=1, and N clients indexed by i and ℓ, with distributed datasets {Dℓi}Ni=1. Denote Dℓ as the aggregated dataset under edge ℓ. Each edge server aggregates models from its clients.
Algorithm 1: Hierarchical Federated Averaging (HierFAVG)
PNG
media_image7.png
754
608
media_image7.png
Greyscale
PNG
media_image5.png
41
183
media_image5.png
Greyscale
With this new architecture, we extend the FAVG to a HierFAVG algorithm. The key steps of the HierFAVG algorithm proceed as follows. After every κ1 local updates on each client, each edge server aggregates its clients’ models. Then after every κ2 edge model aggregations, the cloud server aggregates all the edge servers’ models, which means that the communication with the cloud happens every κ1κ2 local updates.” Here, Liu L. describes that “After every κ1 local updates on each client, each edge server aggregates its clients’ models”. With the iterative nature of this method, Liu L. also shows that any local model from a third or fourth client device can be combined to form an aggregated model (i.e. fifth model). Algorithm 1 and figure 3 illustrates the hierarchical steps of multiple iteration of aggregations of local models into one global model, repeated in many rounds.
More info: See Liu L. in page 2, first paragraph describes “First, by extending the FAVG algorithm to the hierarchical setting, will the new algorithm still converge? Given the two levels of model aggregation (one at the edge, one at the cloud), how often should the models be aggregated at each level? [Moreover], by allowing frequent local updates, can a better latency-energy tradeoff be achieved? In this paper, we address these key questions. First, a rigorous proof is provided to show the convergence of the training algorithm. Through convergence analysis, some [qualitative] guidelines on picking the aggregation frequencies at two levels are also given.”
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Shamshiri and incorporate into the teachings of Liu L. because both references teach generating a model that predicts a respective MTTF for each device based on device characteristics.
One of ordinary skill in the art would be motivated to do so because this would achieve “In the experiments, we also notice that using SGD with momentum can speed up training and improve the final accuracy evidently,” (see Liu L. in page 5, section IV. Experiments, second paragraph on section).
However, Shamshiri in view of Liu L. did not teach “deploying … a …model … based on characteristics of each one of the plurality of first memory devices” or “deploying a second … model …based on characteristics of each one of the plurality of second memory devices”,
However, in an analogous field, Bemani teaches “deploying a first … model …based on characteristics of each one of the plurality of first memory devices” and “deploying a second … model …based on characteristics of each one of the plurality of second memory devices”
See Bemani in pages 5-6, section 2.1, describe "The FedAvg algorithm works
by running the training task on the edge devices, where they share an overall model with the central server that is an average of all the parameters." Bemani here shows running the model or deploying the model.
Further, see Bemani in page 12 last paragraph, part of section 5.1 Distributing CMAPSS for a Collaborative PM describe "each time series consists of 21 sensor observations, three operating settings, a trajectory id, and a cycle count, which the RUL of an engine is estimated from it in terms of the number of operation cycles before the engine runs to failure. The goal is to predict the RUL based on the time series data by the model at the cloud level, which was trained by the federated aggregation from the fog servers and, subsequently, edge devices." Bemani shows running the model on each of the edge devices.
Further, see Bemani in page 7, section 3.1 Network model, describe "We assume that each i asset on a factory site collects measurement data and has information about labeled training samples, such as RUL information of one asset with its sensors measurement." Here, Bemani mentions collecting information or data about measurements of devices of each asset or per device, and is based on characteristics of each of those devices.
Further, see Bemani in page 7, section 3.1. Network Model describe “As depicted in Figure 2, we considered a general FL-supported wireless multi-agent network between fog server and N distributed edge devices, denoted as the set N = {1, . . . , N}. The fog server is directly connected to the cloud server through a wireless link with the nearest base station. The fog server in each factory is also equipped with computational resources to provide communication and computation services to the edge devices.” Bemani here describes multiple devices, indicated by the set N from 1 through N, where N can be any number, or relate to collecting data as well as respective features for each computing device or subsequent devices.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Shamshiri along with the secondary reference of Liu L. with the teachings of Bemani by using the teachings of Shamshiri and Liu L. of generating a model that predicts a respective MTTF for each device based on device characteristics, with Bemani’s teaching of deploying those models for each device.
One of ordinary skill in the art would be motivated to do so because by integrating Bemani’s framework into the methods of Shamshiri and Liu L., one with ordinary skill in the art would achieve the goal of providing “By comparison of the FedLSTM results with other prediction methods in Table 5, it can be confirmed that the performance of the proposed FedLSTM in PM applications has comparable efficiency to the conventional centralized approaches in terms of prediction accuracy,” (see page 20, section 6.2, FedLSTM results, first paragraph in Bemani).
Claim 14:
Regarding claim 14, Shamshiri in view of Liu L., further in view of Bemani, teach the limitations of claim 12.
Further, Shamshiri teaches “14. The method of claim 13, further comprising providing an updated predicted MTTF to the … device ...”
See Shamshiri in page 7434, section II. Aging Monitoring Tags note "The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices. " Shamshiri here uses MTTF as a metric, to determine potential outage of a group of device.
Also, see Shamshiri in page 7433, Introduction, second paragraph, describe “By definition, reliability is the ability of a device to operate as predicted over a specified time. Therefore, by continuously evaluating the device and updating the estimation of the remaining life (and extending the useful life by modifying the usage procedure), the reliability can be increased [8]. This issue further highlights the importance of using aging tags in IoT devices”. Here, Shamshiri mentions continuously updating estimation of the remaining life by MTTF, provided by a MTTF model, this shows that the model is frequently updated.
Further, Liu L. teaches “14. The method of claim 13, further comprising providing an updated predicted … to the host device each time one of the first, the second, the third, the fourth, or the fifth MMTF machine learning models is updated.”
See Liu L. from page 2, Introduction, describe "From the above comparison, we see a necessity in leveraging a cloud server to access the massive training samples, while each edge server enjoys quick model updates with its local clients. This motivates us to propose a client-edge-cloud hierarchical FL system as shown on the right side of Fig. 1, to get the best of both systems. Compared with cloud-based FL, hierarchical FL will significantly reduce the costly communication with the cloud, supplemented by efficient client-edge updates, thereby, resulting a significant reduction in both the runtime and number of local iterations." Here, Liu L. describes the updates from models gets sent to the cloud server or a host device. See algorithm 1 on page for the hierarchical federated learning process.
PNG
media_image8.png
842
691
media_image8.png
Greyscale
In the algorithm, Liu L. mentions that there are multiple rounds of aggregation of the models, and each aggregation is then updated to the cloud server or host. Each round can include subsequent models, such as second, third, fourth, fifth or additional models.
Claim 13, 16, 17 and 18 are rejected under 35 U.S.C. 103 over Shamshiri in view of Liu L., further in view of Bemani, and further in view of Bharti, S. et al., in “Privacy-aware resource sharing in cross-device federated model training for collaborative predictive maintenance,” published on August 30, 2021, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9525091 , (hereafter, Bharti).
Claim 13:
Regarding claim 13, Shamshiri in view of Liu L., further in view of Bemani, teach the limitations of claim 12.
Further, Shamshiri teaches “13. The method of claim 12, further comprising providing the predicted MTTF to a host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both.”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated.” Here, Shamshiri mentions that estimating (i.e. predicting) the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. These groups relate to memory device characteristics associated with a first group of memory devices.
However, Shamshiri in view of Liu L., further in view of Bemani, did not teach “The method of claim 12, further comprising providing the predicted MTTF to a host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both.”
In an analogous field, Bharti teaches “The method of claim 12, further comprising providing the predicted MTTF to a host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both.”
See Bharti in page 120638, part A. FL for collaborative PdM, describe “In a typical FL process, these edge devices participate as FL clients training the global model with their own data and exchanging global model updates with a centralized server. The server then aggregates the received model updates to create an improved global model, which can be distributed among clients for use in their PdM application. However, continuous improvement in the global model is achieved only by clients participating in multiple iterations of the training process. This iterative exchange of model updates instead of raw manufacturing data allows a manufacturer to remain in control and protect the privacy of this data at all times,” Here, Bharti shows that each edge device serve as a host device, which also hosts training of a global model, where the model updates are distributed among clients or devices (i.e. first memory devices or group of devices). Memory device is construed to mean any client or device that are part of a system.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Shamshiri, Liu L., and Bemani, with Bharti, by using the teachings of Shamshiri, Liu L., Bemani, of running a model that predicts a respective MTTF for each device based on device characteristics, and incorporate with Bharti’s teachings of providing a host device hosting training.
One of ordinary skill in the art would be motivated to do so because by integrating Bharti’s framework into the methods of Shamshiri, Liu L., and Bemani, one with ordinary skill in the art would achieve a method that “enables FL clients to maximise available resources within their local network without compromising the benefits of a FL approach (i.e., privacy and shared learning). Experiments performed on the benchmark C-MAPSS data-set demonstrate the advantage of applying SplitPred in the FL process in terms of efficient use of resources, i.e., model convergence time, accuracy, and network load,” (see Bharti in page 120367, abstract).
Claim 16:
Regarding claim 16, Shamshiri in view of Liu L., and further in view of Bemani, teach the limitations of claim 12.
Further, Shamshiri teaches “The method of claim 12, further comprising deploying the first MTTF machine learning model, the second MTTF machine learning model, or both, on memory devices of a second host device different than a first host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated... In this article, we used Vth variation as a criterion for estimating the lifetime of devices since all aging phenomena influence Vth . It is possible to evaluate the aging effects on integrated circuits by both pre- and post-profiling. Preprofiling is a theoretical evaluation of aging effects on transistors’ fundamental parameters … If a sensor obtains the threshold voltage as an output voltage of a node in designing an aging monitoring embedded tag, the results of pre- and post-profiling have the same electrical properties.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. The voltage output relates to the data output information.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to a machine learning model that predicts MTTF.
However, Shamshiri in view of Liu L., and further in view of Bemani, did not teach “The method of claim 12, … of a second host device different than a first host device hosting training of the plurality of first … devices, the plurality of second … devices, or both.”
In an analogous method, Bharti teaches “The method of claim 12, … of a second host device different than a first host device hosting training of the plurality of first … devices, the plurality of second … devices, or both.”
See Bharti in page 120368, Introduction, part A. FL for Collaborative PdM, mention “In traditional FL, the data across these clients can share the same feature space (horizontally partitioned) but belong to a different sample ID space [7]. In the context of collaborative PdM, the asset degradation pattern data can be horizontally partitioned across similar types of assets working under similar operating conditions across multiple manufacturing sites. On the other hand, the data across clients can also share the same sample IDs but differ on the features (vertically partitioned). This is when the same asset is being monitored using different type of sensors across manufacturing sites.” Bharti shows clients belong to different sample ID across multiple manufacturing sites, and clients here are viewed as host devices, which show different devices.
Later, see Bharti in page 120369, section C. Contributions mention “ The model is divided into layers that are allocated to different devices. The framework is not fully edge based and there is no provision of resource sharing among edge devices.” Here, Bharti shows that the global model is distributed to different edge devices across various manufacturing sites for training and inference. Since there are multiple manufacturing sites, each have different type of sensors, these different sensors can be host devices that are different from the host devices already used for training the first group of memory devices. Bharti here shows running a trained model onto devices hosted at another site on different types of sensors, where sensor relates to host device (i.e. second host device different than a first host device).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Shamshiri, Liu L., and Bemani, with Bharti, by using the teachings of Shamshiri, Liu L., Bemani, of running a model that predicts a respective MTTF for each device based on device characteristics, and incorporate with Bharti’s teachings of using a second host device different than a first host device.
One of ordinary skill in the art would be motivated to do so because by integrating Bharti’s framework into the methods of Shamshiri, Liu L., and Bemani, one with ordinary skill in the art would achieve a method that “enables FL clients to maximise available resources within their local network without compromising the benefits of a FL approach (i.e., privacy and shared learning). Experiments performed on the benchmark C-MAPSS data-set demonstrate the advantage of applying SplitPred in the FL process in terms of efficient use of resources, i.e., model convergence time, accuracy, and network load,” (see Bharti in page 120367, abstract).
Claim 17:
Regarding claim 17, Shamshiri in view of Liu L., and further in view of Bemani, teach the limitations of claim 12.
Further, Shamshiri teaches “deploying the third MTTF machine learning model, the fourth MTTF machine learning model … or both on memory devices”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated... In this article, we used Vth variation as a criterion for estimating the lifetime of devices since all aging phenomena influence Vth . It is possible to evaluate the aging effects on integrated circuits by both pre- and post-profiling. Preprofiling is a theoretical evaluation of aging effects on transistors’ fundamental parameters … If a sensor obtains the threshold voltage as an output voltage of a node in designing an aging monitoring embedded tag, the results of pre- and post-profiling have the same electrical properties.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. These devices, both transistors-related and circuit-related devices, are considered to be part of memory devices. The voltage output relates to the data output information.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to a machine learning model that predicts MTTF. Here, Shamshiri shows that ‘implemented the edge ML type’ is construed to also mean deploying the MTTF model.
Further, Liu L. teaches “The method of claim 12, further comprising deploying the third … machine learning model, the … fourth… machine learning model, …”
See Liu L. in page 2, first paragraph mention “Given the two levels of model aggregation (one at the edge, one at the cloud), how often should the models be aggregated at each level.”
Later, see Liu L. in page 3, section II. Federated learning systems, part C. Client-edge-cloud hierarchical FL, describe “With this new architecture, we extend the FAVG to a HierFAVG algorithm. The key steps of the HierFAVG algorithm proceed as follows. After every κ1 local updates on each client, each edge server aggregates its clients’ models. Then after every κ2 edge model aggregations, the cloud server aggregates all the edge servers’ models, which means that the communication with the cloud happens every κ1κ2 local updates.” See Liu in algorithm I. Here, Liu shows that using hierarchical federated learning, this generates two levels of aggregation of models, the first aggregation produces the first iteration of an aggregated or combined model (i.e. third model), and the second aggregation produces the second iteration of an aggregated or combined model (i.e. fifth model).
However, Shamshiri in view of Liu L., and further in view of Bemani, did not teach “… memory devices of a second host device different than a first host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both.”
In an analogous method, Bharti teaches “The method of claim 12, … of a second host device different than a first host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both.”
See Bharti in page 120368, Introduction, part A. FL for Collaborative PdM, mention “In traditional FL, the data across these clients can share the same feature space (horizontally partitioned) but belong to a different sample ID space [7]. In the context of collaborative PdM, the asset degradation pattern data can be horizontally partitioned across similar types of assets working under similar operating conditions across multiple manufacturing sites. On the other hand, the data across clients can also share the same sample IDs but differ on the features (vertically partitioned). This is when the same asset is being monitored using different type of sensors across manufacturing sites.”
Later, see Bharti in page 120369, section C. Contributions mention “ The model is divided into layers that are allocated to different devices. The framework is not fully edge based and there is no provision of resource sharing among edge devices.” Here, Bharti shows that the global model is distributed to different edge devices across various manufacturing sites for training and inference. Since there are multiple manufacturing sites, each have different type of sensors, these different sensors can be host devices that are different from the host devices already used for training the first group of memory devices. Bharti here shows running a trained model onto devices hosted at another site on different types of sensors, where sensor relates to host device (i.e. second host device different than a first host device).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Shamshiri, Liu L., and Bemani, with Bharti, by using the teachings of Shamshiri, Liu L., Bemani, of running a model that predicts a respective MTTF for each device based on device characteristics, and incorporate with Bharti’s teachings of using a second host device different than a first host device.
One of ordinary skill in the art would be motivated to do so because by integrating Bharti’s framework into the methods of Shamshiri, Liu L., and Bemani, one with ordinary skill in the art would achieve a method that “enables FL clients to maximise available resources within their local network without compromising the benefits of a FL approach (i.e., privacy and shared learning). Experiments performed on the benchmark C-MAPSS data-set demonstrate the advantage of applying SplitPred in the FL process in terms of efficient use of resources, i.e., model convergence time, accuracy, and network load,” (see Bharti in page 120367, abstract).
Claim 18:
Regarding claim 18, Shamshiri in view of Liu L., and further in view of Bemani, teach the limitations of claim 12.
Further, Shamshiri teaches “The method of claim 12, further comprising deploying the … MTTF machine learning model on memory devices of a second host device different than a first host device hosting training of the plurality of first memory devices, the plurality of second memory devices, or both”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated... In this article, we used Vth variation as a criterion for estimating the lifetime of devices since all aging phenomena influence Vth . It is possible to evaluate the aging effects on integrated circuits by both pre- and post-profiling. Preprofiling is a theoretical evaluation of aging effects on transistors’ fundamental parameters … If a sensor obtains the threshold voltage as an output voltage of a node in designing an aging monitoring embedded tag, the results of pre- and post-profiling have the same electrical properties.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. These devices, both transistors-related and circuit-related devices, are considered to be part of memory devices. The voltage output relates to the data output information.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to a machine learning model that predicts MTTF. Here, Shamshiri shows that ‘implemented the edge ML type’ is construed to also mean deploying the MTTF model.
Further, Liu L. teaches “The method of claim 12, further comprising deploying the fifth …machine learning model on memory devices …”
See Liu L. in page 2, first paragraph mention “Given the two levels of model aggregation (one at the edge, one at the cloud), how often should the models be aggregated at each level.”
Later, see Liu L. in page 3, section II. Federated learning systems, part C. Client-edge-cloud hierarchical FL, describe “With this new architecture, we extend the FAVG to a HierFAVG algorithm. The key steps of the HierFAVG algorithm proceed as follows. After every κ1 local updates on each client, each edge server aggregates its clients’ models. Then after every κ2 edge model aggregations, the cloud server aggregates all the edge servers’ models, which means that the communication with the cloud happens every κ1κ2 local updates.” See Liu in algorithm I. Here, Liu shows that using hierarchical federated learning, this generates two levels of aggregation of models, the first aggregation produces the first iteration of an aggregated or combined model (i.e. third model), and the second aggregation produces the second iteration of an aggregated or combined model (i.e. fifth model).
PNG
media_image4.png
879
357
media_image4.png
Greyscale
Liu L. mentions algorithm 1, which applies to a second round of aggregation of the model, and this second round of aggregation generates a fifth model.
However, Shamshiri in view of Liu L., and further in view of Bemani, did not teach “The method of claim 12, …of a second host device different than a first host device hosting training of the plurality of … devices...”
In an analogous art, Bharti teaches “The method of claim 12, …of a second host device different than a first host device hosting training of the plurality of … devices...”
See Bharti in page 120368, Introduction, part A. FL for Collaborative PdM, mention “In traditional FL, the data across these clients can share the same feature space (horizontally partitioned) but belong to a different sample ID space [7]. In the context of collaborative PdM, the asset degradation pattern data can be horizontally partitioned across similar types of assets working under similar operating conditions across multiple manufacturing sites. On the other hand, the data across clients can also share the same sample IDs but differ on the features (vertically partitioned). This is when the same asset is being monitored using different type of sensors across manufacturing sites.” Here, Bharti shows clients belong to different sample ID across multiple manufacturing sites, and clients here are viewed as host devices, which show different devices.
Later, see Bharti in page 120369, section C. Contributions mention “ The model is divided into layers that are allocated to different devices. The framework is not fully edge based and there is no provision of resource sharing among edge devices.” Here, Bharti shows that the global model is distributed to different edge devices across various manufacturing sites for training and inference. Since there are multiple manufacturing sites, each have different type of sensors, these different sensors can be host devices that are different from the host devices already used for training the first group of memory devices. Bharti here shows running a trained model onto devices hosted at another site on different types of sensors, where sensor relates to host device (i.e. second host device different than a first host device).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Shamshiri, Liu L., and Bemani, with Bharti, by using the teachings of Shamshiri, Liu L., Bemani, of running a model that predicts a respective MTTF for each device based on device characteristics, and incorporate with Bharti’s teachings of using a second host device different than a first host device.
One of ordinary skill in the art would be motivated to do so because by integrating Bharti’s framework into the methods of Shamshiri, Liu L., and Bemani, one with ordinary skill in the art would achieve a method that “enables FL clients to maximise available resources within their local network without compromising the benefits of a FL approach (i.e., privacy and shared learning). Experiments performed on the benchmark C-MAPSS data-set demonstrate the advantage of applying SplitPred in the FL process in terms of efficient use of resources, i.e., model convergence time, accuracy, and network load,” (see Bharti in page 120367, abstract).
Claims 15, 19, and 20 are rejected under 35 U.S.C. 103 over Shamshiri in view of Liu L., further in view of Bemani, and further in view of Lim, W. et al. in “Federated learning in mobile edge networks: A comprehensive survey Federated learning in mobile edge networks: A comprehensive survey,” published on April 8th 2020, available at: https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9060868 , (hereafter, Lim).
Claim 15:
Regarding claim 15, Shamshiri in view of Liu L., and further in view of Bemani, teach the limitations of claim 12.
Further, Shamshiri teaches “The method of claim 12, further comprising …data from a … MTTF machine learning model and a … MTTF machine learning model into the … MTTF machine learning model … wherein the … MTTF machine learning model is specific to a plurality of … memory devices”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated... In this article, we used Vth variation as a criterion for estimating the lifetime of devices since all aging phenomena influence Vth . It is possible to evaluate the aging effects on integrated circuits by both pre- and post-profiling. Preprofiling is a theoretical evaluation of aging effects on transistors’ fundamental parameters … If a sensor obtains the threshold voltage as an output voltage of a node in designing an aging monitoring embedded tag, the results of pre- and post-profiling have the same electrical properties.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. These devices, both transistors-related and circuit-related devices, are considered to be part of memory devices. The voltage output relates to the data output information.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to a machine learning model that predicts MTTF. Here, Shamshiri shows that ‘implemented the edge ML type’ is construed to also mean deploying the MTTF model.
However, Shamshiri in view of Liu L. and further in view of Bemani, did not teach “… aggregating output data from a sixth … machine learning model and a seventh … machine learning model into the fourth … machine learning model, wherein the sixth … machine learning model is specific to a plurality of third memory devices, and the seventh … machine learning model is specific to a plurality of fourth memory devices.”
In an analogous art, Lim teaches “The method of claim 12, further comprising aggregating output data from a sixth … machine learning model and a seventh … machine learning model into the fourth … machine learning model,”
See algorithm 1, and in page 2036, where Lim describes "Global model aggregation is an integral part of FL. A straightforward and classical algorithm for aggregating the local models is the FedAvg algorithm proposed in [23], which is similar to that of local SGD [64]. The pseudocode for FedAvg is given in Algorithm 1. As described in Step 1 above, the server first initializes the task (lines 11–16). " Since model aggregation is performed iteratively, Lim describes a method that takes future subsequent models such as a local model from a sixth or seventh round of training, and aggregate this into a fourth round of aggregation.
For details, See Lim in page 2036 in section B. Federated Learning describe "Global model aggregation is an integral part of FL. A straightforward and classical algorithm for aggregating the ... At the tth iteration (line 17), the server minimizes the global loss in (4) by the averaging aggregation which is formally defined as ..."
PNG
media_image9.png
793
918
media_image9.png
Greyscale
Further, Lim teaches “wherein the sixth … machine learning model is specific to a plurality of third memory devices, and the seventh … machine learning model is specific to a plurality of fourth memory devices.”
See Lim in section VII, Challenges and future research directions on page 2058 … describe "Cooperative mobile crowd ML: In the existing approaches, mobile devices need to communicate with the server directly and this may increase the energy consumption. In fact, mobile devices nearby can be grouped in a cluster, and the model downloading/uploading between the server and the mobile devices can be facilitated by a “cluster head” that serves as a relay node [205]. The model exchange between the mobile devices and the cluster head can then be done in Device-to-Device (D2D) connections. Such a model can improve the energy efficiency significantly. " Here, Lim shows that each model is associated with a group of devices grouped as a cluster, (i.e. one model per device).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Shamshiri, Liu L., and Bemani, with Lim, by using the teachings of Shamshiri, Liu L., Bemani, of running a model that predicts a respective MTTF for each device based on device characteristics, and incorporate with Lim’s teachings of aggregating output data from a sixth MTTF machine learning model and a seventh MTTF machine learning model into the fourth model.
One of ordinary skill in the art would be motivated to do so because by integrating Lim’s framework into the methods of Shamshiri, Liu L., and Bemani, one with ordinary skill in the art would achieve a method where “FedAvg increases the accuracy eventually since model averaging produces regularization effects similar to dropout [97], which prevents overfitting,” (see Lim in page 2039, section A. Edge and End Computation).
Claim 19:
Regarding claim 19, Shamshiri in view of Liu L., and further in view of Bemani, teach the limitations of claim 12.
Further, Shamshiri teaches “The method of claim 12, further comprising updating the … MTTF machine learning model .”
See Shamshiri in pages 7434 – 7435, in section II. Aging Monitoring Tags describe “The lifetime of a device is determined by the time it will operate reliably. The mean-time-to-failure (MTTF) is one of the most conventional methods for estimating this time for a population of devices… Lifetime prediction systems are also divided into two categories: 1) transistor-based and 2) circuit-based parameters. In the former, the parameters of the target transistor (e.g., threshold voltage, drain current, and transconductance) are evaluated. In the latter, the circuit performance parameters (e.g., the phase difference of two ring oscillators, whereas one of them being under stress) are evaluated... In this article, we used Vth variation as a criterion for estimating the lifetime of devices since all aging phenomena influence Vth . It is possible to evaluate the aging effects on integrated circuits by both pre- and post-profiling. Preprofiling is a theoretical evaluation of aging effects on transistors’ fundamental parameters … If a sensor obtains the threshold voltage as an output voltage of a node in designing an aging monitoring embedded tag, the results of pre- and post-profiling have the same electrical properties.” Shamshiri mentions that estimating the MTTF for a population of devices or group of devices, which are divided into two groups with a first group is transistor-based and second group being circuit-based devices. These devices, both transistors-related and circuit-related devices, are considered to be part of memory devices. The voltage output relates to the data output information.
Further, see Shamshiri in page 7441, section C. Proposed ML-powered framework, describe “The calculation time of the proposed algorithm (in the most advanced version) is negligible for cloud platforms. When cloud resources are not accessible or desirable due to data privacy, calculations must be performed locally. We implemented the edge ML type by reducing the number of epochs and layers of the DNN at the cost of a slight accuracy loss. However, using a neural processing unit (NPU) for a private IoT network increases the accuracy (near to cloud ML) and provides data privacy.” Here, Shamshiri mentions that the calculation time for predicting MTTF is based on an algorithm such as a DNN or deep neural network model, and relates to a machine learning model that predicts MTTF. Here, Shamshiri shows that ‘implemented the edge ML type’ is construed to also mean deploying the MTTF model.
Also, see Shamshiri in page 7433, Introduction, second paragraph, describe “By definition, reliability is the ability of a device to operate as predicted over a specified time. Therefore, by continuously evaluating the device and updating the estimation of the remaining life (and extending the useful life by modifying the usage procedure), the reliability can be increased [8]. This issue further highlights the importance of using aging tags in IoT devices”. Here, Shamshiri mentions continuously updating estimation of the remaining life by MTTF, provided by a MTTF model, this shows that the model is frequently updated.
However, Shamshiri in view of Liu L., and further in view of Bemani, did not teach “The method of claim 12, further comprising updating the fifth … machine learning model in response to a change in the third … machine learning model, the fourth … machine learning model, or both.”
In an analogous art, Lim teaches “The method of claim 12, further comprising updating the fifth … machine learning model in response to a change in the third … machine learning model, the fourth … machine learning model, or both.”
See Lim mention in page 2037, section II Background and fundamentals of federated learning, part C. Statistical challenges mention "In LoAdaBoost FedAvg, participants train the model on their local data and compare the cross-entropy loss with the median loss from the previous training round. If the current cross-entropy loss is higher, the model is retrained before global aggregation so as to increase learning efficiency." Lim shows that if any model information such as model parameter like cross-entropy loss is updated, then the model gets retrained or updated as well.
See algorithm 1, and in page 2036, where Lim describes "Global model aggregation is an integral part of FL. A straightforward and classical algorithm for aggregating the local models is the FedAvg algorithm proposed in [23], which is similar to that of local SGD [64]. The pseudocode for FedAvg is given in Algorithm 1. As described in Step 1 above, the server first initializes the task (lines 11–16). " Since model aggregation is performed iteratively, Lim describes a method that takes future subsequent models such as a local model from a third or fourth and combine to form a fifth model.
For details, See Lim in page 2036 in section B. Federated Learning describe "Global model aggregation is an integral part of FL. A straightforward and classical algorithm for aggregating the ... At the tth iteration (line 17), the server minimizes the global loss in (4) by the averaging aggregation which is formally defined as ..."
PNG
media_image9.png
793
918
media_image9.png
Greyscale
Further, see Lim mention in page 2042, section III. Communication cost, part C. importance based updating, "in fact, the global update is not known in advance before aggregation. As such, the global update made in the previous iteration is used as an estimate for comparison since it was found empirically that more than 99% of the normalized difference of two sequential global updates are smaller than 0.05 in both MNIST CNN and Next-Word-Prediction LSTM." Lim mentions that if an update is made in one of the previous iteration of aggregation (i.e. third model), then the model of the next round of aggregation (i.e. fifth model) gets updated or in this case retrained based on the updates of the previous iteration.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Shamshiri, Liu L., and Bemani, with Lim, by using the teachings of Shamshiri, Liu L., Bemani, of running a model that predicts a respective MTTF for each device based on device characteristics, and incorporate with Lim’s teachings of updating an aggregated model in response to a change to an earlier local model.
One of ordinary skill in the art would be motivated to do so because by integrating Lim’s framework into the methods of Shamshiri, Liu L., and Bemani, one with ordinary skill in the art would achieve a method where “FedAvg increases the accuracy eventually since model averaging produces regularization effects similar to dropout [97], which prevents overfitting,” (see Lim in page 2039, section A. Edge and End Computation).
Claim 20:
Regarding claim 20, Shamshiri in view of Liu L., and further in view of Bemani, teach the limitations of claim 12.
Further, Shamshiri teaches “The method of claim 12, further comprising updating the … MTTF machine learning model …”
However, Shamshiri in view of Liu L., and further in view of Bemani, did not teach “… updating the third … machine learning model in response to a change in the first … machine learning model, the second … machine learning model, or both”
In an analogous method, Lim teaches “… updating the third … machine learning model in response to a change in the first … machine learning model, the second … machine learning model, or both”
See Lim in page 2034, section II, Background and fundamentals of federated learning mention “FL involves the collaborative training of DNN models on end devices. There are, in general, two steps in the FL training process namely (i) local model training on end devices and (ii) global aggregation of updated parameters in the FL server.” Lim describes the process of training local models, then combining those local models into a global model.
Further, see Lim in figure 3, in page 2035 describe using multiple participants or devices, each train their respective local models, and combine them later to aggregated models.
PNG
media_image10.png
854
862
media_image10.png
Greyscale
Further, see Lim in page 2036 , section II part B. Federated learning describe "Step 2 (Local model training and update): Based on the global model wtG , where t denotes the current iteration index, each participant respectively uses its local data and device to update the local model parameters wti . The goal of participant i in iteration t is to find optimal parameters wti that minimize the loss function L(wti) , i.e., wt∗i=argminwtiL(wti).(3)
The updated local model parameters are subsequently sent to the server…
Step 3 (Global model aggregation and update): The server aggregates the local models from participants and then sends the updated global model parameters wt+1G back to the data owners.” Here, Lim shows that each device takes the current global model and trains that model using their own data to create their individual local model. Lim later shows that in another step, that after this training is done, the devices send their updated local model information back to a global server, which then combines these individual updates to create updated global model for the next round. This shows that if a local model (first or second model) gets changed, then this later also changes the global model (i.e. third model).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Shamshiri, Liu L., and Bemani, with Lim, by using the teachings of Shamshiri, Liu L., Bemani, of running a model that predicts a respective MTTF for each device based on device characteristics, and incorporate with Lim’s teachings of updating an aggregated model in response to a change to an earlier local model.
One of ordinary skill in the art would be motivated to do so because by integrating Lim’s framework into the methods of Shamshiri, Liu L., and Bemani, one with ordinary skill in the art would achieve a method where “FedAvg increases the accuracy eventually since model averaging produces regularization effects similar to dropout [97], which prevents overfitting,” (see Lim in page 2039, section A. Edge and End Computation).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENWEI ZENG whose telephone number is (571)272-7111. The examiner can normally be reached Monday-Friday, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WenWei Zeng/Examiner, Art Unit 2146
/SHAHID K KHAN/Primary Examiner, Art Unit 2146