DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement(s) (IDS) submitted on 16 October 2024, 27 October 2025, 09 December 2025, 10 February 2026, and 31 March 2026 is/are being considered by the examiner.
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Applicant has not complied with one or more conditions for receiving the benefit of an earlier filing date under 35 U.S.C. 120 as follows:
The later-filed application must be an application for a patent for an invention which is also disclosed in the prior application (the parent or original nonprovisional application or provisional application). The disclosure of the invention in the parent application and in the later-filed application must be sufficient to comply with the requirements of 35 U.S.C. 112(a) or the first paragraph of pre-AIA 35 U.S.C. 112, except for the best mode requirement. See Transco Products, Inc. v. Performance Contracting, Inc., 38 F.3d 551, 32 USPQ2d 1077 (Fed. Cir. 1994).
The disclosure of the prior-filed application(s), Application No. 17/197,954 and 18/218,319, fail(s) to provide adequate support or enablement in the manner provided by 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph for one or more claims of this application.
Claim 1, and mutatis mutandis claim 15, incorporate embodiments which do not have adequate support or enablement in the specification of the prior-filed application(s). Claim 1 recites “a global machine learning (ML) model,” where the broadest reasonable interpretation of said model is any ML model known in the art which is stored globally and capable of being trained as described therein. In claims 2 and 16, which depend from claims 1 and 15 respectively, applicant further limits the global ML model as being “a particular type of global ML model, and wherein the particular type of global ML model is one of: an audio-based global ML model, a text-based global ML model, or an image-based global ML model”. Based on the principles of claim differentiation, the global ML model described in claim 1 is necessarily broader than any claims which depend thereon. As such, the global ML model, as recited in claims 1 and 15, includes global ML models which are not “an audio-based global ML model, a text-based global ML model, or an image-based global ML model”. Upon review of the prior-filed applications and review of the specification of the instant application, support for a global ML model which is not at least one of “an audio-based global ML model, a text-based global ML model, or an image-based global ML model” was not found. The specification of the as filed parent application does not teach models which are not “an audio-based global ML model, a text-based global ML model, or an image-based global ML model.” The closest support which could be found was paragraph [0080] which states that “Although the method 400 of FIG. 4 is generally described with respect to audio-based models, it should be understood that is for the sake of example and is not meant to be limiting. For instance, the techniques described herein can also be utilized to remote client gradients and transmit the client gradients to a remote system for updating image-based models, text-based models, and/or any other ML model.” Paragraph [0025] is also noteworthy in that it discloses audio-based, image-based, and text based models as possible embodiments for the global ML model. However, no further model types are indicated. Therefore, claims 1 and 15 incorporate new matter and the application is not entitled to the benefit of an earlier filing date under 35 U.S.C. 120.
As these claim limitations are part of the instant application, as filed, said limitations are entitled to the priority as of the filing date of the instant application. However, said limitations are not entitled to the priority date of the prior-filed application(s).
Applicant is required to amend the claims such that the claims are fully supported in the disclosure of the prior-filed application(s) or modify the relationship of the application to the prior-filed application(s) to that of a continuation-in-part.
Appropriate correction is required.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP §§ 706.02(l)(1) - 706.02(l)(3) for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
Claims 1, 11, and 20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1 and 21 of U.S. Patent No. 11,749,261. Although the claims at issue are not identical, they are not patentably distinct from each other because the claims of the issued patent are narrower in scope than that of the instant application. Therefore, the claims of the issued patent anticipate the claims of the instant application. Please see the below mapping with respect to the independent claims.
Instant Application: 18/917,696
U.S. Pat. No: 11,749,261
Claim 1
A remote system comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to be operable to
Claim 21
A remote system comprising: one or more hardware processors; and memory storing instructions that, when executed by the one or more hardware processors, cause the one or more hardware processor to
: receive a plurality of client gradients from a plurality of corresponding client devices
: receive a plurality of client gradients from a plurality of corresponding client devices
, wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices
, wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices
and based on processing corresponding client data that is accessible locally at the given one of the plurality of client devices
based on processing corresponding audio data that captures at least part of a corresponding spoken utterance of a corresponding user of the given one of the plurality of corresponding client devices (as the client gradients are generated locally at a given one of the plurality of corresponding client devices, it must also be at least accessible locally at the same location, at the time of “processing the corresponding client data”)
; generate a plurality of remote gradients
; generate a plurality of remote gradients
, wherein the instructions to generate each of the plurality of remote gradients comprise instructions to
, wherein the instructions to generate each of the plurality of remote gradients comprise instructions to
: obtain remote data that is accessible remotely at the remote system
: obtain additional audio data that captures at least part of an additional spoken utterance of an additional user
; process, using a global machine learning (ML) model stored remotely at the remote system, the remote data to generate predicted output
; process, using a global machine learning (ML) model stored remotely at the remote system, the additional audio data to generate predicted output
; and generate an additional gradient, for inclusion in the plurality of remote gradients
; and generate an additional gradient, for inclusion in the plurality of remote gradients
, based on comparing the additional predicted output to ground truth output corresponding to the remote data
, based on comparing the additional predicted output to ground truth output corresponding to the additional audio data
; select a set of client gradients from among the plurality of client gradients
; select an additional set of remote gradients from among the plurality of remote gradients
; and utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model
; and utilize the set of client gradients and the additional set of remote gradients to update weights of the global ML model
Claim 11
The system of claim 1,
See limitations of claim 21 above
wherein the at least one processor is further operable to: select a set of client gradients from among the plurality of client gradients
; select a set of client gradients from among the plurality of client gradients
; select an additional set of remote gradients from among the plurality of remote gradients
; select an additional set of remote gradients from among the plurality of remote gradients
; and utilizing the set of client gradients, as the plurality of client gradients, and the additional set of remote gradients, as the plurality of remote gradients, to update weights of the global ML model
; and utilize the set of client gradients and the additional set of remote gradients to update weights of the global ML model
Claim 20
A method implemented by one or more processors of a remote system, the method comprising:
A method implemented by one or more processors, the method comprising:
receiving a plurality of client gradients from a plurality of corresponding client devices,
receiving a plurality of client gradients from a plurality of corresponding client devices
wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices
, wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices
and based on processing corresponding client data that is accessible locally at the given one of the plurality of client devices
based on processing corresponding audio data that captures at least part of a corresponding spoken utterance of a corresponding user of the given one of the plurality of corresponding client devices (as the client gradients are generated locally at a given one of the plurality of corresponding client devices, it must also be at least accessible locally at the same location, at the time of “processing the corresponding client data”)
; generating a plurality of remote gradients
;generating a plurality of remote gradients
, wherein generating each of the plurality of remote gradients comprises: obtaining remote data that is accessible remotely at the remote system
, wherein generating each of the plurality of remote gradients comprises: obtaining additional audio data that captures at least part of an additional spoken utterance of an additional user
; processing, using a global machine learning (ML) model stored remotely at the remote system, the remote data to generate predicted output
;processing, using a global machine learning (ML) model stored remotely at the remote system, the additional audio data to generate predicted output
; and generating a remote gradient, for inclusion in the plurality of remote gradients,
; and generating an additional gradient, for inclusion in the plurality of remote gradients
based on comparing the additional predicted output to ground truth output corresponding to the remote data
, based on comparing the additional predicted output to ground truth output corresponding to the additional audio data
; selecting a set of client gradients from among the plurality of client gradients
; selecting an additional set of remote gradients from among the plurality of remote gradients
; and utilizing the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model.
; and utilizing the set of client gradients and the additional set of remote gradients to update weights of the global ML model.
Claims 2-10, and 12-14 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 21 of U.S. 11,749,261 in view of combinations of Liu (U.S. Pat. App. Pub. No. 2022/0261626, hereinafter Liu), Szeto (U.S. Pat. App. Pub. No. 2022/0405644, hereinafter Szeto), Non-Patent Literature to Bonawitz (Bonawitz, K., Eichner, H., Grieskamp, W., Huba, D., Ingerman, A., Ivanov, V., Kiddon, C., Konečný, J., Mazzocchi, S., McMahan, B. and Van Overveldt, T., 2019. Towards federated learning at scale: System design. Proceedings of machine learning and systems, 1, pp.374-388, hereinafter Bonawitz), and Zhang (U.S. Pat. App. Pub. No. 2020/0175422, hereinafter Zhang). The claims of the co-pending application match that of the independent claims of the instant application, but does not teach the limitations of the dependent claims in the combinations as presented herein. However, the each of the recited limitations are taught in the respective cited reference(s), as indicated in the mapping below. It would have been obvious to one or ordinary skilled in the art to have modified the issued patent with the cited references based on the provided motivations presented in the mapping below. To avoid unnecessary repetition, applicant is directed to the claims as mapped in the rejections under 35 USC 102 and 35 USC 103 below, for the itemized mapping and explanations for each claim. Said mapping is incorporated here by reference.
Instant Application: 18/917,696
U.S. Pat. No: 11,749,261
Claim 1
A remote system comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to be operable to
Claim 21
A remote system comprising: one or more hardware processors; and memory storing instructions that, when executed by the one or more hardware processors, cause the one or more hardware processor to
: receive a plurality of client gradients from a plurality of corresponding client devices
: receive a plurality of client gradients from a plurality of corresponding client devices
, wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices
, wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices
and based on processing corresponding client data that is accessible locally at the given one of the plurality of client devices
based on processing corresponding audio data that captures at least part of a corresponding spoken utterance of a corresponding user of the given one of the plurality of corresponding client devices (as the client gradients are generated locally at a given one of the plurality of corresponding client devices, it must also be at least accessible locally at the same location, at the time of “processing the corresponding client data”)
; generate a plurality of remote gradients
; generate a plurality of remote gradients
, wherein the instructions to generate each of the plurality of remote gradients comprise instructions to
, wherein the instructions to generate each of the plurality of remote gradients comprise instructions to
: obtain remote data that is accessible remotely at the remote system
: obtain additional audio data that captures at least part of an additional spoken utterance of an additional user
; process, using a global machine learning (ML) model stored remotely at the remote system, the remote data to generate predicted output
; process, using a global machine learning (ML) model stored remotely at the remote system, the additional audio data to generate predicted output
; and generate an additional gradient, for inclusion in the plurality of remote gradients
; and generate an additional gradient, for inclusion in the plurality of remote gradients
, based on comparing the additional predicted output to ground truth output corresponding to the remote data
, based on comparing the additional predicted output to ground truth output corresponding to the additional audio data
; select a set of client gradients from among the plurality of client gradients
; select an additional set of remote gradients from among the plurality of remote gradients
; and utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model
; and utilize the set of client gradients and the additional set of remote gradients to update weights of the global ML model
Claim 2
The system of claim 1,
See Mapping of claim 2
wherein the global ML model is a particular type of global ML model
, and wherein the particular type of global ML model is one of: an audio-based global ML model, a text-based global ML model, or an image-based global ML model
Claim 3
The system of claim 2,
See Mapping of Claim 3
wherein a type of the remote data that is obtained is based on a type of the plurality of client gradients received from the plurality of client devices
, and wherein the type of the plurality of client gradients received from the plurality of client devices is one of: audio-based gradients for updating an audio-based global ML model, text-based gradients for updating a text-based global ML model, or image-based gradients for updating an image-based global ML model
Claim 4
The system of claim 3,
See Mapping of Claim 4
wherein the at least one processor is further operable to: in response to receiving, as the plurality of client gradients from the plurality of corresponding client devices, the audio-based gradients for updating the audio-based global ML model
: obtain, as the remote data, remote audio data that is accessible remotely at the remote system
Claim 5
The system of claim 3,
See Mapping of Claim 5
wherein the at least one processor is further operable to: in response to receiving, as the plurality of client gradients from the plurality of corresponding client devices, the text-based gradients for updating the text-based global ML model
: obtain, as the remote data, remote textual data that is accessible remotely at the remote system
Claim 6
The system of claim 3,
See Mapping of Claim 6
wherein the at least one processor is further operable to: in response to receiving, as the plurality of client gradients from the plurality of corresponding client devices, the image-based gradients for updating the image-based global ML model
: obtain, as the remote data, remote image data that is accessible remotely at the remote system
Claim 7
The system of claim 1,
See Mapping of Claim 7
wherein the at least one processor is further operable to: subsequent to utilizing the plurality of client gradients and the plurality of remote gradients to update the weights of the global ML model
: transmit, to one or more of the plurality of corresponding client devices, the updated global ML model or the updated global weights of the global ML model.
Claim 8
The system of claim 7,
See Mapping of Claim 8
wherein transmitting the updated global ML model or the updated weights of the global ML model to each of the plurality of corresponding client devices causes each of the plurality of corresponding client devices to replace, in the corresponding local storage, the global ML model with the updated global ML model
or the weights of the global ML model with the updated weights of the updated global ML model.
Claim 9
The system of claim 1,
See Mapping of Claim 9
wherein the at least one processor is further operable to: prior to receiving the plurality of client gradients from the plurality of corresponding client devices: train the global ML model
; and transmit, to each of the plurality of corresponding client devices, the global ML model or the weights of the global ML model.
Claim 10
The system of claim 9,
See Mapping of Claim 10
wherein transmitting the global ML model or the weights of the global ML model to each of the plurality of corresponding client devices causes each of the plurality of corresponding client devices to store, in corresponding local storage, the global ML model or the weights of the global ML model.
Claim 11
The system of claim 1,
See limitations of claim 21 above
wherein the at least one processor is further operable to: select a set of client gradients from among the plurality of client gradients
; select a set of client gradients from among the plurality of client gradients
; select an additional set of remote gradients from among the plurality of remote gradients
; select an additional set of remote gradients from among the plurality of remote gradients
; and utilizing the set of client gradients, as the plurality of client gradients, and the additional set of remote gradients, as the plurality of remote gradients, to update weights of the global ML model
; and utilize the set of client gradients and the additional set of remote gradients to update weights of the global ML model
Claim 12
The system of claim 1,
See Mapping of Claim 12
wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of client gradients to update the weights of the global ML model
; and subsequent to utilizing the plurality of client gradients to update the weights of the global ML model
: utilize the plurality of remote gradients to further update the weights of the global ML model
Claim 13
The system of claim 1,
See Mapping of Claim 13
wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of remote gradients to update the weights of the global ML model
; and subsequent to utilizing the plurality of remote gradients to update the weights of the global ML model
: utilize the plurality of client gradients to further update the weights of the global ML model
Claim 14
The system of claim 1,
See Mapping of Claim 14
wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of client gradients to update first weights of a first instance the global ML model
; utilize, in parallel, the plurality of remote gradients to update to update second weights of a second instance of the global ML model
; and utilize the updated first weights of the first instance of the global ML model and the updated second weights of the second instance of the global ML model to update the weights of the global ML model.
Claim 20
A method implemented by one or more processors of a remote system, the method comprising:
A method implemented by one or more processors, the method comprising:
receiving a plurality of client gradients from a plurality of corresponding client devices,
receiving a plurality of client gradients from a plurality of corresponding client devices
wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices
, wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices
and based on processing corresponding client data that is accessible locally at the given one of the plurality of client devices
based on processing corresponding audio data that captures at least part of a corresponding spoken utterance of a corresponding user of the given one of the plurality of corresponding client devices (as the client gradients are generated locally at a given one of the plurality of corresponding client devices, it must also be at least accessible locally at the same location, at the time of “processing the corresponding client data”)
; generating a plurality of remote gradients
;generating a plurality of remote gradients
, wherein generating each of the plurality of remote gradients comprises: obtaining remote data that is accessible remotely at the remote system
, wherein generating each of the plurality of remote gradients comprises: obtaining additional audio data that captures at least part of an additional spoken utterance of an additional user
; processing, using a global machine learning (ML) model stored remotely at the remote system, the remote data to generate predicted output
;processing, using a global machine learning (ML) model stored remotely at the remote system, the additional audio data to generate predicted output
; and generating a remote gradient, for inclusion in the plurality of remote gradients,
; and generating an additional gradient, for inclusion in the plurality of remote gradients
based on comparing the additional predicted output to ground truth output corresponding to the remote data
, based on comparing the additional predicted output to ground truth output corresponding to the additional audio data
; selecting a set of client gradients from among the plurality of client gradients
; selecting an additional set of remote gradients from among the plurality of remote gradients
; and utilizing the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model.
; and utilizing the set of client gradients and the additional set of remote gradients to update weights of the global ML model.
Claim Objections
Claim 14 is objected to because of the following informalities:
Regarding claim 14, the phrase “...instance the global ML model” should read as “...instance of the global ML model”.
Further regarding claim 14, the phrase “…to update to update” should read as “…to update
Appropriate correction is required.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 15, and 20 is/are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Liu.
Regarding claim 1, Liu discloses A remote system comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to be operable to (The systems, devices and methods described with reference to the “parameter-server model” for “distributed learning”, described with respect to “M distributed computing machines” and a “server”, as implemented using “computer system 2310” including “a processor device 2320, a network interface 2325, a memory 2330” where the processor device in combination with the memory “can be configured to implement the methods, steps, and functions disclosed herein”; Liu, ¶ [0058]-[0059], [0171]-[0172]): receive a plurality of client gradients from a plurality of corresponding client devices (“each distributed computing machine M {a plurality of corresponding client devices} transmits the (optionally compressed) gradient of the local cost function fi, computed in step 204 {a plurality of client gradients}, to the server.”; Liu, ¶ [0062]), wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices (“Using the adversarial perturbation-modified training examples obtained in step 202, in step 204 each distributed computing machine M then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally on each distributed computing machine M.”; Liu, ¶ [0060]) and based on processing corresponding client data that is accessible locally at the given one of the plurality of client devices (The “adversarial perturbation-modified training examples” are derived from “the samples in the local dataset D(i)” which exists for each of the “M distributed computing machines (i.e., distributed workers)”; Liu, ¶ [0058]-[0059]); generate a plurality of remote gradients (Also discloses “a server (e.g., one of the distributed workers could perform the role of the server), which collects local information (e.g., individual gradients of a local cost function) from the other distributed workers to update the parameters θ of a deep neural network-based model.” As the server may be a distributed worker, the server can also calculate a “gradient of a local cost function” where the gradient, as generated by the server-worker performs numerous iteration in processing the local dataset, where each iteration by the server-worker produces a gradient based on a calculated loss {a plurality of remote gradients}.; Liu, ¶ [0058], [0070]), wherein the instructions to generate each of the plurality of remote gradients comprise instructions to: obtain remote data that is accessible remotely at the remote system (Discloses the “M distributed computing machines (i.e., distributed workers) each of which has access to a local dataset D(i)” where the server is “one of the distributed workers”, and where the “local dataset D(i)” for the server is remote data that is accessible remotely at the server {remote system}; Liu, ¶ [0053], [0058]); process, using a global machine learning (ML) model stored remotely at the remote system, the remote data to generate predicted output (Discloses “labels {generate... output} predicted by the model θ {using a global ML model)” for each of “M distributed computing machines (i.e., distributed workers)”, where the server is a “distributed worker,” using “local dataset D(i)”; Liu, ¶ [0058]-[0059]); and generate an additional gradient, for inclusion in the plurality of remote gradients (the server “then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally” where the server’s local gradient is included in the gradients of “each distributed computing machine M”; Liu, ¶ [0060]), based on comparing the additional predicted output to ground truth output corresponding to the remote data (Discloses “supervised adversarial training and semi-supervised adversarial training,” where at least supervised adversarial training relies on comparison of a predicted output to a ground truth for determination of loss and generation of a gradient.; Liu, ¶ [0053], [0060]); and utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model (Discloses that the “server aggregates the gradients of the local cost function fi received from the individual distributed computing machines M” which includes the gradients produced at the server-worker, and “the parameters θ of the deep neural network-based model” are updated.; Liu, ¶ [0062]-[0063]).
Regarding claim 15, Liu discloses A system comprising: by one or more client device processors of a client device (The systems, devices and methods described with reference to the “parameter-server model” for “distributed learning”, described with respect to “M distributed computing machines” and a “server”, as implemented using “computer system 2310” including “a processor device 2320, a network interface 2325, a memory 2330” where the processor device in combination with the memory “can be configured to implement the methods, steps, and functions disclosed herein”; Liu, ¶ [0058]-[0059], [0171]-[0172]): obtain client data that is accessible locally at the client device (The “adversarial perturbation-modified training examples” are derived from “the samples in the local dataset D(i)” which exists for each of the “M distributed computing machines (i.e., distributed workers)”; Liu, ¶ [0058]-[0059]); process, using an on-device machine learning (ML) model stored locally on the client device, the client data to generate predicted output (Discloses “labels {generate... output} predicted by the model θ {using a global ML model)” for each of “M distributed computing machines (i.e., distributed workers)” using “local dataset D(i)”; Liu, ¶ [0058]-[0059]); generate a client gradient based on the predicted output (“Using the adversarial perturbation-modified training examples obtained in step 202, in step 204 each distributed computing machine M then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally on each distributed computing machine M.”; Liu, ¶ [0060]); and transmit, to a remote system and from the client device, the client gradient (“each distributed computing machine M {a plurality of corresponding client devices} transmits the (optionally compressed) gradient of the local cost function fi, computed in step 204 {a plurality of client gradients}, to the server.”; Liu, ¶ [0062]); by one or more remote system processors of the remote system: obtain remote data that is accessible remotely at the remote system (Discloses the “M distributed computing machines (i.e., distributed workers) each of which has access to a local dataset D(i)” where the server is “one of the distributed workers”, and where the “local dataset D(i)” for the server is remote data that is accessible remotely at the server {remote system}; Liu, ¶ [0053], [0058]); process, using a global ML model stored remotely at the remote system and that is a global counterpart of the on-device ML model, the remote data to generate additional predicted output (Discloses “labels {generate... output} predicted by the model θ {using a global ML model)” for each of “M distributed computing machines (i.e., distributed workers)”, where the server is a “distributed worker,” using “local dataset D(i)”; Liu, ¶ [0058]-[0059]); generate a remote gradient based on the additional predicted output (the server “then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally” where the server’s local gradient is included in the gradients of “each distributed computing machine M”; Liu, ¶ [0060]); and utilize at least the client gradient and the remote gradient to update weights of the global ML model (Discloses that the “server aggregates the gradients of the local cost function fi received from the individual distributed computing machines M” which includes the gradients produced at the server-worker, and “the parameters θ of the deep neural network-based model” are updated.; Liu, ¶ [0062]-[0063]).
Regarding claim 20, Liu discloses A method implemented by one or more processors of a remote system (The systems, devices and methods described with reference to the “parameter-server model” for “distributed learning”, described with respect to “M distributed computing machines” and a “server”, as implemented using “computer system 2310” including “a processor device 2320, a network interface 2325, a memory 2330” where the processor device in combination with the memory “can be configured to implement the methods, steps, and functions disclosed herein”; Liu, ¶ [0058]-[0059], [0171]-[0172]), the method comprising: receiving a plurality of client gradients from a plurality of corresponding client devices, (“each distributed computing machine M {a plurality of corresponding client devices} transmits the (optionally compressed) gradient of the local cost function fi, computed in step 204 {a plurality of client gradients}, to the server.”; Liu, ¶ [0062]) wherein each of the plurality of client gradients is generated locally at a given one of the plurality of corresponding client devices (“Using the adversarial perturbation-modified training examples obtained in step 202, in step 204 each distributed computing machine M then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally on each distributed computing machine M.”; Liu, ¶ [0060]) and based on processing corresponding client data that is accessible locally at the given one of the plurality of client devices (The “adversarial perturbation-modified training examples” are derived from “the samples in the local dataset D(i)” which exists for each of the “M distributed computing machines (i.e., distributed workers)”; Liu, ¶ [0058]-[0059]); generating a plurality of remote gradients, (Also discloses “a server (e.g., one of the distributed workers could perform the role of the server), which collects local information (e.g., individual gradients of a local cost function) from the other distributed workers to update the parameters θ of a deep neural network-based model.” As the server may be a distributed worker, the server can also calculate a “gradient of a local cost function” where the gradient, as generated by the server-worker performs numerous iteration in processing the local dataset, where each iteration by the server-worker produces a gradient based on a calculated loss {a plurality of remote gradients}.; Liu, ¶ [0058], [0070]) wherein generating each of the plurality of remote gradients comprises: obtaining remote data that is accessible remotely at the remote system (Discloses the “M distributed computing machines (i.e., distributed workers) each of which has access to a local dataset D(i)” where the server is “one of the distributed workers”, and where the “local dataset D(i)” for the server is remote data that is accessible remotely at the server {remote system}; Liu, ¶ [0053], [0058]); processing, using a global machine learning (ML) model stored remotely at the remote system, the remote data to generate predicted output (Discloses “labels {generate... output} predicted by the model θ {using a global ML model)” for each of “M distributed computing machines (i.e., distributed workers)”, where the server is a “distributed worker,” using “local dataset D(i)”; Liu, ¶ [0058]-[0059]); and generating a remote gradient, for inclusion in the plurality of remote gradients, (the server “then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally” where the server’s local gradient is included in the gradients of “each distributed computing machine M”; Liu, ¶ [0060]) based on comparing the additional predicted output to ground truth output corresponding to the remote data (Discloses “supervised adversarial training and semi-supervised adversarial training,” where at least supervised adversarial training relies on comparison of a predicted output to a ground truth for determination of loss and generation of a gradient.; Liu, ¶ [0053], [0060]); and utilizing the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model (Discloses that the “server aggregates the gradients of the local cost function fi received from the individual distributed computing machines M” which includes the gradients produced at the server-worker, and “the parameters θ of the deep neural network-based model” are updated.; Liu, ¶ [0062]-[0063]).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-6, 12-14, and 16-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu as applied to claim(s) 1 and 15 above, and further in view of Szeto.
Regarding claim 2, the rejection of claim 1 is incorporated. Liu discloses all of the elements of the current invention as stated above. Liu further discloses wherein the global ML model is a particular type of global ML model (The deep neural network model {global machine learning model} is necessarily a particular type of neural network based model, as particular type includes all types of neural network based models {machine learning models}.; Liu, ¶ [0046]), and wherein the particular type of global ML model is... an image-based global ML model (In one example, the “deep neural network” is “for image classification”; Liu, ¶ [0046]). However, Liu fail(s) to expressly recite wherein the particular type of global ML model is one of: an audio-based global ML model... [or] a text-based global ML model.
Szeto teaches systems and method for distributed machine learning using proxy data. (Szeto, ¶ [0018]). Regarding claim 2, Szeto teaches wherein the particular type of global ML model is one of: an audio-based global ML model... [or] a text-based global ML model (“Machine learning algorithms 295 can include … a neural network algorithm” where the “training may be supervised, semi-supervised, or unsupervised” and modalities of the analyzed data can include “audio data” and “text data”; Szeto, ¶ [0069])..
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the particular type of global ML model is one of: an audio-based global ML model... [or] a text-based global ML model. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 3, the rejection of claim 2 is incorporated. Liu discloses all of the elements of the current invention as stated above. Liu further discloses wherein a type of the remote data that is obtained is based on a type of the plurality of client gradients received from the plurality of client devices (The deep neural network model {global machine learning model} is necessarily a particular type of neural network based model, as particular type includes all types of neural network based models {machine learning models}.; Liu, ¶ [0046]), and wherein the type of the plurality of client gradients received from the plurality of client devices is... image-based gradients for updating an image-based global ML model (In one example, the “deep neural network” is “for image classification”; Liu, ¶ [0046]). However, Liu fail(s) to expressly recite wherein the type of the plurality of client gradients received from the plurality of client devices is one of: audio-based gradients for updating an audio-based global ML model…[or] text-based gradients for updating a text-based global ML model.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 3, Szeto teaches wherein the type of the plurality of client gradients received from the plurality of client devices is one of: audio-based gradients for updating an audio-based global ML model…[or] text-based gradients for updating a text-based global ML model (“Machine learning algorithms 295 can include … a neural network algorithm” where the “training may be supervised, semi-supervised, or unsupervised” and modalities of the analyzed data can include “audio data” and “text data”, where supervised. As read in combination with the deep neural network based model gradients described in Liu, the combination teaches “audio...modality” based gradient {audio-based gradients} for updating an audio-based “deep neural network based model” {global ML model}, and a “text...modality” based gradient {text-based gradients} for updating a text-based “deep neural network based model” {global ML model}; Szeto, ¶ [0069]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the type of the plurality of client gradients received from the plurality of client devices is one of: audio-based gradients for updating an audio-based global ML model…[or] text-based gradients for updating a text-based global ML model. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 4, the rejection of claim 3 is incorporated. Liu discloses all of the elements of the current invention as stated above. Liu further discloses wherein the at least one processor is further operable to:… [after] receiving, as the plurality of client gradients from the plurality of corresponding client devices, the audio-based gradients for updating the audio-based global ML model (Discloses, at “step 204 each distributed computing machine M then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally on each distributed computing machine M,” which as understood in the context of the audio modality described in Szeto is an audio based gradient for updating the audio-based “deep neural network-based model” of Liu.; Liu, ¶ [0053], [0060]): obtain, as the remote data, remote audio data that is accessible remotely at the remote system (As indicated with respect to claim 1, the “M distributed computing machines (i.e., distributed workers)” each have “access to a local dataset D(i)” where the server is “one of the distributed workers”, and where the “local dataset D(i)” for the server is remote data that is accessible remotely at the server {remote system}. As read in the context of the audio modality of Szeto, the “local dataset D(i)” for the server is remote audio data which is accessible remotely at the remote system.; Liu, ¶ [0058]). However, Liu fail(s) to expressly recite wherein the above steps are performed in response to receiving the audio based gradients.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 4, Szeto teaches wherein the at least one processor is further operable to: in response to receiving, as the plurality of client gradients from the plurality of corresponding client devices, the audio-based gradients for updating the audio-based global ML model (“the salient private data features are transmitted to a global modeling engine that aggregates such salient features from many private entities” where the “global modeling engine receives the salient private data features and locally re-instantiates the private data distributions in memory”, where the salient private data features “can... include the parameters of the trained actual model that is trained on the actual private data” and “the global modeling engine generates proxy data from the salient private data features, for example, by using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data”. In the context of Liu, the salient private data features are understood as at least including the respective client gradients.; Szeto, ¶ [0045], [0074], [0115]): obtain, as the remote data, remote audio data that is accessible remotely at the remote system (the system includes “using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data {obtain…remote…data that is accessible remotely at the remote system}” where said data can include “audio data” and/or “text data”; Szeto, ¶ [0069], [0115]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the above steps are performed in response to receiving the audio based gradients. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 5, the rejection of claim 3 is incorporated. Liu discloses all of the elements of the current invention as stated above. Liu further discloses wherein the at least one processor is further operable to:… [after] receiving, as the plurality of client gradients from the plurality of corresponding client devices, the text-based gradients for updating the text-based global ML model (Discloses, at “step 204 each distributed computing machine M then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally on each distributed computing machine M,” which as understood in the context of the text modality described in Szeto is a text based gradient for updating the text-based “deep neural network-based model” of Liu. Further, in the context of a later iteration, all outputs from a previous iteration must be received prior to the next iteration {in response to...}; Liu, ¶ [0053], [0060]): obtain, as the remote data, remote textual data that is accessible remotely at the remote system (As indicated with respect to claim 1, the “M distributed computing machines (i.e., distributed workers)” each have “access to a local dataset D(i)” where the server is “one of the distributed workers”, and where the “local dataset D(i)” for the server is remote data that is accessible remotely at the server {remote system}. As read in the context of the “text...modality” of Szeto, the “local dataset D(i)” for the server is remote text data which is accessible remotely at the remote system.; Liu, ¶ [0058]). However, Liu fail(s) to expressly recite wherein the above steps are performed in response to receiving the text based gradients.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 5, Szeto teaches wherein the at least one processor is further operable to: in response to receiving, as the plurality of client gradients from the plurality of corresponding client devices, the text-based gradients for updating the text-based global ML model (“the salient private data features are transmitted to a global modeling engine that aggregates such salient features from many private entities” where, in one example, the “global modeling engine receives the salient private data features and locally re-instantiates the private data distributions in memory”, where the salient private data features “can... include the parameters of the trained actual model that is trained on the actual private data” and “the global modeling engine generates proxy data from the salient private data features”. In the context of Liu, the salient private data features are understood as at least including the respective client gradients.; Szeto, ¶ [0045], [0074], [0113]-[0115]): obtain, as the remote data, remote textual data that is accessible remotely at the remote system (the system can further include “using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data {obtain…remote…data that is accessible remotely at the remote system}” where said data can include “audio data” and/or “text data”; Szeto, ¶ [0069], [0115]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the above steps are performed in response to receiving the text based gradients. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 6, the rejection of claim 3 is incorporated. Liu discloses all of the elements of the current invention as stated above. Liu further discloses wherein the at least one processor is further operable to:… [after] receiving, as the plurality of client gradients from the plurality of corresponding client devices, the image-based gradients for updating the image-based global ML model (Discloses, at “step 204 each distributed computing machine M then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally on each distributed computing machine M,” which as understood in the context of “image classification”, is an image based gradient for updating the image-based “deep neural network-based model” of Liu.; Liu, ¶ [0046], [0053], [0060]): obtain, as the remote data, remote image data that is accessible remotely at the remote system (As indicated with respect to claim 1, the “M distributed computing machines (i.e., distributed workers)” each have “access to a local dataset D(i)” where the server is “one of the distributed workers”, and where the “local dataset D(i)” for the server is remote data that is accessible remotely at the server {remote system}. As read in the context of “image classification”, the “local dataset D(i)” for the server is remote image data which is accessible remotely at the remote system.; Liu, ¶ [0046], [0058]). However, Liu fail(s) to expressly recite wherein the above steps are performed in response to receiving the image based gradients.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 6, Szeto teaches wherein the at least one processor is further operable to: in response to receiving, as the plurality of client gradients from the plurality of corresponding client devices, the image-based gradients for updating the image-based global ML model (“the salient private data features are transmitted to a global modeling engine that aggregates such salient features from many private entities” where the “global modeling engine receives the salient private data features and locally re-instantiates the private data distributions in memory”, where the salient private data features “can... include the parameters of the trained actual model that is trained on the actual private data” and “the global modeling engine generates proxy data from the salient private data features”. In the context of Liu, the salient private data features are understood as at least including the respective client gradients.; Szeto, ¶ [0045], [0074], [0115]): obtain, as the remote data, remote image data that is accessible remotely at the remote system (the system includes “using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data {obtain…remote…data that is accessible remotely at the remote system}” where said data can include the image data, as described in Liu.; Szeto, ¶ [0069], [0115]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the above steps are performed in response to receiving the image based gradients. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 12, the rejection of claim 1 is incorporated. Liu discloses all of the elements of the current invention as stated above. However, Liu fail(s) to expressly recite wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of client gradients to update the weights of the global ML model; and subsequent to utilizing the plurality of client gradients to update the weights of the global ML model: utilize the plurality of remote gradients to further update the weights of the global ML model.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 12, Szeto teaches wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of client gradients to update the weights of the global ML model (“Operation 510 begins by configuring a private data server operating as a modeling engine to receive model instructions (e.g., from a private data server 124 or from central/global server 130) to create a trained actual model 240 {a first instance of the global ML model} from at least some local private data and according to an implementation of at least one machine learning algorithm,” where the private data servers are the client devices and the local private data is the plurality of client gradients, and the system includes “training the implementation of the machine learning algorithm on the local private data.”; Szeto, ¶ [0099]-[0100], [0112]); and subsequent to utilizing the plurality of client gradients to update the weights of the global ML model: utilize the plurality of remote gradients to further update the weights of the global ML model (Operation 510 is performed for each of the “private data servers,” where, in the context of Liu, the private data servers can include both the workers and the server-worker. The plurality of remote gradients in Liu corresponds to the private data generated by the server-worker. As such, the server-worker also includes “configuring a private data server operating as a modeling engine to receive model instructions (e.g., from a private data server 124 or from central/global server 130) to create a trained actual model 240 {a second instance of the global ML model} from at least some local private data {the plurality of remote gradients}.” Further, as disclosed, “In some embodiments, the trained actual model 240” as maintained at each of the private data servers, “is updated in real-time, on a daily, weekly, bimonthly, monthly, quarterly, or annual basis” and “as soon as new private data is available, it can be incorporated into the trained actual models and the trained proxy models.” Thus, the performance of the local updates in real time using local data can occur in any order (including the training with remote gradients subsequent to client gradients). As each of these can be immediately incorporated into the “trained proxy model {global ML model}”, this results in updating the weights of the “trained proxy model {global ML model}” first with the plurality of client gradients and then subsequently with the plurality of remote gradients.; Szeto, ¶ [0048], [0069], [0099]-[0100], [0112])..
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of client gradients to update the weights of the global ML model; and subsequent to utilizing the plurality of client gradients to update the weights of the global ML model: utilize the plurality of remote gradients to further update the weights of the global ML model. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 13, the rejection of claim 1 is incorporated. Liu discloses all of the elements of the current invention as stated above. However, Liu fail(s) to expressly recite wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of remote gradients to update the weights of the global ML model; and subsequent to utilizing the plurality of remote gradients to update the weights of the global ML model: utilize the plurality of client gradients to further update the weights of the global ML model.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 13, Szeto teaches wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of remote gradients to update the weights of the global ML model (“Operation 510 begins by configuring a private data server operating as a modeling engine to receive model instructions (e.g., from a private data server 124 or from central/global server 130) to create a trained actual model 240 {a first instance of the global ML model} from at least some local private data and according to an implementation of at least one machine learning algorithm,” where the private data servers are the client devices and the local private data is the plurality of client gradients, and the system includes “training the implementation of the machine learning algorithm on the local private data.”; Szeto, ¶ [0099]-[0100], [0112]); and subsequent to utilizing the plurality of remote gradients to update the weights of the global ML model: utilize the plurality of client gradients to further update the weights of the global ML model (Operation 510 is performed for each of the “private data servers,” where, in the context of Liu, the private data servers can include both the workers and the server-worker. The plurality of remote gradients in Liu corresponds to the private data generated by the server-worker. As such, the server-worker also includes “configuring a private data server operating as a modeling engine to receive model instructions (e.g., from a private data server 124 or from central/global server 130) to create a trained actual model 240 {a second instance of the global ML model} from at least some local private data {the plurality of remote gradients}.” Further, as disclosed, “In some embodiments, the trained actual model 240” as maintained at each of the private data servers, “is updated in real-time, on a daily, weekly, bimonthly, monthly, quarterly, or annual basis” and “as soon as new private data is available, it can be incorporated into the trained actual models and the trained proxy models.” Thus, the performance of the local updates in real time using local data can occur in any order (including the training with client gradients subsequent to remote gradients). As each of these can be immediately incorporated into the “trained proxy model {global ML model}” as they are received, this results in updating the weights of the “trained proxy model {global ML model}” first with the plurality of remote gradients and then subsequently with the plurality of client gradients.; Szeto, ¶ [0048], [0069], [0099]-[0100], [0112]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of remote gradients to update the weights of the global ML model; and subsequent to utilizing the plurality of remote gradients to update the weights of the global ML model: utilize the plurality of client gradients to further update the weights of the global ML model. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 14, the rejection of claim 1 is incorporated. Liu discloses all of the elements of the current invention as stated above. Liu further discloses utilize the updated first weights of the first instance of the global ML model and the updated second weights of the second instance of the global ML model to update the weights of the global ML model (Discloses that the “server aggregates the gradients of the local cost function fi received from the individual distributed computing machines M” which includes the gradients produced at the server-worker, and “the parameters θ of the deep neural network-based model” are updated; Liu, ¶ [0062]-[0063]). However, Liu fail(s) to expressly recite wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of client gradients to update first weights of a first instance the global ML model; [and] utilize, in parallel, the plurality of remote gradients to update to update second weights of a second instance of the global ML model.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 14, Szeto teaches wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of client gradients to update first weights of a first instance the global ML model (“Operation 510 begins by configuring a private data server operating as a modeling engine to receive model instructions (e.g., from a private data server 124 or from central/global server 130) to create a trained actual model 240 {a first instance of the global ML model} from at least some local private data and according to an implementation of at least one machine learning algorithm,” where the private data servers are the client devices and the local private data is the plurality of client gradients, and the system includes “training the implementation of the machine learning algorithm on the local private data.”; Szeto, ¶ [0099]-[0100], [0112]); utilize, in parallel, the plurality of remote gradients to update to update second weights of a second instance of the global ML model (Operation 510 is performed for each of the “private data servers,” where, in the context of Liu, the private data servers can include both the workers and the server-worker. The plurality of remote gradients in Liu corresponds to the private data generated by the server-worker. As such, the server-worker also includes “configuring a private data server operating as a modeling engine to receive model instructions (e.g., from a private data server 124 or from central/global server 130) to create a trained actual model 240 {a second instance of the global ML model} from at least some local private data {the plurality of remote gradients}” where the performance of the local updates are not dependent on one another, and thus can occur in any order (including in parallel); Szeto, ¶ [0099]-[0100], [0112]); and utilize the updated first weights of the first instance of the global ML model and the updated second weights of the second instance of the global ML model to update the weights of the global ML model (“the salient private data features are transmitted to a global modeling engine that aggregates such salient features from many private entities” where the salient private data features “can... include the parameters of the trained actual model that is trained on the actual private data” and “the global modeling engine generates proxy data from the salient private data features” where “proxy data 360” can include “actual model parameters, or other information related to private data 322 {the updated first weights... and the updated second weights}” and “the global modeling engine creates a trained proxy model from the set of proxy data by training the same type of or implementation of the machine learning algorithm used to create the trained actual model. “; Szeto, ¶ [0088], [0113]-[0115], [0117]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the instructions to utilize the plurality of client gradients and the plurality of remote gradients to update weights of the global ML model comprise instructions to: utilize the plurality of client gradients to update first weights of a first instance the global ML model; [and] utilize, in parallel, the plurality of remote gradients to update to update second weights of a second instance of the global ML model. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 16, the rejection of claim 15 is incorporated. Liu discloses all of the elements of the current invention as stated above. Liu further discloses wherein the global ML model is a particular type of global ML model (The deep neural network model {global machine learning model} is necessarily a particular type of neural network based model, as particular type includes all types of neural network based models {machine learning models}.; Liu, ¶ [0046]), and wherein the particular type of global ML model is... an image-based global ML model (In one example, the “deep neural network” is “for image classification”; Liu, ¶ [0046]). However, Liu fail(s) to expressly recite wherein the particular type of global ML model is one of: an audio-based global ML model... [or] a text-based global ML model.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 16, Szeto teaches wherein the particular type of global ML model is one of: an audio-based global ML model... [or] a text-based global ML model (“Machine learning algorithms 295 can include … a neural network algorithm” where the “training may be supervised, semi-supervised, or unsupervised” and modalities of the analyzed data can include “audio data” and “text data”; Szeto, ¶ [0069]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the particular type of global ML model is one of: an audio-based global ML model... [or] a text-based global ML model. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 17, the rejection of claim 16 is incorporated. Liu discloses all of the elements of the current invention as stated above. Liu further discloses wherein a type of the remote data that is obtained is based on a type of the plurality of client gradients received from the plurality of client devices (The deep neural network model {global machine learning model} is necessarily a particular type of neural network based model, as particular type includes all types of neural network based models {machine learning models}.; Liu, ¶ [0046]), and wherein the type of the plurality of client gradients received from the plurality of client devices is... image-based gradients for updating an image-based global ML model (In one example, the “deep neural network” is “for image classification”; Liu, ¶ [0046]). However, Liu fail(s) to expressly recite wherein the type of the plurality of client gradients received from the plurality of client devices is one of: audio-based gradients for updating an audio-based global ML model…[or] text-based gradients for updating a text-based global ML model.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 17, Szeto teaches wherein the type of the plurality of client gradients received from the plurality of client devices is one of: audio-based gradients for updating an audio-based global ML model…[or] text-based gradients for updating a text-based global ML model (“Machine learning algorithms 295 can include … a neural network algorithm” where the “training may be supervised, semi-supervised, or unsupervised” and modalities of the analyzed data can include “audio data” and “text data”, where supervised. As read in combination with the deep neural network based model gradients described in Liu, the combination teaches “audio...modality” based gradient {audio-based gradients} for updating an audio-based “deep neural network based model” {global ML model}, and a “text...modality” based gradient {text-based gradients} for updating a text-based “deep neural network based model” {global ML model}; Szeto, ¶ [0069]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the type of the plurality of client gradients received from the plurality of client devices is one of: audio-based gradients for updating an audio-based global ML model…[or] text-based gradients for updating a text-based global ML model. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Regarding claim 18, the rejection of claim 17 is incorporated. Liu discloses all of the elements of the current invention as stated above. Liu further discloses wherein the at least one processor is further operable to:… [after] receiving, as the plurality of client gradients from the plurality of corresponding client devices, the audio-based gradients for updating the audio-based global ML model (Discloses, at “step 204 each distributed computing machine M then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally on each distributed computing machine M,” which as understood in the context of the audio modality described in Szeto is an audio based gradient for updating the audio-based “deep neural network-based model” of Liu.; Liu, ¶ [0053], [0060]): obtain, as the remote data, remote audio data that is accessible remotely at the remote system (As indicated with respect to claim 1, the “M distributed computing machines (i.e., distributed workers)” each have “access to a local dataset D(i)” where the server is “one of the distributed workers”, and where the “local dataset D(i)” for the server is remote data that is accessible remotely at the server {remote system}. As read in the context of the audio modality of Szeto, the “local dataset D(i)” for the server is remote audio data which is accessible remotely at the remote system.; Liu, ¶ [0058])… [after] receiving, as the plurality of client gradients from the plurality of corresponding client devices, the text-based gradients for updating the text-based global ML model (Discloses, at “step 204 each distributed computing machine M then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally on each distributed computing machine M,” which as understood in the context of the text modality described in Szeto is a text based gradient for updating the text-based “deep neural network-based model” of Liu. Further, in the context of a later iteration, all outputs from a previous iteration must be received prior to the next iteration {in response to...}; Liu, ¶ [0053], [0060]): obtain, as the remote data, remote textual data that is accessible remotely at the remote system (As indicated with respect to claim 1, the “M distributed computing machines (i.e., distributed workers)” each have “access to a local dataset D(i)” where the server is “one of the distributed workers”, and where the “local dataset D(i)” for the server is remote data that is accessible remotely at the server {remote system}. As read in the context of the “text...modality” of Szeto, the “local dataset D(i)” for the server is remote text data which is accessible remotely at the remote system.; Liu, ¶ [0058]); and… [after] receiving, as the plurality of client gradients from the plurality of corresponding client devices, the image-based gradients for updating the image-based global ML model (Discloses, at “step 204 each distributed computing machine M then computes a (local) gradient of the local cost function fi (in Equation 2) with respect to the parameters θ of the deep neural network-based model stored locally on each distributed computing machine M,” which as understood in the context of “image classification”, is an image based gradient for updating the image-based “deep neural network-based model” of Liu.; Liu, ¶ [0046], [0053], [0060]): obtain, as the remote data, remote image data that is accessible remotely at the remote system (As indicated with respect to claim 1, the “M distributed computing machines (i.e., distributed workers)” each have “access to a local dataset D(i)” where the server is “one of the distributed workers”, and where the “local dataset D(i)” for the server is remote data that is accessible remotely at the server {remote system}. As read in the context of “image classification”, the “local dataset D(i)” for the server is remote image data which is accessible remotely at the remote system.; Liu, ¶ [0046], [0058]). However, Liu fail(s) to expressly recite wherein the above steps are performed in response to receiving the text based gradients, the audio based gradients or the image based gradients.
The relevance of Szeto is described above with relation to claim 2. Regarding claim 18, Szeto teaches wherein the at least one processor is further operable to: in response to receiving, as the plurality of client gradients from the plurality of corresponding client devices, the audio-based gradients for updating the audio-based global ML model (“the salient private data features are transmitted to a global modeling engine that aggregates such salient features from many private entities” where the “global modeling engine receives the salient private data features and locally re-instantiates the private data distributions in memory”, where the salient private data features “can... include the parameters of the trained actual model that is trained on the actual private data” and “the global modeling engine generates proxy data from the salient private data features, for example, by using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data”. In the context of Liu, the salient private data features are understood as at least including the respective client gradients.; Szeto, ¶ [0045], [0074], [0115]): obtain, as the remote data, remote audio data that is accessible remotely at the remote system (the system includes “using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data {obtain…remote…data that is accessible remotely at the remote system}” where said data can include “audio data” and/or “text data”; Szeto, ¶ [0069], [0115]); in response to receiving, as the plurality of client gradients from the plurality of corresponding client devices, the text-based gradients for updating the text-based global ML model (“the salient private data features are transmitted to a global modeling engine that aggregates such salient features from many private entities” where the “global modeling engine receives the salient private data features and locally re-instantiates the private data distributions in memory”, where the salient private data features “can... include the parameters of the trained actual model that is trained on the actual private data” and “the global modeling engine generates proxy data from the salient private data features, for example, by using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data”. In the context of Liu, the salient private data features are understood as at least including the respective client gradients.; Szeto, ¶ [0045], [0074], [0115]): obtain, as the remote data, remote textual data that is accessible remotely at the remote system (the system includes “using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data {obtain…remote…data that is accessible remotely at the remote system}” where said data can include “audio data” and/or “text data”; Szeto, ¶ [0069], [0115]); and in response to receiving, as the plurality of client gradients from the plurality of corresponding client devices, the image-based gradients for updating the image-based global ML model (“the salient private data features are transmitted to a global modeling engine that aggregates such salient features from many private entities” where, in one example, the “global modeling engine receives the salient private data features and locally re-instantiates the private data distributions in memory”, where the salient private data features “can... include the parameters of the trained actual model that is trained on the actual private data” and “the global modeling engine generates proxy data from the salient private data features”. In the context of Liu, the salient private data features are understood as at least including the respective client gradients.; Szeto, ¶ [0045], [0074], [0113]-[0115]): obtain, as the remote data, remote image data that is accessible remotely at the remote system (the system can further include “using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data {obtain…remote…data that is accessible remotely at the remote system}” where said data can include “audio data” and/or “text data”; Szeto, ¶ [0069], [0115]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Szeto to include wherein the above steps are performed in response to receiving the text based gradients, the audio based gradients or the image based gradients. The “generation of proxy data” for the training of a centralized model, as described in Szeto, for a wide variety of data formats as known in the art, such as audio data or text data, provides “a synthetic equivalent of actual data” which can be more compact in size, parameterized, and obfuscated, while reducing knowledge lost during the de-identification process of prior art distributed learning models, which provides the known benefits of increased quality training data and reduced transmission costs, as recognized by Szeto. (Szeto, ¶ [0077]).
Claims 7-10 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu as applied to claim(s) 1 and 15 above, and further in view of Bonawitz.
Regarding claim 7, the rejection of claim 1 is incorporated. Liu disclose all of the elements of the current invention as stated above. Liu further discloses wherein the at least one processor is further operable to:... utilizing the plurality of client gradients and the plurality of remote gradients to update the weights of the global ML model (“ in step 208, each distributed computing machine M transmits the (optionally compressed) gradient of the local cost function fi, computed in step 204, to the server. The server aggregates the gradients of the local cost function fi received from the individual distributed computing machines M” and “then transmits an aggregated gradient back to the distributed computing machines M.”; Liu, ¶ [0062])However, Liu fail(s) to expressly recite wherein the at least one processor is further operable to: subsequent to utilizing the plurality of client gradients and the plurality of remote gradients to update the weights of the global ML model: transmit, to one or more of the plurality of corresponding client devices, the updated global ML model or the updated global weights of the global ML model.
Bonawitz teaches “a scalable production system for Federated Learning in the domain of mobile devices, based on TensorFlow.” (Bonawitz, ¶ Abstract). Regarding claim 7, Bonawitz teaches wherein the at least one processor is further operable to: subsequent to utilizing the plurality of client gradients and the plurality of remote gradients to update the weights of the global ML model (“Our system enables one to train a deep neural network, using TensorFlow (Abadi et al., 2016), on data stored on the phone which will never leave the device. The weights are combined in the cloud with Federated Averaging, constructing a global model”; Bonawitz, ¶ pg. 1, col. 2, lines 11-20): transmit, to one or more of the plurality of corresponding client devices, the updated global ML model or the updated global weights of the global ML model (The global model is then “pushed back to phones for inference,” which is the transmission the updated global ML model or the updated global weights of the global ML model to the phones {the plurality of corresponding client devices}; Bonawitz, ¶ pg. 1, col. 2, lines 11-20).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Bonawitz to include wherein the at least one processor is further operable to: subsequent to utilizing the plurality of client gradients and the plurality of remote gradients to update the weights of the global ML model: transmit, to one or more of the plurality of corresponding client devices, the updated global ML model or the updated global weights of the global ML model. Bonawitz discloses a federated learning system including secure aggregation, which overcomes “numerous practical issues” including “device availability that correlates with the local data distribution in complex ways (e.g., time zone dependency); unreliable device connectivity and interrupted execution; orchestration of lock-step execution across devices with varying availability; and limited device storage and compute resources,” while maintaining the security of the transmission itself, which addresses practicality problems with implementation of the system described in Liu, which would be known to one skilled in the art while maintaining or improving privacy of the data, as recognized by Bonawitz. (Bonawitz, ¶ pg. 1, col. 2, lines 11-32, pg. 6, col. 1, lines 32-44).
Regarding claim 8, the rejection of claim 7 is incorporated. Liu disclose all of the elements of the current invention as stated above. However, Liu fail(s) to expressly recite wherein transmitting the updated global ML model or the updated weights of the global ML model to each of the plurality of corresponding client devices causes each of the plurality of corresponding client devices to replace, in the corresponding local storage, the global ML model with the updated global ML model or the weights of the global ML model with the updated weights of the updated global ML model.
The relevance of Bonawitz is described above with relation to claim 7. Regarding claim 8, Bonawitz teaches wherein transmitting the updated global ML model or the updated weights of the global ML model to each of the plurality of corresponding client devices causes each of the plurality of corresponding client devices to replace, in the corresponding local storage, the global ML model with the updated global ML model or the weights of the global ML model with the updated weights of the updated global ML model (the global model is then “pushed back to phones for inference,” where “pushed back” is with reference to transmission of the global model to the phones {updated global ML model}, and where “for inference” means both replacing the current model for the purposes of inference with the updated model in the corresponding local storage at the phone {client device}, and use of updated model for inference.; Bonawitz, ¶ pg. 1, col. 2, lines 11-20).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Bonawitz to include wherein transmitting the updated global ML model or the updated weights of the global ML model to each of the plurality of corresponding client devices causes each of the plurality of corresponding client devices to replace, in the corresponding local storage, the global ML model with the updated global ML model or the weights of the global ML model with the updated weights of the updated global ML model. Bonawitz discloses a federated learning system including secure aggregation, which overcomes “numerous practical issues” including “device availability that correlates with the local data distribution in complex ways (e.g., time zone dependency); unreliable device connectivity and interrupted execution; orchestration of lock-step execution across devices with varying availability; and limited device storage and compute resources,” while maintaining the security of the transmission itself, which addresses practicality problems with implementation of the system described in Liu, which would be known to one skilled in the art while maintaining or improving privacy of the data, as recognized by Bonawitz. (Bonawitz, ¶ pg. 1, col. 2, lines 11-32, pg. 6, col. 1, lines 32-44).
Regarding claim 9, the rejection of claim 1 is incorporated. Liu disclose all of the elements of the current invention as stated above. Liu further discloses wherein the at least one processor is further operable to: prior to receiving the plurality of client gradients from the plurality of corresponding client devices: train the global ML model (Discloses the use of “distributed adversarial training pre-trained model” where pre-training indicates that the model is trained prior to receiving gradients of the local cost function {the plurality of client gradients} from the workers {corresponding client devices}; Liu, ¶ [0091]). However, Liu fail(s) to expressly recite transmit, to each of the plurality of corresponding client devices, the global ML model or the weights of the global ML model.
The relevance of Bonawitz is described above with relation to claim 7. Regarding claim 9, Bonawitz teaches transmit, to each of the plurality of corresponding client devices, the global ML model or the weights of the global ML model (After being trained, the global model is then “pushed back to phones for inference,” which is the transmission the global ML model or the global weights of the global ML model to the phones {the plurality of corresponding client devices}; Bonawitz, ¶ pg. 1, col. 2, lines 11-20).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Bonawitz to include transmit, to each of the plurality of corresponding client devices, the global ML model or the weights of the global ML model. Bonawitz discloses a federated learning system including secure aggregation, which overcomes “numerous practical issues” including “device availability that correlates with the local data distribution in complex ways (e.g., time zone dependency); unreliable device connectivity and interrupted execution; orchestration of lock-step execution across devices with varying availability; and limited device storage and compute resources,” while maintaining the security of the transmission itself, which addresses practicality problems with implementation of the system described in Liu, which would be known to one skilled in the art while maintaining or improving privacy of the data, as recognized by Bonawitz. (Bonawitz, ¶ pg. 1, col. 2, lines 11-32, pg. 6, col. 1, lines 32-44).
Regarding claim 10, the rejection of claim 9 is incorporated. Liu disclose all of the elements of the current invention as stated above. However, Liu fail(s) to expressly recite wherein transmitting the global ML model or the weights of the global ML model to each of the plurality of corresponding client devices causes each of the plurality of corresponding client devices to store, in corresponding local storage, the global ML model or the weights of the global ML model.
The relevance of Bonawitz is described above with relation to claim 7. Regarding claim 10, Bonawitz teaches wherein transmitting the global ML model or the weights of the global ML model to each of the plurality of corresponding client devices causes each of the plurality of corresponding client devices to store, in corresponding local storage, the global ML model or the weights of the global ML model (the global model is then “pushed back to phones for inference,” where “pushed back” is with reference to transmission of the global model to the phones {updated global ML model}, and where “for inference” means at least storage of the updated model in the corresponding local storage at the phone {client device}, and use of updated model for inference.; Bonawitz, ¶ pg. 1, col. 2, lines 11-20).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Bonawitz to include wherein transmitting the global ML model or the weights of the global ML model to each of the plurality of corresponding client devices causes each of the plurality of corresponding client devices to store, in corresponding local storage, the global ML model or the weights of the global ML model. Bonawitz discloses a federated learning system including secure aggregation, which overcomes “numerous practical issues” including “device availability that correlates with the local data distribution in complex ways (e.g., time zone dependency); unreliable device connectivity and interrupted execution; orchestration of lock-step execution across devices with varying availability; and limited device storage and compute resources,” while maintaining the security of the transmission itself, which addresses practicality problems with implementation of the system described in Liu, which would be known to one skilled in the art while maintaining or improving privacy of the data, as recognized by Bonawitz. (Bonawitz, ¶ pg. 1, col. 2, lines 11-32, pg. 6, col. 1, lines 32-44).
Regarding claim 19, the rejection of claim 15 is incorporated. Liu disclose all of the elements of the current invention as stated above. Liu further discloses wherein the at least one processor is further operable to: prior to receiving the plurality of client gradients from the plurality of corresponding client devices: train the global ML model (Discloses the use of “distributed adversarial training pre-trained model” where pre-training indicates that the model is trained prior to receiving gradients of the local cost function {the plurality of client gradients} from the workers {corresponding client devices}; Liu, ¶ [0091]). However, Liu fail(s) to expressly recite transmit, to each of the plurality of corresponding client devices, the global ML model or the weights of the global ML mode; and subsequent to utilizing the plurality of client gradients and the plurality of remote gradients to update the weights of the global ML model: transmit, to one or more of the plurality of corresponding client devices, the updated global ML model or the updated global weights of the global ML model.
The relevance of Bonawitz is described above with relation to claim 7. Regarding claim 19, Bonawitz teaches transmit, to each of the plurality of corresponding client devices, the global ML model or the weights of the global ML model (After being trained, the global model is then “pushed back to phones for inference,” which is the transmission the global ML model or the global weights of the global ML model to the phones {the plurality of corresponding client devices}; Bonawitz, ¶ pg. 1, col. 2, lines 11-20); and subsequent to utilizing the plurality of client gradients and the plurality of remote gradients to update the weights of the global ML model (“Our system enables one to train a deep neural network, using TensorFlow (Abadi et al., 2016), on data stored on the phone which will never leave the device. The weights are combined in the cloud with Federated Averaging, constructing a global model” which is understood in the context of the gradients of Liu as the training of a global model “in the cloud” {at the server} using the gradients generated both at the server-worker and the client devices to update the weights.; Bonawitz, ¶ pg. 1, col. 2, lines 11-20): transmit, to one or more of the plurality of corresponding client devices, the updated global ML model or the updated global weights of the global ML model (The global model is then “pushed back to phones for inference,” which is the transmission the updated global ML model or the updated global weights of the global ML model to the phones {the plurality of corresponding client devices}; Bonawitz, ¶ pg. 1, col. 2, lines 11-20).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Bonawitz to include transmit, to each of the plurality of corresponding client devices, the global ML model or the weights of the global ML mode; and subsequent to utilizing the plurality of client gradients and the plurality of remote gradients to update the weights of the global ML model: transmit, to one or more of the plurality of corresponding client devices, the updated global ML model or the updated global weights of the global ML model. Bonawitz discloses a federated learning system including secure aggregation, which overcomes “numerous practical issues” including “device availability that correlates with the local data distribution in complex ways (e.g., time zone dependency); unreliable device connectivity and interrupted execution; orchestration of lock-step execution across devices with varying availability; and limited device storage and compute resources,” while maintaining the security of the transmission itself, which addresses practicality problems with implementation of the system described in Liu, which would be known to one skilled in the art while maintaining or improving privacy of the data, as recognized by Bonawitz. (Bonawitz, ¶ pg. 1, col. 2, lines 11-32, pg. 6, col. 1, lines 32-44).
Claim 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu as applied to claim 1 above, and further in view of Zhang.
Regarding claim 11, the rejection of claim 1 is incorporated. Liu disclose all of the elements of the current invention as stated above. However, Liu fail(s) to expressly recite wherein the at least one processor is further operable to: select a set of client gradients from among the plurality of client gradients; select an additional set of remote gradients from among the plurality of remote gradients; and utilizing the set of client gradients, as the plurality of client gradients, and the additional set of remote gradients, as the plurality of remote gradients, to update weights of the global ML model.
Zhang teaches “asynchronous gradient weight compression in a distributed machine learning system.” (Zhang, ¶ [0001]). Regarding claim 11, Zhang teaches wherein the at least one processor is further operable to: select a set of client gradients from among the plurality of client gradients (“gradient weight compression system 102 can send (e.g., via transmit component 502 and/or network 116) a windowed concatenated compressed gradient weight to each polling learner 114a, 114b, 114N” where the “compression component 108 can compute an updated model gradient weight... that can constitute a windowed concatenated compressed gradient weight” where “such a windowed concatenated compressed gradient weight” is computed “using only the one or more compressed gradient weights that can be identified by pointer component 110 as being not present in a previously computed model gradient weight {selecting a set of client gradients}” where the “gradient weights... not present in a previously computed model gradient weight” are selected from the group of compressed gradient weights {from among a plurality of client gradients}.; Zhang, ¶ [0097]-[0098]); select an additional set of remote gradients from among the plurality of remote gradients (In the context of the server-worker of Liu, the “windowed concatenated compressed gradient weight” comprising the “gradient weights... not present in a previously computed model gradient weight” includes those produced by all workers, where the “gradient weights... not present in a previously computed model gradient weight” produced by the server-worker are the additional set of remote gradients selected from among the plurality of remote gradients; Zhang, ¶ [0097]-[0098]); and utilizing the set of client gradients, as the plurality of client gradients, and the additional set of remote gradients, as the plurality of remote gradients, to update weights of the global ML model (“each learner 114a, 114b, 114N” where the “learners 114a, 114b, 114N can comprise a server device” can “unpack compressed gradient weights of the windowed concatenated compressed gradient weight” and “each learner 114a, 114b, 114N can update its weights.”; Zhang, ¶ [0045], [0101]-[0102]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the distributed training techniques of Liu to incorporate the teachings of Zhang to include wherein the at least one processor is further operable to: select a set of client gradients from among the plurality of client gradients; select an additional set of remote gradients from among the plurality of remote gradients; and utilizing the set of client gradients, as the plurality of client gradients, and the additional set of remote gradients, as the plurality of remote gradients, to update weights of the global ML model. Zhang discloses the identification of gradients not present in a set in an asynchronous learner system where “identification and/or computation operations can constitute aggressive backward compression (e.g., from a parameter server to leaner entities), which can be implemented in a distributed machine learning model… to enable computationally inexpensive calculation of concatenated compressed gradient weights of such a distributed machine learning model… where such concatenated compressed gradient weights can be transferred via reduced computational costs (e.g., reduced by a factor of 32) between a parameter server (e.g., gradient weight compression system 102) and one or more remote learner entities of the distributed machine learning system (e.g., learners 114a, 114b, 114N) without compromising accuracy of model parameters of such a system,” which provides the known benefits of improved efficiency and data availability, as recognized by Zhang. (Zhang, ¶ [0104], [0106]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Peterson (U.S. Pat. App. Pub. No. 2021/0073677) discloses techniques for domain adaptation of a machine learning (ML) model which impose differential privacy onto federated learning by the ML model.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sean E. Serraguard whose telephone number is (313)446-6627. The examiner can normally be reached 07:00-17:00 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel C. Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Sean E Serraguard/Primary Examiner, Art Unit 2657