DETAILED ACTION
1. This is in response to the application No. 19/209,181 filed on 05/15/2025. Claims 1-20 are submitted for examination. Claims 1, 8 and 15 are independent.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
3. This application filed on 05/15/2025 is a Continuation of application No. 18202435, filed on 05/26/202 , now U.S. Patent # 12373598. Application No. 18202435 claims Priority from Provisional Application 63350333, filed 06/08/2022. Application No. 18202435 also claims Priority from Provisional Application 63346812, filed 05/27/2022
Information Disclosure Statement
4. No information disclosure statements (IDS) is submitted for this application.
Drawings
5. The drawings filed on May 15, 2025 are accepted.
Specification
6. The specification filed on May 15, 2025 is also accepted.
Internet Communications
7. Applicant is encouraged to submit a written authorization for Internet communications (PTO/SB/439, http:/www.uspto.gov/sites/default/files/documents/sb0439.pdf) in the instant patent application to authorize the examiner to communicate with the applicant via email. The authorization will allow the examiner to better practice compact prosecution. The written authorization can be submitted via one of the following methods only: (1) Central Fax, which can be found in the Conclusion section of this Office action; (2) regular postal mail; (3) EFS WEB; or (4) the service window on the Alexandria campus. EFS web is the recommended way to submit the form since this allows the form to be entered into the file wrapper within the same day (system dependent). Written authorization submitted via other methods, such as direct fax to the examiner or email, will not be accepted. See MPEP § 502.03.
Claim Rejections - 35 USC § 102
8. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
9. Claims 1-2, 7-9 and 14-16 are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Hugh Brendan McMahan et al (McMahan) (US Publication No. 20190227980 A1) (Pub. Date: July 25, 2019)
The following is referring to the independent claim 1:
As per independent claim 1, McMahan discloses a system for training a computer model, comprising: one or more processors; and a non-transitory computer-readable medium having instructions executable by the one or more processors for [Para. 0004, “The computing system can include one or more server computing devices. The one or more server computing devices can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions”]:
determining a set of gradients by applying the computer model with a set of current model parameter values to training data samples [Para. 0092-0094, “select a batch b of size B from k's examples; return update Δk=ClipFn(−η∇
PNG
media_image1.png
38
17
media_image1.png
Greyscale
(θ; b)); para. 0094 expressly permits sampling users or training examples and para. 0070-0073, repeat the user-update operation for each sampled user. The term, ∇
PNG
media_image1.png
38
17
media_image1.png
Greyscale
(θ; b) is gradient determined by applying the loss of the computer model, at current parameter values θ, to the training example in batch b. Repeating that operation for sampled users or training examples produced the claimed set of gradients The multiplication by −η merely changes the sign and scale of each gradient before clipping]
determining at least one adjusted gradient by, for at least one gradient in the set of gradients [Para. 0077-0079 and 0090-0093 and 0101, : Δk=ClipFn(−η∇
PNG
media_image1.png
38
17
media_image1.png
Greyscale
(θ; b)); where ClipFn returns π(Δ,S) and equation 4 defines π(Δ,S)=def Δ·min (1,S/IIΔII). Note, the input to the clipping function is a gradient multiplied by the learning rate. The output of the clipping function is therefore an adjusted version of that gradient. When the magnitude is reduced; otherwise it preserved.]
setting a scaling factor to one of: a reference bound or a magnitude of the at least one gradient; [Para. 0101 equation 4, π(Δ,S)=def Δ·min (1,S/IIΔII). Algebraically, the same equation is π(Δ,S)= Δ,S/max(S, IIΔII2). Note let the scaling factor be d= max(S, IIΔII2). If the gradient magnitude is no greater than the reference bound, then d=S. If the magnitude exceeds the reference bound, then d= IIΔII2 .Thus, the value selected for d is necessary one of the reference bound or the magnitude of the gradient, the same as it is claimed.]
and
determining an adjusted gradient by adjusting the at least one gradient based on a ratio of a clipping bound to the scaling factor [[Para. 0101 equation 4, π(Δ,S)=def Δ·min (1,S/IIΔII)= Δ,S/max(S, IIΔII2).; Note: when d= max(S, IIΔII2, the adjusted gradient is Δ (S/d). The multiplier S/d corresponds to the limitation ratio of the clipping bound S to the scaling factor d. Claim 1 does not require the reference bound and clipping bound to be different. So S may satisfy both terms.]
determining a model update gradient based on the at least one adjusted gradient and added noise; [Para. 0031, “a bounded-sensitivity data-weighted average “ of the clipped local updates. Para. 0037, “and the bounded-sensitivity data-weighted average of the local updates can include adding a noise component to the bounded-sensitivity data-weighted average of the local updates and “the noise component can be a Gaussian noise” Note: In the stochastic-gradient embodiment, each local updates is a clipped, learning rate scaled gradient. The server combines those adjusted gradients into an average and adds noise. The resulting noisy aggregate is therefore a model-update gradient based on adjusted gradients and added noise.];
updating the current model parameter values based on the model update gradient [Para. 0036-0037, determining the updated machine-learned model from the noisy aggregate. The training pseudocode also gives that…]
As per independent claim 8, Independent claim 8 is a method version of a system claim 1 and has the same scope as independent claim 1. Thus, claim 8 is rejected for same reason as claim 1.
As per independent claim 15, Independent claim 15 is a non-transitory computer readable medium version of a system claim 1 and has the same scope as independent claim 1. Thus, claim 15 is rejected for same reason as claim 1.
As per dependent claim 2, McMahan discloses a system for training a computer model as applied to claim 1 above. Furthermore, McMahan discloses the system wherein , wherein determining the adjusted gradient for the at least one gradient having the magnitude of the gradient higher than the reference bound comprises adjusting the gradient to a magnitude substantially equal to the clipping bound. [Para. 0101 equation 4, π(Δ,S)=def Δ·min (1, S/IIΔII2) and IIΔII2 is greater than S, the output is ΔS/IIΔII2 whose magnitude is exactly S. Note.The “reference bound” and “clipping bound” both read on S. A gradient above S is adjusted to the magnitude S. which is more precise than “substantially equal”]
As per dependent claim 9, dependent claim 9 is a method version of a system claim 2 and has the same scope as dependent claim 2. Thus, claim 9 is rejected for same reason as claim 2.
As per dependent claim 16, dependent claim 16 is a non-transitory computer readable medium version of a system claim 2 and has the same scope as dependent claim 2. Thus, claim 16 is rejected for same reason as claim 2.
As per dependent claim 7, McMahan discloses a system for training a computer model as applied to claim 1 above. Furthermore, McMahan discloses the system, wherein determining the model update gradient based on the at least one adjusted gradient and added noise includes averaging or summing the at least one adjusted gradient.[Para. 0031 teaches the average of the local updates and para. 0033-0034 disclose calculating that average from weighted sum of the local updates]
As per dependent claim 14, dependent claim 14 is a method version of a system claim 7 and has the same scope as dependent claim 7. Thus, claim 14 is rejected for same reason as claim 7.
Claim Rejections - 35 USC § 103
10. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
11. Claims 3-6, 10-13 and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Hugh Brendan McMahan et al (McMahan) (US Publication No. 20190227980 A1) (Pub. Date: July 25, 2019) in view of NPL document titled, “"Differentially Private Learning with Adaptive Clipping” by Galen Andrew et al (Andrew) (Publication Date: Nov. 2021)
As per dependent claim 3, McMahan discloses a system for training a computer model as applied to claim 1 above. McMahan doesn’t explicitly disclose the following underlined or bolded claim limitation: “ modifying the reference bound based on the set of gradients for use of the modified reference bound with another batch of training data samples”
However, Andrew discloses; “modifying the reference bound based on the set of gradients for use of the modified reference bound with another batch of training data samples” [Page 4, 2.1. defines the current bound Ct , computes bt i = I||∆t i||2≤Ct and then calculates Ct+1 =C · exp(−ηC(¯ b− γ)).The next iteration calls the training function using Ct+1. The local function states “local data split into bathes” Note, the current set of update magnitudes is compared Ct . The resulting statistic changes the clipping bound to Ct+1, which is then used in the next training round and its batches]
McMahan and Andrew are analogous/in the same field of endeavor as they both are directed to learning with clipping.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the gradient based embodiment of McMahan by applying adaptive rule such as “modifying the reference bound based on the set of gradients for use of the modified reference bound with another batch of training data samples” as per teachings of Andrew to update the reference bound based on the present set of gradients for another batch or round and provide adaptive clipping outperforming the best fixed clip. [See Andrew, abstract, “The method tracks the quantile closely, uses a negligible amount of privacy budget, is compatible with other federated learning technologies such as compression and secure aggregation, and has a straightforward joint DP analysis with DP-FedAvg. Experiments demonstrate that adaptive clipping to the median update norm works well across a range of realistic federated learning tasks, sometimes outperforming even the best fixed clip chosen in hindsight, and without the need to tune any clipping hyperparameter”]
As per dependent claim 10, dependent claim 10 is a method version of a system claim 3 and has the same scope as dependent claim 3. Thus, claim 10 is rejected for same reason as claim 3.
As per dependent claim 17, dependent claim 17 is a non-transitory computer readable medium version of a system claim 3 and has the same scope as dependent claim 3. Thus, claim 17 is rejected for same reason as claim 3.
As per dependent claim 4, the combination of McMahan and Andrew discloses a system for training a computer model as applied to claim 3 above. Furthermore, Andrew discloses the system, wherein modifying the reference bound includes increasing or decreasing the reference bound based on a number of gradients in the set of gradients having a magnitude above the reference bound [page 4, 2.1, “m be the number of users in a round and let γ ∈ [0,1] denote the target quantile;. Ct be the clipping threshold, and ηC be the learning rate. It defines bt i = I||∆t i||2≤Ct. Therefore the number above the bound is At= m -∑bt i .Note: If too many magnitudes exceed Ct, then bt is less than γ, making the exponent positive and increasing the bound. If two few exceed it, then bt is greater than γ, making the exponent negative and decreasing the bound. The update therefore increases or decreases the bound based on the number above it.]
As per dependent claim 11, dependent claim 11 is a method version of a system claim 4 and has the same scope as dependent claim 4. Thus, claim 11 is rejected for same reason as claim 4.
As per dependent claim 18, dependent claim 18 is a non-transitory computer readable medium version of a system claim 4 and has the same scope as dependent claim 4. Thus, claim 18 is rejected for same reason as claim 4.
As per dependent claim 5, the combination of McMahan and Andrew discloses a system for training a computer model as applied to claim 3 above. Furthermore, Andrew discloses the system, wherein modifying the reference bound includes modifying the reference bound with randomized noise. [Page 4, 2.1, See table for Algorithm 1 Andrew states adding Gaussian noise to the sum and defines bt on the left column as shown on the last equation for bt in algorithm 1, “DPFedAvg-M with adaptive clipping” and followed by the equation for Ct+1 shown on the top, right column on 4, table algorithm 1, “DPFedAvg-M with adaptive clipping” Note: the randomized Gaussian value is incorporated into bt and bt is incorporated into the new bound. The reference bound is therefore modified with randomized noise.]
As per dependent claim 12, dependent claim 12 is a method version of a system claim 5 and has the same scope as dependent claim 5. Thus, claim 12 is rejected for same reason as claim 5.
As per dependent claim 19, dependent claim 19 is a non-transitory computer readable medium version of a system claim 5 and has the same scope as dependent claim 5. Thus, claim 19 is rejected for same reason as claim 5.
As per dependent claim 6, the combination of McMahan and Andrew discloses a system for training a computer model as applied to claim 3 above. Furthermore, Andrew discloses the system wherein modifying the reference bound comprises applying an exponential function based on: a number of gradients in the set of gradients having a magnitude higher than the reference bound by a threshold value; a randomized noise; a number of training data samples in the batch; and a clipping learning rate. [Page 4, 2.1, See table for Algorithm 1 Andrew states adding Gaussian noise to the sum and defines bt on the left column as shown on the last equation for bt in algorithm 1, “DPFedAvg-M with adaptive clipping” and followed by the equation for Ct+1 shown on the top, right column on 4, table algorithm 1, “DPFedAvg-M with adaptive clipping” Note: the randomized Gaussian value is incorporated into bt and bt is incorporated into the new bound. By substitution you can rewrite the equation expressly depends on the number above the bound, the target above the bound threshold and randomized noise; m the number of samples or participants in the batch or round and the clipping learning rate.]
As per dependent claim 13, dependent claim 13 is a method version of a system claim 6 and has the same scope as dependent claim 6. Thus, claim 13 is rejected for same reason as claim 6.
As per dependent claim 20, dependent claim 20 is a non-transitory computer readable medium version of a system claim 6 and has the same scope as dependent claim 6. Thus, claim 20 is rejected for same reason as claim 6.
Conclusion
12. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
a. US Publication No. 2023/0351042 A1 to De discloses methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for privacy-sensitive training of a neural network. In one aspect, a method includes training a set of neural network parameters of the neural network on a set of training data over multiple training iterations to optimize an objective function. Each training iteration includes: sampling a batch of network inputs from the set of training data; determining a clipped gradient for each network input in the batch of network inputs; and updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs.
b. US Publication No. 2022/0231648 A1 to Li discloses techniques for improved machine learning using gradient pruning, comprising computing, using a first batch of training data, a first gradient tensor comprising a gradient for each parameter of a parameter tensor for a machine learning model; identifying a first subset of gradients in the first gradient tensor based on a first gradient criteria; and updating a first subset of parameters in the parameter tensor based on the first subset of gradients in the first gradient tensor.
US Publication No. 2022/0318412 A1 to Guo disclose Privacy-aware pruning in machine learning that provides techniques for improved machine learning using private variational dropout. A set of parameters of a global machine learning model is updated based on a local data set, and the set of parameters is pruned based on pruning criteria. A noise-augmented set of gradients is computed for a subset of parameters remaining after the pruning, based in part on a noise value, and the noise-augmented set of gradients is transmitted to a global model server.
See the other cited prior arts.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAMSON B LEMMA whose telephone number is 571-272-3806. The examiner can normally be reached on M-F 8am-10pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Shaw Yin Chen can be reached on to 571-272-8878. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SAMSON B LEMMA/Primary Examiner, Art Unit 2498