CTNF 18/564,160 CTNF 93309 DETAILED ACTION 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Notice for all US Patent Applications filed on or after March 16, 2013 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Status of the Claims This communication is in response to communications received on 11/27/23. Claim(s) 4-7, 9-12, and 14-20 is/are amended, claim(s) none is/are cancelled, claim(s) none is/are new, and applicant does not provide any information on where support for the amendments can be found in the instant specification. Therefore, Claims 1-20 is/are pending and have been addressed below. Information Disclosure Statement The information disclosure statement(s) (IDS) submitted on 2/15/24, 10/30/24, and 2/20/25 was/were considered by the examiner. Priority 02-27 AIA Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. IL294292 , filed on 06/26/2022 . Response to Arguments There are no arguments. Claim Rejections - 35 USC § 101 Claim(s) 1-20 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter as noted below. The limitation(s) below for representative claim(s) 1, 19, and 20 that, under its broadest reasonable interpretation, is directed to updating a neural network. Step 1 : The claim(s) as drafted, is/are a process (claim(s) 19 recites a series of steps) and system (claim(s) 1-18, 20 recites a series of components). Step 2A – Prong 1 : The claimed invention is directed to an abstract idea without significantly more. The claim(s) recite(s): Claim 19: repeatedly performing, by each of a plurality of worker computing units , operations comprising: obtaining current values of the set of neural network parameters from a central memory ; sampling a batch of network inputs from a set of training data; determining a respective gradient corresponding to each network input, comprising, for each network input: processing the network input using the neural network , in accordance with current values of the set of neural network parameters, to generate a network output; and determining a gradient of an objective function with respect to the set of neural network parameters when the objective function is evaluated on the network output; determining an aggregated gradient based on the gradients corresponding to the network inputs; identifying a proper subset of a set of gradient values included in the aggregated gradient as target gradient values to be combined with random noise; generating a noisy gradient by combining random noise with the target gradient values in the aggregated gradient; and updating the current values of the set of neural network parameters stored in the central memory using the noisy gradient . Claim(s) 1 and 20: same analysis as claim(s) 19. Dependent claims 2-18 recite the same or similar abstract idea(s) as independent claim(s) 1, 19, and 20 with merely a further narrowing of the abstract idea(s): . The identified limitations of the independent and dependent claims above fall well-within the groupings of subject matter identified by the courts as being abstract concepts of: mathematical relationships, mathematical formulas or equations, or mathematical calculations because the invention is directed to the application of mathematical processes as they are associated with updating a neural network. Step 2A – Prong 2 : This judicial exception is not integrated into a practical application because: The additional elements unencompassed by the abstract idea include worker computing units, memory , neural network (claim(s) 1, 19, 20), system (claim(s) 1), one or more computer storage media, one or more computers (claim(s) 20), neural network (claim(s) 10-13, 15-18). The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements as described above with respect to Step 2A Prong 2 fails to describe: Improvements to the functioning of a computer, or to any other technology or technical field - see MPEP 2106.05(a) Applying or using a judicial exception to effect a particular treatment or prophylaxis for a disease or medical condition – see Vanda Memo Applying the judicial exception with, or by use of, a particular machine – see MPEP 2106.05(b) Effecting a transformation or reduction of a particular article to a different state or thing - see MPEP 2106.05(c) Applying or using the judicial exception in some other meaningful way beyond generally linking the use of the judicial exception to a particular technological environment, such that the claim as a whole is more than a drafting effort designed to monopolize the exception - see MPEP 2106.05(e) and Vanda Memo. Thus the additional elements as described above with respect to Step 2A Prong 2 are merely (as additionally noted by instant specification [0117]) invoked as a tool and/or general purpose computer to apply instructions of an abstract idea in a particular technological environment, and/or mere application of an abstract idea in a particular technological environment and merely limiting the use of an abstract idea to a particular technological field do not integrate an abstract idea into a practical application (MPEP 2106.05(f)&(h)). Step 2B : The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus the additional elements as described above with respect to Step 2A Prong 2 are merely (as additionally noted by instant specification [0117]) invoked as a tool and/or a general purpose computer to apply instructions of an abstract idea in a particular technological environment, and/or mere application of an abstract idea in a particular technological environment and merely limiting the use of an abstract idea to a particular technological field do not integrate an abstract idea into a practical application and thus similarly the combination and arrangement of the above identified additional elements when analyzed under Step 2B also fails to necessitate a conclusion that the claims amount to significantly more than the abstract idea for the same reasons as set forth above (MPEP 2106.05(f)&(h)). Claim Rejections - 35 USC § 102 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. 07-12-aia AIA (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 07-15-aia AIA Claim(s) 1, 7-11, 19, and 20 is/are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by De et al. (US 2023/0351042 A1) . Regarding claims 1, 19, and 20 (currently amended) , De teaches a method performed by one or more computers for privacy-sensitive training of a neural network having a set of neural network parameters, the method comprising : {a system for privacy-sensitive training of a neural network having a set of neural network parameters, the system comprising: a central memory that is configured to store current values of the set of neural network parameters; and – claim 1 } {one or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for privacy-sensitive training of a neural network having a set of neural network parameters, the operations comprising the operations comprising: – claim 20 } repeatedly performing, by each of a plurality of worker computing units, operations comprising [the limitation is interpreted based on broadest reasonable interpretation of instant specification [0006], then see at least [0006] “The method includes: training a set of neural network parameters of the neural network on a set of training data over multiple training iterations to optimize an objective function, including, at each training iteration: sampling a batch of network inputs from the set of training data;”; [0131] system 100 process data such as neural network and its components from memory “The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. … The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.”] : obtaining current values of the set of neural network parameters from a central memory [see at least [0085] “Training system 100 processes each augmented network input … using the neural network 110, in accordance with current values of the network parameters w(t) 122.t, to generate a respective network output ŷj 136.j for the augmented network input 134.j.”; [0131] system 100 process data such as neural network and its components from memory “The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. … The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.”; Fig. 1 and [0065, 0077] where current/initial parameters are 112.t and updated parameters are 112.(t+1) after training “Particularly, training system 100 updates the values of the network parameters w(t) 112.t at each training iteration (t) according to a privatized update rule (see examples below) to progressively minimize (or maximize) the objective function 140 with respect to the network parameters 112.” and as noted in Fig 1 to the left of second 110 “Network Parameters 112.tt(+1)” after training; [0066] “The neural network 110 is parametrized by a set of network parameters 112 and is configured to process a network input to generate a network output.”; [0067] “For example, the set of neural network parameters 112 may include weights (and biases) of multiple neural network layers of the neural network 110, e.g., weights of one or more feedforward neural network layers, convolutional weights of one or more convolutional neural network layers, parameter matrices and parameter vectors of one or more recurrent neural network layers, etc.”] ; sampling a batch of network inputs from a set of training data [see at least [0113] “Training system samples a batch of network inputs from the set of training data (220). In some implementations, at each training iteration, the batch of network inputs includes at least 4000 network inputs.”] ; determining a respective gradient corresponding to each network input, comprising, for each network input: processing the network input using the neural network, in accordance with current values of the set of neural network parameters, to generate a network output; and determining a gradient of an objective function with respect to the set of neural network parameters when the objective function is evaluated on the network output; determining an aggregated gradient based on the gradients corresponding to the network inputs; identifying a proper subset of a set of gradient values included in the aggregated gradient as target gradient values to be combined with random noise; generating a noisy gradient by combining random noise with the target gradient values in the aggregated gradient; and updating the current values of the set of neural network parameters stored in the central memory using the noisy gradient [for the limitations above, see at least [0006] “The method includes: training a set of neural network parameters of the neural network on a set of training data over multiple training iterations to optimize an objective function, including, at each training iteration: sampling a batch of network inputs from the set of training data; determining a clipped gradient for each network input in the batch of network inputs, including, for each network input in the batch of network inputs: generating multiple augmented versions of the network input, wherein each augmented version of the network input results from applying a respective augmentation transformation to the network input; determining, for each of the multiple augmented versions of the network input, a gradient of the objective function for the augmented version of the network input; determining a combined gradient for the network input by combining the gradients determined for the multiple augmented versions of the network input; and generating the clipped gradient for the network input by clipping the combined gradient for the network input; and updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs.”; [0010-0011] “In some implementations, for one or more of the network inputs, generating the clipped gradient for the network input includes scaling the combined gradient for the network input to cause a norm of the combined gradient for the network input to satisfy a clipping threshold. In some implementations, scaling the combined gradient for the network input to cause the norm of the combined gradient for the network input to satisfy the clipping threshold includes scaling the combined gradient for the network input by a scaling factor defined as a ratio of: (i) the clipping threshold, and (ii) the norm of the combined gradient for the network input.”; [0012-0013] “In some implementations, the method further includes, before updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs: generating a set of noise parameters, including randomly sampling the noise parameters from a noise distribution; and applying the noise parameters to the clipped gradients for the network inputs in the batch of network inputs. In some implementations, the noise distribution includes a Gaussian noise distribution.”]. Regarding claim 7 (currently amended) , De teaches the system of claim 1, wherein generating the noisy gradient by combining random noise with the target gradient values in the aggregated gradient comprises, for each target gradient value in the aggregated gradient: adding a respective random noise value to the target gradient value [see at least [0012-0013] “In some implementations, the method further includes, before updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs: generating a set of noise parameters, including randomly sampling the noise parameters from a noise distribution; and applying the noise parameters to the clipped gradients for the network inputs in the batch of network inputs. In some implementations, the noise distribution includes a Gaussian noise distribution.”]. Regarding claim 8 , De teaches the system of claim 7, wherein the random noise value is sampled from a Gaussian distribution [see at least [0012-0013] “In some implementations, the method further includes, before updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs: generating a set of noise parameters, including randomly sampling the noise parameters from a noise distribution; and applying the noise parameters to the clipped gradients for the network inputs in the batch of network inputs. In some implementations, the noise distribution includes a Gaussian noise distribution.”]. Regarding claim 9 (currently amended) , De teaches the system of claim 1, wherein determining the aggregated gradient based on the gradients corresponding to the network inputs comprises: generating the aggregated gradient as an average of the gradients corresponding to the network inputs [see at least [0009] “In some implementations, determining the combined gradient for the network input includes averaging the gradients determined for the plurality of augmented versions of the network input.”]. Regarding claim 10 (currently amended) , De teaches the system of claim 1, wherein for each network input, determining the gradient of the objective function with respect to the set of neural network parameters when the objective function is evaluated on the network output comprises: backpropagating the gradient of the objective function through the set of neural network parameters [see at least [0088] “For example, training system 100 can use backpropagation to determine the gradient 142.j for each augmented network input 134.j.”]. Regarding claim 11 (currently amended) , De teaches the system of claim 1 , wherein updating the current values of the set of neural network parameters stored in the central memory using the noisy gradient comprises: updating the current values of the set of neural network parameters using the noisy gradient by a gradient descent update rule [see at least [0099-0100] “Training system 100 updates the neural network parameters w (t+1) 122.(t+1) for the training iteration (t) using the clipped gradients 146.i. For example, training system 100 can implement a first-order optimization technique using the privatized gradient to establish a privatized update rule of the form: … (8) where ηt is the learning rate (or step-size) for the training iteration. The learning rate can be the same or differ between training iterations. For example, training system 100 can use a constant learning rate for each training iteration, e.g., using ηt=η with η<1, or decay the learning rate at each training iteration, … . The privatized update rule of Eq. (8) is a type of stochastic gradient descent (SGD) technique but training system 100 can use a similar privatized update rule in combination with other first-order optimization techniques, such as SGD with momentum or Adam.”] . Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-23-aia AIA The factual inquiries set forth in Graham v. John Deere Co. , 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 07-37-05 It has been held that a prior art reference must either be in the field of applicant’s endeavor or, if not, then be reasonably pertinent to the particular problem with which the applicant was concerned, in order to be relied upon as a basis for rejection of the claimed invention. See In re Oetiker , 977 F.2d 1443, 24 USPQ2d 1443 (Fed. Cir. 1992). 07-21-aia AIA Claim (s) 2-3 is/are rejected under 35 U.S.C. 103 as being unpatentable over De et al. (US 2023/0351042 A1) in view of Han et al. (US 2023/0351189 A1) . Regarding claim 2 , De teaches the system of claim 1, wherein for each network input, determining the gradient corresponding to the network input comprises: clipping the gradient corresponding to the network input based on clipping threshold [see at least [0006] “determining a combined gradient for the network input by combining the gradients determined for the multiple augmented versions of the network input; and generating the clipped gradient for the network input by clipping the combined gradient for the network input; and updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs.”; [0010-0011] “In some implementations, for one or more of the network inputs, generating the clipped gradient for the network input includes scaling the combined gradient for the network input to cause a norm of the combined gradient for the network input to satisfy a clipping threshold. In some implementations, scaling the combined gradient for the network input to cause the norm of the combined gradient for the network input to satisfy the clipping threshold includes scaling the combined gradient for the network input by a scaling factor defined as a ratio of: (i) the clipping threshold, and (ii) the norm of the combined gradient for the network input.”; [0012-0013] “In some implementations, the method further includes, before updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs: generating a set of noise parameters, including randomly sampling the noise parameters from a noise distribution; and applying the noise parameters to the clipped gradients for the network inputs in the batch of network inputs. In some implementations, the noise distribution includes a Gaussian noise distribution.”]. De teaches updating a neural network but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as updating a neural network, Han discloses predefined clipping threshold [see at least [0010] “In the method of training the binarized neural network and the memory device according to example embodiments, when the binarized neural network is trained, the weight set may be updated, and the parameterized weight clipping scheme in which the range of the clipping function applied to the weight set is adaptively and dynamically changed may be used. For example, the range of the clipping function may be changed based on gradient descent in the backpropagation process. Accordingly, a problem caused by gradient mismatch (or missing) in the binarized neural network, e.g., a dead weight problem, may be reduced, the binarized neural network may be efficiently trained with a reduced or minimum decrease in accuracy, and the accuracy of the binarized neural network may be improved or enhanced as compared to a conventional training scheme.”; [0033] “In the method of training the binarized neural network according to example embodiments, a forward propagation process is performed on the binarized neural network using a clipping function (block S100).”; claim 10 and [0036] “The clipping function may represent a function that, when a magnitude (or amplitude) of a specific weight is out of a predetermined range, changes the corresponding weight to a predetermined threshold value corresponding to the predetermined range. The efficient training may be performed by applying, employing, or using the clipping function when the binarized neural network is trained. For example, the clipping function may be applied to a weight set that includes a plurality of weights (or weight elements).”; [0038] “Thereafter, a backpropagation process is performed to update the weight set and to change a range of the clipping function (block S200).”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify De with Han to include the limitation(s) above as disclosed by Han. Doing so would improve De’s (De) neural network updating by further defining the clipping methodology [see at least Han [0003] ]. Furthermore, all of the claimed elements were known in the prior arts of a) De and b) Han and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention. Regarding claim 3 , modified De teaches the system of claim 2, as well as predefined clipping threshold and De teaches wherein for each network input, clipping the gradient corresponding to the network input based on the clipping threshold comprises: scaling the gradient to cause a norm of the gradient to satisfy the clipping threshold [see at least [0006] “determining a combined gradient for the network input by combining the gradients determined for the multiple augmented versions of the network input; and generating the clipped gradient for the network input by clipping the combined gradient for the network input; and updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs.”; [0010-0011] “In some implementations, for one or more of the network inputs, generating the clipped gradient for the network input includes scaling the combined gradient for the network input to cause a norm of the combined gradient for the network input to satisfy a clipping threshold. In some implementations, scaling the combined gradient for the network input to cause the norm of the combined gradient for the network input to satisfy the clipping threshold includes scaling the combined gradient for the network input by a scaling factor defined as a ratio of: (i) the clipping threshold, and (ii) the norm of the combined gradient for the network input.”; [0012-0013] “In some implementations, the method further includes, before updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs: generating a set of noise parameters, including randomly sampling the noise parameters from a noise distribution; and applying the noise parameters to the clipped gradients for the network inputs in the batch of network inputs. In some implementations, the noise distribution includes a Gaussian noise distribution.”] . 07-21-aia AIA Claim (s) 4-6 is/are rejected under 35 U.S.C. 103 as being unpatentable over De et al. (US 2023/0351042 A1) in view of Ergun published December 23, 2021 (reference U on the Notice of References Cited) . Regarding claim 4 (currently amended) , De teaches the system of claim 1, . De teaches aggregated gradients but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as aggregated gradients, Ergun discloses wherein the aggregated gradient is defined by a sparse array of numerical values [see at least [pg 1] “Secure aggregation is a popular protocol in privacy-preserving federated learning, which allows model aggregation without revealing the individual models in the clear. On the other hand, conventional secure aggregation protocols incur a significant communication overhead, which can become a major bottleneck in real- world bandwidth-limited applications. Towards addressing this challenge, in this work we propose a lightweight gradient sparsification framework for secure aggregation, in which the server learns the aggregate of the sparsified local model updates from a large number of users, but without learning the individual parameters. Our theoretical analysis demonstrates that the proposed framework can significantly reduce the communication overhead of secure aggregation while ensuring comparable computational complexity. We further identify a trade-off between privacy and communication efficiency due to sparsification. Our experiments demonstrate that our framework reduces the communication overhead by up to 7.8x, while also speeding up the wall clock training time by 1.13x, when compared to conventional secure aggregation benchmarks.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify De with Ergun to include the limitation(s) above as disclosed by Ergun. Doing so would improve De’s (De) neural network updating by further defining the clipping methodology [see at least Ergun [0003] ]. Furthermore, all of the claimed elements were known in the prior arts of a) De and b) Ergun and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention. Regarding claim 5 (currently amended) , De teaches the system of claim 1, as well as the noisy gradient . De teaches aggregated gradients but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as aggregated gradients, Ergun discloses wherein the gradient is defined by a sparse array of numerical values [see at least [pg 1] “Secure aggregation is a popular protocol in privacy-preserving federated learning, which allows model aggregation without revealing the individual models in the clear. On the other hand, conventional secure aggregation protocols incur a significant communication overhead, which can become a major bottleneck in real-world bandwidth-limited applications. Towards addressing this challenge, in this work we propose a lightweight gradient sparsification framework for secure aggregation, in which the server learns the aggregate of the sparsified local model updates from a large number of users, but without learning the individual parameters. Our theoretical analysis demonstrates that the proposed framework can significantly reduce the communication overhead of secure aggregation while ensuring comparable computational complexity. We further identify a trade-off between privacy and communication efficiency due to sparsification. Our experiments demonstrate that our framework reduces the communication overhead by up to 7.8x, while also speeding up the wall clock training time by 1.13x, when compared to conventional secure aggregation benchmarks.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify De with Ergun to include the limitation(s) above as disclosed by Ergun. Doing so would improve De’s (De) neural network updating by further defining the clipping methodology [see at least Ergun [0003] ]. Furthermore, all of the claimed elements were known in the prior arts of a) modified De and b) Ergun and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention. Regarding claim 6 (currently amended) , De teaches the system of claim 1, wherein identifying the proper subset of the set of gradient values included in the aggregated gradient as target gradient values to be combined with random noise comprises: values in the aggregated gradient; and selecting a gradient value in the aggregated gradient as a target gradient value only if the gradient value is included in the aggregated gradient [for the limitations above, see at least [0006] “determining a combined gradient for the network input by combining the gradients determined for the multiple augmented versions of the network input; and generating the clipped gradient for the network input by clipping the combined gradient for the network input; and updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs.”; [0010-0011] “In some implementations, for one or more of the network inputs, generating the clipped gradient for the network input includes scaling the combined gradient for the network input to cause a norm of the combined gradient for the network input to satisfy a clipping threshold. In some implementations, scaling the combined gradient for the network input to cause the norm of the combined gradient for the network input to satisfy the clipping threshold includes scaling the combined gradient for the network input by a scaling factor defined as a ratio of: (i) the clipping threshold, and (ii) the norm of the combined gradient for the network input.”; [0012-0013, 0095, 0115] “In some implementations, the method further includes, before updating the neural network parameters using the clipped gradients for the network inputs in the batch of network inputs: generating a set of noise parameters, including randomly sampling the noise parameters from a noise distribution; and applying the noise parameters to the clipped gradients for the network inputs in the batch of network inputs. In some implementations, the noise distribution includes a Gaussian noise distribution.”]. De teaches aggregated gradients but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as aggregated gradients, Ergun discloses comprises: identifying a set of non-zero gradient values in the aggregated gradient; and in the set of non-zero gradient values [see at least [pg 1] “Secure aggregation is a popular protocol in privacy-preserving federated learning, which allows model aggregation without revealing the individual models in the clear. On the other hand, conventional secure aggregation protocols incur a significant communication overhead, which can become a major bottleneck in real-world bandwidth-limited applications. Towards addressing this challenge, in this work we propose a lightweight gradient sparsification framework for secure aggregation, in which the server learns the aggregate of the sparsified local model updates from a large number of users, but without learning the individual parameters. Our theoretical analysis demonstrates that the proposed framework can significantly reduce the communication overhead of secure aggregation while ensuring comparable computational complexity. We further identify a trade-off between privacy and communication efficiency due to sparsification. Our experiments demonstrate that our framework reduces the communication overhead by up to 7.8x, while also speeding up the wall clock training time by 1.13x, when compared to conventional secure aggregation benchmarks.”; [pg 9] “C. Sparsified Gradient Construction Using the additive and multiplicative masks, user … constructs a sparsified masked gradient … . More specifically, for each non-zero element in bij for a given j ∈ [N], user I adds the corresponding element from rij to its quantized local gradient yi if i<j, and subtracts it if i>j. The key property of this process is to ensure that once the sparsified masked gradients are aggregated at the server, the pairwise additive masks cancel out.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify De with Ergun to include the limitation(s) above as disclosed by Ergun. Doing so would improve De’s (De) neural network updating by further defining the clipping methodology [see at least Ergun [0003] ]. Furthermore, all of the claimed elements were known in the prior arts of a) modified De and b) Ergun and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention . 07-21-aia AIA Claim (s) 12-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over De et al. (US 2023/0351042 A1) in view of Li et al. (US 2020/0372076 A1) . Regarding claim 12 (currently amended) , De teaches the system of claim 1 , . De teaches updating machine learning but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as updating machine learning, Li discloses wherein the neural network is configured to receive a network input that includes features values of a categorical feature, wherein the set of neural network parameters define a respective embedding corresponding to each possible value of the categorical feature [see at least [0006] “wherein during the training: the machine learning model is configured to process an input that comprises one or more possible categorical feature values of respective categorical features by performing operations comprising: for only those possible categorical feature values included in the input that are specified as active by the output sequence, mapping the possible categorical feature value to a corresponding embedding that is iteratively adjusted during the training;”; [0074] “For example, for each categorical feature, the set of shared parameters corresponding to the categorical feature may be represented as a two-dimensional (2-D) array of numerical values. The number of rows in the array may be equal to the number of possible feature values of the categorical feature, and the number of columns may be equal to the maximum allowable embedding dimensionality for possible values of the categorical feature.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify De with Li to include the limitation(s) above as disclosed by Li. Doing so would improve De’s (De) neural network updating by further defining how the updating occurs [see at least Li [0002-0004] ]. Furthermore, all of the claimed elements were known in the prior arts of a) De and b) Li and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention. Regarding claim 13 , modified De teaches the system of claim 12, . Modified De teaches updating machine learning but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as updating machine learning, Li discloses wherein the neural network comprises an embedding layer that is configured to map each categorical feature value included in the network input to a corresponding embedding [see at least [0006] “wherein during the training: the machine learning model is configured to process an input that comprises one or more possible categorical feature values of respective categorical features by performing operations comprising: for only those possible categorical feature values included in the input that are specified as active by the output sequence, mapping the possible categorical feature value to a corresponding embedding that is iteratively adjusted during the training;”; [0074] “For example, for each categorical feature, the set of shared parameters corresponding to the categorical feature may be represented as a two-dimensional (2-D) array of numerical values. The number of rows in the array may be equal to the number of possible feature values of the categorical feature, and the number of columns may be equal to the maximum allowable embedding dimensionality for possible values of the categorical feature.”; [0071] “In some implementations, the controller neural network 110 may be configured to generate an output that defines both: (i) a categorical feature specification, and (ii) an architecture of the machine learning model. The data defining the architecture of the machine learning model may specify, e.g., the number of neural network layers used by the prediction system of the machine learning model to process the embeddings of the active categorical feature values to generate an output.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify modified De with Li to include the limitation(s) above as disclosed by Li. Doing so would improve modified De’s (De) neural network updating by further defining how the updating occurs [see at least Li [0002-0004] ]. Furthermore, all of the claimed elements were known in the prior arts of a) modified De and b) Li and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention. Regarding claim 14 (currently amended) , De teaches the system of claim 12 , . Modified De teaches updating machine learning but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as updating machine learning, Li discloses wherein the categorical feature has at least 100,000 possible categorical feature values [see at least [0031] “This may result in unacceptable computational resource consumption and poor performance of the machine learning model, e.g., for machine learning models that perform large-scale recommendation tasks (e.g., recommending videos or webpages to users) by processing categorical features having large numbers (e.g., millions or billions) of possible categorical feature values.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify modified De with Li to include the limitation(s) above as disclosed by Li. Doing so would improve modified De’s (De) neural network updating by further defining how the updating occurs [see at least Li [0002-0004] ]. Furthermore, all of the claimed elements were known in the prior arts of a) modified De and b) Li and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention. Regarding claim 15 (currently amended) , De teaches the system of claim 12 , . Modified De teaches updating machine learning but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as updating machine learning, Li discloses wherein the neural network is configured to receive a network input includes feature values of the categorical feature that characterize a previous search query of a user, and the neural network is configured to generate a network output that characterizes a predicted next search query of the user [see at least [0043] “In one example, the machine learning model 102 may be configured to process an input that characterizes a previous textual search query of a user to generate an output that specifies a predicted next search query of the user. The categorical features in the input to the machine learning model may include, e.g.: the previous search query, uni-grams of the previous search query, bi-grams of the previous search query, and tri-grams of the previous search query. A n-gram (e.g., uni-gram, bi-gram, or tri-gram) of a search query refers to a sequence of n consecutive characters in the search query.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify modified De with Li to include the limitation(s) above as disclosed by Li. Doing so would improve modified De’s (De) neural network updating by further defining how the updating occurs [see at least Li [0002-0004] ]. Furthermore, all of the claimed elements were known in the prior arts of a) modified De and b) Li and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention. Regarding claim 16 (currently amended) , De teaches the system of claim 12 , . Modified De teaches updating machine learning but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as updating machine learning, Li discloses wherein the neural network is configured to receive a network input that includes feature values of the categorical feature that characterize previous videos watched by a user, and the neural network is configured to generate a network output that characterizes a predicted next video watched by the user [see at least [0045] “In another example, the machine learning model 102 may be configured to process an input that characterizes previous videos watched by a user to generate an output that characterizes a predicted next video to be watched by the user (e.g., on a video-sharing platform). The categorical features in the input to the machine learning model may include a categorical feature specifying identifiers (IDs) of the previous videos watched by the user, where the possible feature values of the categorical feature include a respective ID corresponding to each of multiple videos.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify modified De with Li to include the limitation(s) above as disclosed by Li. Doing so would improve modified De’s (De) neural network updating by further defining how the updating occurs [see at least Li [0002-0004] ]. Furthermore, all of the claimed elements were known in the prior arts of a) modified De and b) Li and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention. Regarding claim 17 (currently amended) , De teaches the system of claim 12 , . Modified De teaches updating machine learning but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as updating machine learning, Li discloses wherein the neural network is configured to receive a network input that includes feature values of the categorical feature that characterize previous webpages visited by a user, and the neural network is configured to generate a network output that characterizes a predicted next webpage visited by the user [see at least [0046] “In another example, the machine learning model 102 may be configured to process an input that characterizes previous webpages visited by a user to generate an output that characterizes a predicted next webpage to be visited by the user. The categorical features in the input to the machine learning model may include a categorical feature specifying IDs of previous websites visited by the user, where the possible feature values of the categorical feature include a respective ID corresponding to each of multiple webpages.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify modified De with Li to include the limitation(s) above as disclosed by Li. Doing so would improve modified De’s (De) neural network updating by further defining how the updating occurs [see at least Li [0002-0004] ]. Furthermore, all of the claimed elements were known in the prior arts of a) modified De and b) Li and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention. Regarding claim 18 (currently amended) , De teaches the system of claim 12 , . Modified De teaches updating machine learning but doesn’t/don’t explicitly teach however, in the field pertinent to the particular problem with which the applicant was concerned such as updating machine learning, Li discloses wherein the neural network is configured to receive a network input that includes feature values of the categorical feature that characterizes previous products associated with a user, and the neural network is configured to generate a network output that characterizes a predicted next product associated with the user [see at least [0047] “In another example, the machine learning model 102 may be configured to process an input that characterizes products associated with a user, e.g., products that were previously purchased by the user, or products that the user previously viewed on an online platform, to generate an output that characterizes other products that may be of interest to the user. The categorical features in the input to the machine learning model may include a categorical feature specifying IDs of products associated with the user, where the possible feature values of the categorical feature include a respective ID corresponding to each of multiple products. The output of the machine learning model may include a respective score for each product in a set of multiple products, where the score for each product characterizes a likelihood that the product is of interest to the user (e.g., should be recommended to the user).”]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify modified De with Li to include the limitation(s) above as disclosed by Li. Doing so would improve modified De’s (De) neural network updating by further defining how the updating occurs [see at least Li [0002-0004] ]. Furthermore, all of the claimed elements were known in the prior arts of a) modified De and b) Li and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention . Conclusion When responding to the office action, any new claims and/or limitations should be accompanied by a reference as to where the new claims and/or limitations are supported in the original disclosure, see MPEP § 2163.04 (I)(B). 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. De et al. – WO 2023/209192 A1 (relevant because it teaches same as US2023/0351042 A1) Choi et al. – US 2023/0169333 A1 (relevant because it teaches [0056] “In parallel processing the neural network 120 through the distributed training, gradient clipping may be performed to stably update the neural network 120 and achieve fast convergence. Gradient clipping is a scheme of reducing an exploding gradient issue by clipping a gradient for a gradient value to not exceed a predetermined value in an error backpropagation process. The gradient clipping may be performed based on a threshold, which may refer to, for example, if a calculated gradient is greater than a threshold, clipping the calculated gradient according to the threshold so that a maximum value of the gradient is limited to the threshold. In an example, the value of the gradient for which the gradient clipping is performed may be adjusted to the threshold. There may be a lower threshold and/or an upper threshold for a threshold for gradient clipping. Here, the lower threshold may be referred to as a minimum clip value, and the upper threshold may be referred to as a maximum clip value. The gradient clipping may include gradient scaling.”) Ding et al. – Privacy-Preserving Feature Extraction via Adversarial Training (relevant because it teaches “Specifically, we use an intermediate layer to separate the entire neural network into two parts, which are respectively deployed on the user device and the cloud server. … Therefore, we also propose an approach to achieve Privacy-preserving Feature Extraction based on Adversarial Training (P-FEAT), where the goal of privacy attacking tasks and the goal of target tasks are adversarial in terms of sensitive attributes. By imposing privacy constraints during the feature extraction, we can reduce the contribution of the extracted features to the privacy leakage. In this way, privacy protection capability of the encoder can be further strengthened.”) Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES WEBB whose telephone number is (313)446-6615. The examiner can normally be reached on M-F 10-3. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jerry O’Connor can be reached on (571) 272-6787. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAMES WEBB/Examiner, Art Unit 3624 Application/Control Number: 18/564,160 Page 2 Art Unit: 3624 Application/Control Number: 18/564,160 Page 3 Art Unit: 3624 Application/Control Number: 18/564,160 Page 4 Art Unit: 3624 Application/Control Number: 18/564,160 Page 5 Art Unit: 3624 Application/Control Number: 18/564,160 Page 6 Art Unit: 3624 Application/Control Number: 18/564,160 Page 7 Art Unit: 3624 Application/Control Number: 18/564,160 Page 8 Art Unit: 3624 Application/Control Number: 18/564,160 Page 9 Art Unit: 3624 Application/Control Number: 18/564,160 Page 10 Art Unit: 3624 Application/Control Number: 18/564,160 Page 11 Art Unit: 3624 Application/Control Number: 18/564,160 Page 12 Art Unit: 3624 Application/Control Number: 18/564,160 Page 13 Art Unit: 3624 Application/Control Number: 18/564,160 Page 14 Art Unit: 3624 Application/Control Number: 18/564,160 Page 16 Art Unit: 3624 Application/Control Number: 18/564,160 Page 17 Art Unit: 3624 Application/Control Number: 18/564,160 Page 18 Art Unit: 3624 Application/Control Number: 18/564,160 Page 19 Art Unit: 3624 Application/Control Number: 18/564,160 Page 20 Art Unit: 3624 Application/Control Number: 18/564,160 Page 21 Art Unit: 3624 Application/Control Number: 18/564,160 Page 22 Art Unit: 3624 Application/Control Number: 18/564,160 Page 23 Art Unit: 3624 Application/Control Number: 18/564,160 Page 24 Art Unit: 3624 Application/Control Number: 18/564,160 Page 25 Art Unit: 3624 Application/Control Number: 18/564,160 Page 26 Art Unit: 3624 Application/Control Number: 18/564,160 Page 27 Art Unit: 3624 Application/Control Number: 18/564,160 Page 28 Art Unit: 3624 Application/Control Number: 18/564,160 Page 29 Art Unit: 3624 Application/Control Number: 18/564,160 Page 30 Art Unit: 3624 Application/Control Number: 18/564,160 Page 31 Art Unit: 3624 Application/Control Number: 18/564,160 Page 32 Art Unit: 3624 Application/Control Number: 18/564,160 Page 33 Art Unit: 3624 Application/Control Number: 18/564,160 Page 34 Art Unit: 3624 Application/Control Number: 18/564,160 Page 35 Art Unit: 3624 Application/Control Number: 18/564,160 Page 36 Art Unit: 3624