Prosecution Insights
Last updated: October 01, 2026
Application No. 18/750,688

METHOD AND APPARATUS FOR TRAINING INTELLIGENT MODEL

Non-Final OA §103
Filed
Jun 21, 2024
Priority
Dec 22, 2021 — CN 202111582987.9 +1 more
Examiner
BALAKRISHNAN, VIJAY MURALI
Art Unit
Tech Center
Assignee
Huawei Technologies Co., Ltd.
OA Round
1 (Non-Final)
41%
Grant Probability
Moderate
1-2
OA Rounds
1y 8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 41% of resolved cases
41%
Career Allowance Rate
11 granted / 27 resolved
-19.3% vs TC avg
Strong +73% interview lift
Without
With
+73.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
14 currently pending
Career history
44
Total Applications
across all art units

Statute-Specific Performance

§101
27.5%
-12.5% vs TC avg
§103
36.6%
-3.4% vs TC avg
§102
12.6%
-27.4% vs TC avg
§112
23.3%
-16.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 27 resolved cases

Office Action

§103
DETAILED ACTION This nonfinal action is in response to application 18/750,688 filed on 06/21/2024, which is a continuation of international application PCT/CN2022/140797 filed 12/21/2022 with priority to foreign application CN202111582987.9 filed on 12/22/2021. Claims 21-40 are pending in the application. Claims 21, 31, and 40 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) filed 08/26/2024 and 03/26/2025 have been fully considered by the examiner. Specification The specification is objected to because the title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: “Method and Apparatus for Federated Learning Using Gradient Quantization and a Shared Feature Constraint”. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 21-24, 31-33, and 39 are rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (“Federated Multi-Task Learning”, published 2017), hereinafter Smith, in view of Xu et al. (“FedV: Privacy-Preserving Federated Learning over Vertically Partitioned Data”, available conference 11/15/2021) hereinafter Xu, and Sun et al. (“Lazily Aggregated Quantized Gradient Innovation for Communication-Efficient Federated Learning”, available online 10/23/2020), hereinafter Sun. Regarding claim 21, Smith teaches A method, performed by a first participant node, (“In this work, we propose a modeling approach that differs significantly from prior work on federated learning, where the aim thus far has been to train a single global model across the network [25, 36, 26]. Instead, we address statistical challenges in the federated setting by learning separate models for each node, {w1; : : : ;wm}. This can be naturally captured through a multi-task learning (MTL) framework, where the goal is to consider fitting separate but related models simultaneously [14, 2, 57, 28]” [Smith page 1 Introduction]) the method comprising: receiving first information from a central node, wherein the first information indicates an inter-feature constraint variable, and the inter-feature constraint variable represents a constraint relationship between different features (“The matrix PNG media_image1.png 30 162 media_image1.png Greyscale models relationships amongst tasks… MTL problems differ based on their assumptions on R, which takes PNG media_image2.png 42 51 media_image2.png Greyscale as input and promotes some suitable structure amongst the tasks.” [Smith page 3 General Multi-Task Learning Setup]; see Algorithm 1 – PNG media_image3.png 473 1008 media_image3.png Greyscale [Smith page 4]; Updating of PNG media_image2.png 42 51 media_image2.png Greyscale (i.e., inter-feature constraint promoting structure amongst node tasks) occurs centrally (line 11 of Algorithm 1), and PNG media_image2.png 42 51 media_image2.png Greyscale is continually received from central node by PNG media_image4.png 25 8 media_image4.png Greyscale nodes (i.e., participant nodes) for each iteration (lines 5 and 6 of Algorithm 1) via updating of node-local variables with PNG media_image5.png 32 52 media_image5.png Greyscale (lines 7 and 8 of Algorithm 1), which depends on PNG media_image2.png 42 51 media_image2.png Greyscale via term PNG media_image6.png 42 57 media_image6.png Greyscale (see equation 4 PNG media_image7.png 87 1056 media_image7.png Greyscale where PNG media_image8.png 46 250 media_image8.png Greyscale and “R* depends on PNG media_image2.png 42 51 media_image2.png Greyscale ” [Smith page 4 Federated Update of W])), wherein the central node and a plurality of participant node groups jointly perform training of an intelligent model (“Given data PNG media_image9.png 37 151 media_image9.png Greyscale from m nodes, multi-task learning fits separate weight vectors PNG media_image10.png 38 107 media_image10.png Greyscale to the data for each task (node) through arbitrary convex loss functions PNG media_image11.png 27 31 media_image11.png Greyscale (e.g., the hinge loss for SVM models)… The matrix PNG media_image1.png 30 162 media_image1.png Greyscale models relationships amongst tasks, and is either known a priori or estimated while simultaneously learning task models” [Smith page 3 General Multi-Task Learning Setup]; “Remark 4. MOCHA can be modified to solve problems when there are tasks that are shared among nodes. In this case, each node still solves a data local sub-problem based on its own data for the task, but the central node needs to do an additional aggregation step to add the results for all the nodes that share the data of each task. This reduces the size of matrix and simplifies its update” [Smith pages 13-14 Appendix – Reducing the Size of PNG media_image2.png 42 51 media_image2.png Greyscale by Sharing Tasks]; To reduce burden of central matrix updates, nodes may share the same task (i.e., form a participant node group)), wherein participant nodes are in each one of the participant node group, and the first participant node is in a first node group of the plurality of participant node groups; ([Smith pages 13-14 Appendix – Reducing the Size of PNG media_image2.png 42 51 media_image2.png Greyscale by Sharing Tasks] as detailed above) obtaining first update information based on the inter-feature constraint variable and first sample data using an inference model, wherein the first update information corresponds to the inter-feature constraint variable (see lines 7-8 of Algorithm 1 [Smith page 4] as detailed above – node-local variable updates depend on PNG media_image5.png 32 52 media_image5.png Greyscale , calculated via PNG media_image7.png 87 1056 media_image7.png Greyscale , i.e., uses locally stored sample data Xt, and depends on PNG media_image2.png 42 51 media_image2.png Greyscale via term PNG media_image6.png 42 57 media_image6.png Greyscale for node-local optimization objective) However, Smith does not expressly teach the intelligent model consist[ing] of a plurality of feature models corresponding to a plurality of features of an inference target, wherein participant nodes train a respective one of the feature models, and the first participant node trains a first feature model, and obtaining gradient information based on a model parameter of the first feature model. In the same field of endeavor, Xu teaches implementation of a federated learning framework (“Privacy-preserving vertical FL is challenging because complete sets of labels and features are not owned by one entity. Existing approaches for vertical FL require multiple peer-to-peer communications among parties, leading to lengthy training times, and are restricted to (approximated) linear models and just two parties. To close this gap, we propose FedV, a framework for secure gradient computation in vertical settings for several widely used ML models such as linear models, logistic regression, and support vector machines” [Xu Abstract]) wherein the intelligent model consists of a plurality of feature models corresponding to a plurality of features of an inference target, (“Let P = {𝑝𝑖 }𝑖 ∈[𝑛] be the set of 𝑛 parties in VFL. Let D[𝑋,𝑌 ] be the training dataset across the set of parties P, where 𝑋 ∈ R𝑑 represents the feature set and 𝑌 ∈ R denotes the labels. We assume that except for the identifier features, there are no overlapping training features between any two parties’ local datasets, and these datasets can form the “global” dataset D.” [Xu page 182 Vertical Federated Learning]; “Under the VFL setting, the gradient computation at each training epoch relies on (i) the parties’ collaboration to exchange their “partial model” with each other, or (ii) exposing their data to the aggregator to compute the final gradient update” [Xu page 183 Gradient Descent in Vertical FL]; Each party in the VFL framework has a respective partial model corresponding to its local features) wherein participant nodes train a respective one of the feature models, and the first participant node trains a first feature model (“For this purpose, the active party and all other passive parties perform slightly different pre-processing steps before invoking FedV-SecGrad. The active party, 𝑝1, appends a vector with labels 𝑦 to obtain 𝑥 (𝑖) 𝑝1 𝑤𝑝1 − 𝑦(𝑖) as its ‘partial model update’. For the passive party 𝑝2, its ‘partial model update’ is defined by 𝑥 (𝑖) 𝑝2 𝑤𝑝2” [Xu page 185 FedV-SecGrad Process]) and obtaining gradient information based on a model parameter of the first feature model (“FedV-SecGrad is a generic approach to securely compute gradients of an ML objective with a prediction function that can be written as 𝑓 (𝑥;𝑤) := 𝑔(𝑤⊺𝑥), where 𝑔 : R → R is a differentiable function, 𝑥 and 𝑤 denote the feature vector and the model weights vector, respectively.” [Xu page 185 VERTICAL TRAINING: FEDV-SECGRAD]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the intelligent model consist[ing] of a plurality of feature models corresponding to a plurality of features of an inference target, wherein participant nodes train a respective one of the feature models, and the first participant node trains a first feature model, and obtaining gradient information based on a model parameter of the first feature model as taught by Xu into Smith because they are both directed towards federated learning frameworks. Given that Smith already suggests implementation of aggregation across node-local updates to reduce burden of central matrix updates [Smith pages 13-14 Appendix – Reducing the Size of PNG media_image2.png 42 51 media_image2.png Greyscale by Sharing Tasks], incorporating the teachings of Xu would further enable an efficient aggregation step that minimizes need for peer-to-peer communication and thereby reduces training time [Xu Abstract]. However, the combination does not expressly teach sending the first gradient information to the central node. In the same field of endeavor, Sun teaches a federated learning framework that send[s] the first gradient information to the central node (“This paper focuses on communication-efficient federated learning problem, and develops a novel distributed quantized gradient approach, which is characterized by adaptive communications of the quantized gradients. Specifically, the federated learning builds upon the server-worker infrastructure, where the workers calculate local gradients and upload them to the server; then the server obtain the global gradient by aggregating all the local gradients and utilizes it to update the model parameter” [Sun Abstract]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated sending the first gradient information to the central node as taught by Sun into the combination because Smith and Sun are both directed towards federated learning frameworks. While Smith does not calculate global model updates, it does teach aggregation of node-local updates via the central node [Smith pages 13-14 Appendix – Reducing the Size of PNG media_image2.png 42 51 media_image2.png Greyscale by Sharing Tasks], wherein said updates further amount to loss minimization objectives optimizable via gradient descent. As such, incorporating the teachings of Sun would allow greater efficiency during aggregation of local information by reducing communication burden and skipping unnecessary or less informative gradient communications between nodes [Sun Abstract]. Regarding claim 22, the combination of Smith, Xu, and Sun teaches the limitations of parent claim 21, and Xu further teaches receiving a first identifier set from the central node, (see Procedure 2: FedV-SecGrad including lines 1 and 6 [Xu page 186]; Aggregator (i.e., central node) generates batch indices {1,…,m}, and queries each party (i.e., participant) using PNG media_image12.png 20 192 media_image12.png Greyscale ) wherein the first identifier set comprises an identifier of sample data of the inter-feature constraint variable, and the inter-feature constraint variable is selected by the central node; (see Procedure 2: FedV-SecGrad [Xu page 186] as detailed above; “Each party uses the one-time-password chain associated with the training epoch to randomly select the samples that are going to be included in a batch for the given batch index, as shown in line 15” [Xu page 186 Detailed Execution of FedV-SecGrad Process]) and wherein obtaining first gradient information based on the inter-feature constraint variable, the model parameter of the first feature model, and the first sample data using the gradient inference model comprises: determining that a sample data set of the first participant node comprises the first sample data based on the first sample data corresponding to a first identifier, wherein the first identifier belongs to the first identifier set (see Procedure 2: FedV-SecGrad [Xu page 186] as detailed above; Party has local dataset PNG media_image13.png 28 30 media_image13.png Greyscale , and after receiving batch index PNG media_image14.png 21 30 media_image14.png Greyscale it uses PNG media_image15.png 22 257 media_image15.png Greyscale to get batch samples (i.e., sample data set)); and obtaining the first gradient information based on the inter-feature constraint variable, the model parameter of the first feature model, and the first sample data using the gradient inference model ([Xu page 185 VERTICAL TRAINING: FEDV-SECGRAD] as detailed above; see Procedure 2: FedV-SecGrad [Xu page 186] as detailed above; Execution of aggregation procedure results in acquisition of batch gradient (lines 12 and 13 of Procedure 2)). Regarding claim 23, the combination of Smith, Xu, and Sun teaches the limitations of parent claim 21, and Sun further teaches wherein sending the first gradient information to the central node comprises: sending first target gradient information that is quantized to the central node, wherein the first target gradient information comprises the first gradient information (“ PNG media_image16.png 43 77 media_image16.png Greyscale is the quantized gradient that coarsely approximates the local gradient PNG media_image17.png 28 82 media_image17.png Greyscale … Similar to GD however, only when all the local quantized gradients PNG media_image18.png 27 93 media_image18.png Greyscale are collected, the server can update θ” [Sun page 2032 Context and Contributions in a Nutshell]) Regarding claim 24, the combination of Smith, Xu, and Sun teaches the limitations of parent claim 23, and Sun further teaches obtaining second residual gradient information based on the first target gradient information and the first target gradient information that is quantized, wherein the second residual gradient information is a residual amount that is in the first target gradient information and that is not sent to the central node (“With PNG media_image19.png 31 167 media_image19.png Greyscale PNG media_image20.png 31 70 media_image20.png Greyscale denoting the local quantization error, it is clear that the quantization error is not larger that half of the length of the interval that each value covers… The aggregated quantized gradient is PNG media_image21.png 37 221 media_image21.png Greyscale … that is PNG media_image22.png 35 196 media_image22.png Greyscale ” [Sun page 2033 Gradient Innovation-Based Quantization]). Regarding claim 31, it is a method claim that substantially corresponds to the method of claim 21, which is already taught by the combination of Smith, Xu, and Sun as detailed above. Xu further teaches training the first feature model based on the inter-feature constraint variable and model training data, ([Xu page 185 FedV-SecGrad Process] as detailed in claim 21 above) and Sun further teaches obtain[ing] second gradient information; and sending the second gradient information to the central node ([Sun Abstract] as detailed in claim 21 above; an additional worker may send another (i.e. second) local gradient to the central node). Consequently, claim 31 is rejected for the same reasons as claim 21. Regarding claims 32-33, they are method claims that substantially correspond to the method of claims 23-24. Consequently, they are rejected for the same reasons as claims 23-24. Regarding claim 39, the combination of Smith, Xu, and Sun teaches the limitations of parent claim 31, and Xu further teaches receiving third information from the central node, wherein the third information indicates an updated parameter of the first feature model; (“The aggregator queries the parties with the current model weights,𝑤𝑝𝑖 . To reduce data transfer and protect against inference attacks1, the aggregator only sends each party the weights that pertain to its partial feature set. We denote these partial model weights as 𝑤𝑝𝑖 in line 2.” [Xu page 186]) and updating a parameter of the first feature model based on the third information (“The active party, 𝑝1, appends a vector with labels 𝑦 to obtain 𝑥 (𝑖) 𝑝1 𝑤𝑝1 − 𝑦(𝑖) as its ‘partial model update’. For the passive party 𝑝2, its ‘partial model update’ is defined by 𝑥 (𝑖) 𝑝2 𝑤𝑝2 . Each party 𝑝𝑖 encrypts its ‘partial model update’ using the MIFE encryption algorithm with its encryption key skMIFE 𝑝𝑖 , and sends it to the aggregator” [Xu page 185]). Claims 25-30 and 34-38 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Smith, Xu, and Sun, as applied to claims 23 and 32 above, further in view of Wang et al., (“Quantized Federated Learning Under Transmission Delay and Outage Constraints”, available online 11/11/2021), hereinafter Wang. Regarding claim 25, the combination of Smith, Xu, and Sun teaches the limitations of parent claim 23, and Sun further teaches determining a first threshold based on first quantization noise information, wherein the first quantization noise information represents a loss introduced by quantization encoding and decoding on the first target gradient information (“Now it only remains to design the selection criterion to decide which worker to upload the quantized gradient or its innovation. We propose the following communication criterion: worker m ∈ M skips the upload at iteration k, if it satisfies PNG media_image23.png 202 585 media_image23.png Greyscale where PNG media_image24.png 36 173 media_image24.png Greyscale are predetermined constants, and PNG media_image25.png 32 290 media_image25.png Greyscale is the quantization error of the local gradient when last time worker m uploads gradient innovation to the server” [Sun pages 2033-2034 Gradient Innovation-Based Aggregation]; Squared norm of the difference between current and previously uploaded quantized gradients is evaluated against right hand side of equation (9a) (i.e., threshold) which includes quantization errors (i.e., noise information)); and wherein sending the first target gradient information that is quantized to the central node comprises: determining that a metric value of the first target gradient information is greater than the first threshold; and sending the first target gradient information that is quantized to the central node in response to determining that the metric value of the first target gradient information is greater than the first threshold ([Sun pages 2033-2034 Gradient Innovation-Based Aggregation] as detailed above; If the squared norm of the difference between current and previously uploaded quantized gradients (i.e., metric value) is not less than (i.e., is greater than) the threshold, then the worker uploads the quantized gradient) However, the combination does not expressly teach further determining a first threshold based on channel resource information. In the same field of endeavor, Wang teaches a federated learning framework (“In this paper, we consider such non-ideal wireless channels, and carry out the first analysis showing that the FL convergence can be severely jeopardized by TO and QE, but intriguingly can be alleviated if the clients have uniform outage probabilities. These insightful results motivate us to propose a robust FL scheme, named FedTOE, which performs joint allocation of wireless resources and quantization bits across the clients to minimize the QE while making the clients have the same TO probability” [Wang Abstract]) that determines a first threshold based on channel resource information (“As suggested by Corollary 1 that it is crucial to maintain a uniform outage probability across the clients, we enforce the constraint qi = qmax for all i = 1, · · · ,N, where qmax ∈ (0, 0.5] is a preset target outage probability value. Then, by (19), it remains to reduce the effect of quantization errors. Therefore, aiming at improving the learning performance, we choose to minimize the accumulative average PNG media_image26.png 36 230 media_image26.png Greyscale in the RHS of (19) under the constraints of uniform outage probability and transmission delay.3” [Wang page 329 Proposed FedTOE]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated determining a first threshold based on channel resource information as taught by Wang into the combination because both Smith and Wang are directed towards federated learning frameworks. Incorporating the teachings of Wang would enable the combined FL framework to consider and account for limitations of non-ideal wireless channels on quantization error [Wang Abstract]. Regarding claim 26, the combination of Smith, Xu, Sun, and Wang teaches the limitations of parent claim 25, and Sun further teaches when the metric value of the first target gradient information is less than or equal to the first threshold, determining not to send the first target gradient information that is quantized to the central node ([Sun pages 2033-2034 Gradient Innovation-Based Aggregation] as detailed above; If the squared norm of the difference between current and previously uploaded quantized gradients (i.e., metric value) is less the threshold, then the worker skips the quantized gradient) Regarding claim 27, the combination of Smith, Xu, Sun, and Wang teaches the limitations of parent claim 26, and Sun further teaches in response to determining that the metric value of the first target gradient information is less than the first threshold, determining third residual gradient information, wherein the third residual gradient information is the first target gradient information (([Sun pages 2033-2034 Gradient Innovation-Based Aggregation] as detailed above; When the quantized gradient is skipped, the information is by nature retained as skipped or delayed, i.e., residual information) Regarding claim 28, the combination of Smith, Xu, Sun, and Wang teaches the limitations of parent claim 25, and Sun further teaches obtaining the first quantization noise information based on the first target gradient information With PNG media_image19.png 31 167 media_image19.png Greyscale PNG media_image20.png 31 70 media_image20.png Greyscale denoting the local quantization error, it is clear that the quantization error is not larger that half of the length of the interval that each value covers… The aggregated quantized gradient is PNG media_image21.png 37 221 media_image21.png Greyscale , and the aggregated quantization error is PNG media_image27.png 27 147 media_image27.png Greyscale PNG media_image28.png 42 175 media_image28.png Greyscale ” [Sun page 2033 Gradient Innovation-Based Quantization]. Wang further teaches obtaining quantization noise information based on the channel resource information and communication cost information, ([Wang page 329 Proposed FedTOE] as detailed in claim 25 above) and wherein the communication cost information indicates a communication cost weight of a communication resource, and the communication resource comprises transmission power or transmission bandwidth (“By (9), this yields the following resource allocation problem…where Wtotal is the total bandwidth of the uplink channel, Pmax is the maximum transmit power of each client, τi is the average uplink transmission delay per communication round of client i, τmax is the constraint on uplink transmission delay, and Z+ is the positive integer set” [Wang page 329 Proposed FedTOE]). Regarding claim 29, the combination of Smith, Xu, Sun, and Wang teaches the limitations of parent claim 25, and Wang further teaches determining a transmission bandwidth or a transmission power based on the first quantization noise information, communication cost information, the channel resource information, and the first target gradient information, wherein the communication cost information indicates a communication cost weight of a communication resource, and the communication resource comprises the transmission power or the transmission bandwidth; ([Wang page 329 Proposed FedTOE] as detailed above; see also Algorithm 2 FedTOE: Algorithm to Solve – PNG media_image29.png 60 655 media_image29.png Greyscale [Wang page 330]) and determining the first threshold based on the first quantization noise information and the communication resource ([Sun pages 2033-2034 Gradient Innovation-Based Aggregation] and ([Wang page 329 Proposed FedTOE] as detailed in claim 25 above). Regarding claim 30, the combination of Smith, Xu, Sun, and Wang teaches the limitations of parent claim 29, and Wang further teaches receiving second information from the central node, wherein the second information indicates the communication cost information ([Wang page 329 Proposed FedTOE] and Algorithm 2 FedTOE: Algorithm to Solve [Wang page 330] as detailed in claim 29 above; The server (i.e., central node) coordinates client-specific communication cost information, wherein said information is implicitly received from a user with access to the server). Regarding claims 34-38, they are method claims that substantially correspond to the methods of claims 25-29, which are already taught by the combination of Smith, Xu, Sun, and Wang as detailed above. Consequently, they are rejected for the same reasons. Claim 40 is rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (“Federated Multi-Task Learning”, published 2017), hereinafter Smith, in view of Xu et al. (“FedV: Privacy-Preserving Federated Learning over Vertically Partitioned Data”, available conference 11/15/2021) hereinafter Xu. Regarding claim 40, Smith teaches A method, performed by a central node, (“In this work, we propose a modeling approach that differs significantly from prior work on federated learning, where the aim thus far has been to train a single global model across the network [25, 36, 26]. Instead, we address statistical challenges in the federated setting by learning separate models for each node, {w1; : : : ;wm}. This can be naturally captured through a multi-task learning (MTL) framework, where the goal is to consider fitting separate but related models simultaneously [14, 2, 57, 28]” [Smith page 1 Introduction]; “Note that solving for PNG media_image30.png 22 27 media_image30.png Greyscale is not dependent on the data and therefore can be computed centrally” [Smith page 3 MOCHA: A Framework for Federated Multi-Task Learning]) the method comprising: determining an inter-feature constraint variable, wherein the inter-feature constraint variable represents a constraint relationship between different features, (“The matrix PNG media_image1.png 30 162 media_image1.png Greyscale models relationships amongst tasks… MTL problems differ based on their assumptions on R, which takes PNG media_image2.png 42 51 media_image2.png Greyscale as input and promotes some suitable structure amongst the tasks.” [Smith page 3 General Multi-Task Learning Setup]; see Algorithm 1 – PNG media_image3.png 473 1008 media_image3.png Greyscale [Smith page 4]; Updating of PNG media_image2.png 42 51 media_image2.png Greyscale (i.e., inter-feature constraint promoting structure amongst node tasks) occurs centrally (line 11 of Algorithm 1), and PNG media_image2.png 42 51 media_image2.png Greyscale is continually received from central node by PNG media_image4.png 25 8 media_image4.png Greyscale nodes (i.e., participant nodes) for each iteration (lines 5 and 6 of Algorithm 1) via updating of node-local variables with PNG media_image5.png 32 52 media_image5.png Greyscale (lines 7 and 8 of Algorithm 1), which depends on PNG media_image2.png 42 51 media_image2.png Greyscale via term PNG media_image6.png 42 57 media_image6.png Greyscale (see equation 4 PNG media_image7.png 87 1056 media_image7.png Greyscale where PNG media_image8.png 46 250 media_image8.png Greyscale and “R* depends on PNG media_image2.png 42 51 media_image2.png Greyscale ” [Smith page 4 Federated Update of W])) wherein the central node and a plurality of participant node groups jointly perform training of an intelligent model, model (“Given data PNG media_image9.png 37 151 media_image9.png Greyscale from m nodes, multi-task learning fits separate weight vectors PNG media_image10.png 38 107 media_image10.png Greyscale to the data for each task (node) through arbitrary convex loss functions PNG media_image11.png 27 31 media_image11.png Greyscale (e.g., the hinge loss for SVM models)… The matrix PNG media_image1.png 30 162 media_image1.png Greyscale models relationships amongst tasks, and is either known a priori or estimated while simultaneously learning task models” [Smith page 3 General Multi-Task Learning Setup]; “Remark 4. MOCHA can be modified to solve problems when there are tasks that are shared among nodes. In this case, each node still solves a data local sub-problem based on its own data for the task, but the central node needs to do an additional aggregation step to add the results for all the nodes that share the data of each task. This reduces the size of matrix and simplifies its update” [Smith pages 13-14 Appendix – Reducing the Size of PNG media_image2.png 42 51 media_image2.png Greyscale by Sharing Tasks]; To reduce burden of central matrix updates, nodes may share the same task (i.e., form a participant node group)) wherein participant nodes are in each one of the participant node groups; ([Smith pages 13-14 Appendix – Reducing the Size of PNG media_image2.png 42 51 media_image2.png Greyscale by Sharing Tasks] as detailed above) and sending first information to a participant node in the plurality of participant node groups, wherein the first information comprises the inter-feature constraint variable. (see lines 7-8 of Algorithm 1 [Smith page 4] as detailed above – node-local variable updates depend on PNG media_image5.png 32 52 media_image5.png Greyscale , calculated via PNG media_image7.png 87 1056 media_image7.png Greyscale , i.e., uses locally stored sample data Xt, and depends on PNG media_image2.png 42 51 media_image2.png Greyscale via term PNG media_image6.png 42 57 media_image6.png Greyscale for node-local optimization objective). However, Smith does not expressly teach the intelligent model consist[ing] of a plurality of feature models corresponding to a plurality of features of an inference target, wherein participant nodes train a respective one of the feature models. In the same field of endeavor, Xu teaches implementation of a federated learning framework (“Privacy-preserving vertical FL is challenging because complete sets of labels and features are not owned by one entity. Existing approaches for vertical FL require multiple peer-to-peer communications among parties, leading to lengthy training times, and are restricted to (approximated) linear models and just two parties. To close this gap, we propose FedV, a framework for secure gradient computation in vertical settings for several widely used ML models such as linear models, logistic regression, and support vector machines” [Xu Abstract]) wherein the intelligent model consists of a plurality of feature models corresponding to a plurality of features of an inference target, (“Let P = {𝑝𝑖 }𝑖 ∈[𝑛] be the set of 𝑛 parties in VFL. Let D[𝑋,𝑌 ] be the training dataset across the set of parties P, where 𝑋 ∈ R𝑑 represents the feature set and 𝑌 ∈ R denotes the labels. We assume that except for the identifier features, there are no overlapping training features between any two parties’ local datasets, and these datasets can form the “global” dataset D.” [Xu page 182 Vertical Federated Learning]; “Under the VFL setting, the gradient computation at each training epoch relies on (i) the parties’ collaboration to exchange their “partial model” with each other, or (ii) exposing their data to the aggregator to compute the final gradient update” [Xu page 183 Gradient Descent in Vertical FL]; Each party in the VFL framework has a respective partial model corresponding to its local features) wherein participant nodes train a respective one of the feature models, and the first participant node trains a first feature model (“For this purpose, the active party and all other passive parties perform slightly different pre-processing steps before invoking FedV-SecGrad. The active party, 𝑝1, appends a vector with labels 𝑦 to obtain 𝑥 (𝑖) 𝑝1 𝑤𝑝1 − 𝑦(𝑖) as its ‘partial model update’. For the passive party 𝑝2, its ‘partial model update’ is defined by 𝑥 (𝑖) 𝑝2 𝑤𝑝2” [Xu page 185 FedV-SecGrad Process]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the intelligent model consist[ing] of a plurality of feature models corresponding to a plurality of features of an inference target, wherein participant nodes train a respective one of the feature models as taught by Xu into Smith because they are both directed towards federated learning frameworks. Given that Smith already suggests implementation of aggregation across node-local updates to reduce burden of central matrix updates [Smith pages 13-14 Appendix – Reducing the Size of PNG media_image2.png 42 51 media_image2.png Greyscale by Sharing Tasks], incorporating the teachings of Xu would further enable an efficient aggregation step that minimizes need for peer-to-peer communication and thereby reduces training time [Xu Abstract]. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to VIJAY M BALAKRISHNAN whose telephone number is (571) 272-0455. The examiner can normally be reached 10am-5pm EST Mon-Thurs. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER WELCH can be reached on (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /V.M.B./ Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Jun 21, 2024
Application Filed
Oct 18, 2024
Response after Non-Final Action
Sep 23, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743623
INFORMATION PROCESSING DEVICE AND MACHINE LEARNING METHOD THAT OPTIMIZE A DECODING PROCESS USING BACK-PROPAGATION
4y 0m to grant Granted Sep 22, 2026
Patent 12731026
METHOD AND SYSTEM FOR PROGRAM SAMPLING USING NEURAL NETWORK
3y 10m to grant Granted Sep 08, 2026
Patent 12711407
REASONING METHOD BASED ON STRUCTURAL ATTENTION MECHANISM FOR KNOWLEDGE-BASED QUESTION ANSWERING AND COMPUTING APPARATUS FOR PERFORMING THE SAME
3y 8m to grant Granted Aug 18, 2026
Patent 12645933
Method and System for Training a Neural Network for Generating Universal Adversarial Perturbations
4y 7m to grant Granted Jun 02, 2026
Patent 12619871
INTERPRETABLE NEURAL NETWORK ARCHITECTURE USING CONTINUED FRACTIONS
3y 11m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
41%
Grant Probability
99%
With Interview (+73.3%)
3y 11m (~1y 8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 27 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month