Prosecution Insights
Last updated: August 17, 2026
Application No. 17/948,392

Data Processing Method and Apparatus

Final Rejection §103
Filed
Sep 20, 2022
Priority
Mar 20, 2020 — CN 202010202053.7 +1 more
Examiner
MAHARAJ, DEVIKA S
Art Unit
2100
Tech Center
2100 — Computer Architecture & Software
Assignee
Huawei Technologies Co., Ltd.
OA Round
2 (Final)
56%
Grant Probability
Moderate
3-4
OA Rounds
8m
Est. Remaining
65%
With Interview

Examiner Intelligence

Grants 56% of resolved cases
56%
Career Allowance Rate
48 granted / 86 resolved
+0.8% vs TC avg
Moderate +9% lift
Without
With
+9.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 7m
Avg Prosecution
22 currently pending
Career history
111
Total Applications
across all art units

Statute-Specific Performance

§101
30.0%
-10.0% vs TC avg
§103
46.4%
+6.4% vs TC avg
§102
10.5%
-29.5% vs TC avg
§112
10.5%
-29.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 86 resolved cases

Office Action

§103
DETAILED ACTION 1. This communication is in response to the amendments filed on October 31, 2025 for Application No. 17/948,392 in which Claims 1-20 are presented for examination. Notice of Pre-AIA or AIA Status 2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments 3. The amendments filed on October 31, 2025 have been considered. Claims 1, 3, 8, 10, 15-16, and 20 have been amended. Thus, Claims 1-20 are pending and presented for examination. 4. Applicant’s arguments and corresponding amendments filed October 31, 2025 with respect to the specification objection of the title have been fully considered and are persuasive. Thus, the specification objection has been withdrawn. 5. Applicant’s arguments and corresponding amendments filed October 31, 2025 with respect to the 35 U.S.C. 101 rejection have been fully considered and are persuasive. Thus, the 35 U.S.C. 101 rejection has been withdrawn. 6. Applicant’s arguments filed October 31, 2025 with respect to the 35 U.S.C. 103 rejection have been fully considered but they are not persuasive. Applicant’s Arguments on Pgs. 11-12 of Arguments/Remarks state: “The combination of Zhang and Gierach fails to render obvious claims 1-20 because the combination of Zhang and Gierach fails to obtain a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold. Claim 1 reads: […] As shown above, claim 1 requires obtaining a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold. Independent claims 8 and 15 contain similar limitations. Support for the amendments can be found in previously presented claim 3. To reject previously presented claim 3, the Examiner acknowledges that Zhang fails to disclose these limitations, but asserts that paragraph 32 of Gierach discloses similar limitations. See Office Action, p. 24-25. However, Gierach does not obtain a third model by deleting a feature interaction item corresponding to an optimized architecture parameter whose value is less than a threshold: […] As shown above, Gierach discloses identifying and removing second-order cross terms based on a statistical ratio (B/A) compared to a threshold (TH4). Gierach's approach is static and heuristic-driven, with no model-driven training or learnable parameters involved. In contrast, the present claims recite a method that incorporates trainable architecture parameters into an FM-based model, optimize those parameters based on data, and then refine the model by deleting interaction terms associated with low-valued optimized parameters, not based on heuristics. Thus, Gierach fails to obtain a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold, as claimed. For at least these reasons, the Applicant submits that the rejection of claims 1-20 under 35 U.S.C. § 103 as being rendered obvious by the combination of Zhang and Gierach is improper. The Applicant respectfully requests to withdraw the rejections.” Examiner respectfully disagrees. As disclosed by the rejection of previously presented claim 3, which recites substantially the same limitations as the newly added limitations of the Independent claims, Examiner cited Gierach Par. [0019] and Par. [0032] which recite “The second K-th order attribution model is effectively the K-th order attribution model with the set of insignificant second order cross terms removed. The modified second K-th order attribution model is effectively the K-th order attribution model with the set of insignificant second order cross terms removed and with the first function of the second order cross terms replaced by the third function.” & “In addition, the second step may also include classifying the second order cross term with the two independent variables X_i and X_j as unimportant if a monotonic non-decreasing function of the fraction B/A is less than a threshold ‘TH4’. A second order cross term being classified as unimportant by either the first step and/or the second step may be identified as insignificant.” – therefore, in order to obtain a k-th order attribution model (third model, which is based on a preceding first/second k-th order attribution model as indicated above), a second feature interaction item corresponding to one of the optimized architecture parameters (second order cross terms with two independent variables) may be deleted/removed. Contrary to Applicant’s arguments, this approach is not static, and also incorporates trainable parameters – See for example Gierach Claim 1 which states “[…] second order cross terms each comprising the first function of two of the M independent variables weighted by one of the (M)(M-1)/2 second order model parameters minus the set of insignificant second order cross terms of the k-th order attribution model […]”. Thus, the variables/terms as presented by Gierach are associated with a tunable weight value, indicating that these terms are learnable/trainable parameters and not associated with only a static approach. Nonetheless, it must be noted that nothing in the instant claim language states that the obtained “third model” must consist of trainable parameters – instead, the claim language states that a third model is obtained by deleting second feature interaction items and loosely based on the first and second models which comprise trainable values. As the third model is obtained merely based on the first and second models without significantly more, this does not immediately indicate that the third model comprises such trainable values, only that it comprises particular feature interaction items. Hence, Zhang in view of Gierach still teaches the instant limitations. Thus, the 35 U.S.C. 103 rejection is maintained. Claim Rejections - 35 USC § 103 7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 8. Claims 1-4, 6-7, 8-9, 11, 13-14, 15, 17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (hereinafter Zhang) (“Field-Aware Neural Factorization Machine for Click-Through Rate Prediction”), in view of Gierach et al. (hereinafter Gierach) (US PG-PUB 20180075482). Regarding Claim 1, Zhang teaches: A method (Zhang, Pg. 1, Abstract, “This paper combines traditional feature combination methods and deep neural networks to automate feature combinations to improve the accuracy of click through rate prediction. We propose a mechanism named ’Field-aware Neural Factorization Machine’ (FNFM).”, thus, methods are disclosed) comprising: adding first architecture parameters to first feature interaction items in a first model to obtain second model, wherein the first model is a factorization machine (FM)-based model, and wherein each of the first architecture parameters is a trainable value representing an importance of a corresponding feature interaction item; (Zhang teaches first architecture parameters Vi (Page 4 Section 2) which are added to interaction items (vector x on Page 3). Both the first model and the second model are understood to be factorization machines: “The FNFM model uses the idea of factorization machine to learn the expression of second-order interactive features (reads on second architecture parameters) in the form of hidden vector products.” [Page 4], in which the second model is the “learned” version of the first model with the expression of second-order interaction features. The embedding vector is understood to be related to the importance of the feature interaction item, as taught in “Space Optimized Mode” on page 2, where each vi is tied to each feature item and is an embedding value for it. Further, the architecture parameters are trainable, as it is stated on Pg. 4 Section C. that the FNFM model uses the idea of factorization machines in order to learn the interactive features) performing a first optimization on second architecture parameters in the second model to obtain optimized architecture parameters (Zhang Pg. 4 section 2 teaches equation 12, which teaches optimizing architecture parameters by using an equation involving the summation of all aij (second architecture parameters) to generate optimized architecture parameters fbi(Vx).) Zhang does not distinctly disclose: and obtaining, based on the optimized architecture parameters and based on the first model or the second model, a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold. However, Gierach teaches: and obtaining, based on the optimized architecture parameters and based on the first model or the second model, a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold. (Gierach teaches: “The modified second K-th order attribution model (reads on third model) is effectively the K-th order attribution model (reads on first model) with the set of insignificant second order cross terms (reads on feature interaction item deletion) removed” [0019]. Furthermore, it teaches doing so based on the optimized architecture parameters: “M independent variables which include at least zero-th order terms, first order linear terms and second order cross terms” and “ M independent variables weighted by one of the M first order model parameters (reads on optimized architecture parameters by weighing parameters), which are identical to the corresponding terms in the K-th order attribution mode.” [0019]. The weights combined with the variables are understood to be the optimized architecture parameters as the weights of interaction items provide the “importance of a corresponding feature interaction item” and the M independent variables are the feature interaction items. These are considered optimized as they are generated by a calculation by the model for the purpose of using it with the model (optimized to work with the model). The weights generate the linear combination which is used to generate the second k-th order attribution model and removes the insignificant second order cross terms (feature items) as taught by paragraph 0019. “In addition, the second step may also include classifying the second order cross term with the two independent variables X_i and X_j as unimportant if a monotonic non-decreasing function of the fraction B/A is less than a threshold ‘TH4’.” [0032] As explained, the feature interaction item that is deleted is the second order cross terms which is deleted by the fractional comparison B/A which is formed by taking the ratio of two values created by the architecture parameter and is less than a ratio TH4) It would have been obvious to one of ordinary skill in the art to modify the method of claim 1, as disclosed by Zhang to also include obtaining, based on the optimized architecture parameters and based on the first model or the second model, a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold, as disclosed by Gierach. One of ordinary skill in the art would have been motivated to make this modification to enable the improvement of removing features which do not match a certain threshold, in order to maintain and/or enhance quality of the system. (Gierach, Par. [0032], “In addition, the second step may also include classifying the second order cross term with the two independent variables X_i and X_j as unimportant if a monotonic non-decreasing function of the fraction B/A is less than a threshold ‘TH4’. A second order cross term being classified as unimportant by either the first step and/or the second step may be identified as insignificant.”) Regarding Claim 2, Zhang as modified by Gierach teaches all the limitations of Claim 1, and Zhang further teaches: The method of claim 1, wherein the first optimization allows the optimized architecture parameters to be sparse. (Zhang teaches: “In order to be able to cross and combine dense numerical features with sparse category features, dense numerical features can also be compressed into low-dimensional spaces by embedding techniques” [Page 3] which teaches a step in the optimization of the architecture parameters which allows the data to be handled when sparse (xi which generates ai which is used then used to generate the optimized architecture parameters as shown in claim 1.) Regarding Claim 3, Zhang as modified by Gierach teaches all the limitations of Claim 2, and Gierach further teaches: The method of claim 2, wherein the threshold is based on an application requirement (Gierach, Par. [0032], “In addition, the second step may also include classifying the second order cross term with the two independent variables X_i and X_j as unimportant if a monotonic non-decreasing function of the fraction B/A is less than a threshold ‘TH4’. A second order cross term being classified as unimportant by either the first step and/or the second step may be identified as insignificant.”, thus, the threshold TH4 is based on an application requirement for classifying variables as important/unimportant). The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 4, Zhang as modified by Gierach teaches all the limitations of Claim 1, but Zhang does not further teach: The method of claim 1, wherein a value of at least one of the first architecture parameters is equal to zero after completing the first optimization. Zhang teaches sparsifying input features (See Pg. 3 Section 3) but does not distinctly describe the values of the vector V (vector of architecture parameters) used in equation 12 to calculate the optimized parameters. However, Gierach does teach wherein a value of at least one of the first architecture parameters is equal to zero after completing the first optimization (“The first order model parameter 424 may be a scalar which may be different for different first order cross terms. It may be positive, negative or a zero.” This teaches the weight parameter related to the importance of the feature, to being a value of zero (Paragraph 0019 teaches the first order model parameter being used as a weight parameter). Which teaches the parameter value to be zero before optimization and having it unchanged after optimization.) It would have been obvious to one of ordinary skill in the art to modify the method of claim 1, as disclosed by Zhang in view of Gierach to include wherein a value of at least one of the first architecture parameters is equal to zero after completing the first optimization, as disclosed by Gierach. One of ordinary skill in the art would be motivated to make this modification to enable the improvement of labeling the feature item parameters as insignificant (Gierach, Par. [0219], “The first order model parameter 424 may be set to zero for the insignificant first order cross terms.” in which insignificant terms have a weight of zero) Regarding Claim 6, Zhang as modified by Gierach teaches all the limitations of Claim 1, and Zhang further teaches: The method of claim 1, further comprising performing a second optimization on model parameters in the second model, wherein the second optimization comprises scalarization processing on the model parameters. (Zhang teaches a further step of optimization within the Normalization layer on Page 4. Also as taught by Zhang: “Batch normalization [19] is an adaptive reparameterization method that is mainly used to solve the problem that the gradient of the deep neural network disappears during the training process. Batch normalization can transform the data into a statistical distribution with a mean of 0 variance of 1”. [Page 4] The transformation of the data into a distribution is understood to be an scalarization process and as taught in the quotation above, the gradient of the deep neural network is understood to be model parameters of the second model) Regarding Claim 7, Zhang as modified by Gierach teaches all the limitations of Claim 6, and further teaches: The method of claim 6, wherein the second optimization further comprises batch normalization (BN) processing on the model parameters, (As taught by Zhang: “Batch normalization [19] is an adaptive reparameterization method that is mainly used to solve the problem that the gradient of the deep neural network disappears during the training process” [Page 4] which teaches using batch normalization on the parameters within a deep neural network.) wherein performing the first optimization and performing the second optimization comprises simultaneously performing the first optimization and the second optimization using the same training data, (Zhang teaches layer “BI-INTERACTION ATTENTION LAYER” and layer “NORMALIZATION LAYER” on Page 4 which are performed right after each other and involve the same training data input which is input in the input layer as seen on page 3 of Zhang.) and wherein the method further comprises training the third model to obtain a click-through rate (CTR) prediction model or a conversion rate (CVR) prediction model. (Gierach teaches the model being used as click-through rate prediction model. Figures 16A teaches the generation of the third model (Second K-th order attribution model). Figured 16B which directly connects to the creation of the model includes step 1634 of the system generating model parameters the parameters are then used with the attribution server 100 as taught by paragraph 0110. Finally Paragraph 0159-160 teach that the attribution server tracks the number of clicks involve with an advertising campaign, which is understood to be the click-through rate (how often are elements of the campaign being clicked). To summarize, the third model that is trained in Gierach is then a model that is used to generate parameters for an attribution server. Therefore, the third model is a click-through rate model as it generated parameters directly for click-through rate prediction.) The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 8, Zhang teaches: A method (Zhang, Pg. 1, Abstract, “This paper combines traditional feature combination methods and deep neural networks to automate feature combinations to improve the accuracy of click through rate prediction. We propose a mechanism named ’Field-aware Neural Factorization Machine’ (FNFM).”, thus, methods are disclosed) comprising: adding first architecture parameters to first feature interaction items in a first model to obtain to obtain a second model, wherein the first model is a factorization machine (FM)-based model, and wherein each of the first architecture parameters is a trainable value reprsenting an importance of a corresponding feature interaction item; (Zhang teaches first architecture parameters Vi (Page 4 Section 2) which are added to interaction items (vector x on Page 3). Both the first model and the second model are understood to be factorization machines: “The FNFM model uses the idea of factorization machine to learn the expression of second-order interactive features (reads on second architecture parameters) in the form of hidden vector products.” [Page 4], in which the second model is the ”learned” version of the first model with the expression of second-order interaction features. The embedding vector is understood to be related to the importance of the feature interaction item, as taught in “Space Optimized Mode” on page 2, where each vi is tied to each feature item and is an embedding value for it. Further, the architecture parameters are trainable, as it is stated on Pg. 4 Section C. that the FNFM model uses the idea of factorization machines in order to learn the interactive features) performing a first optimization on second architecture parameters in the second model to obtain optimized architecture parameters; (Zhang teaches equation 12, which teaches optimizing architecture parameters by using an equation involving the summation of all aij (second architecture parameters) to generate optimized architecture parameters fbi(Vx).) Zhang does not distinctly disclose: obtaining, based on the optimized architecture parameters and based on the first model or the second model, a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold. training the third model using a training sample of a target object to obtain a click-through rate (CTR) prediction model or a conversion rate (CVR) prediction model; inputting data of the target object into the CTR prediction model or the CVR prediction model to obtain a prediction result of the target object; and determining a recommendation result of the target object based on the prediction result. However, Gierach teaches: obtaining, based on the optimized architecture parameters and based on the first model or the second model, a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold (Gierach teaches: “The modified second K-th order attribution model (reads on third model) is effectively the K-th order attribution model (reads on first model) with the set of insignificant second order cross terms (reads on feature interaction item deletion) removed” [0019]. Furthermore, it teaches doing so based on the optimized architecture parameters: “M independent variables which include at least zero-th order terms, first order linear terms and second order cross terms” and “ M independent variables weighted by one of the M first order model parameters (reads on optimized architecture parameters by weighing parameters), which are identical to the corresponding terms in the K-th order attribution mode.” [0019]. The weights combined with the variables are understood to be the optimized architecture parameters as the weights of interaction items provide the “importance of a corresponding feature interaction item” and the M independent variables are the feature interaction items. These are considered optimized as they are generated by a calculation by the model for the purpose of using it with the model (optimized to work with the model). The weights generate the linear combination which is used to generate the second k-th order attribution model and removes the insignificant second order cross terms (feature items) as taught by paragraph 0019. “In addition, the second step may also include classifying the second order cross term with the two independent variables X_i and X_j as unimportant if a monotonic non-decreasing function of the fraction B/A is less than a threshold ‘TH4’.” [0032] As explained, the feature interaction item that is deleted is the second order cross terms which is deleted by the fractional comparison B/A which is formed by taking the ratio of two values created by the architecture parameter and is less than a ratio TH4) training the third model using a training sample of a target object to obtain a click-through rate (CTR) prediction model or a conversion rate (CVR) prediction model; (Gierach teaches: “K-th order attribution model 410 to estimate the dependent variable Y_1 (e.g., dependent variable 414) in an optimal way in terms of some goodness-of-fit measure, based on the marketing data used as training data in the regression analysis” [0227] which teaches the training data used to train the third model being marketing data. Gierach also teaches the trained model being used as click-through rate prediction model. Figures 16A teaches the generation of the third model (Second K-th order attribution model). Figured 16B which directly connects to the creation of the model includes step 1634 of the system generating model parameters the parameters are then used with the attribution server 100 as taught by paragraph 0110. Finally Paragraph 0159-160 teach that the attribution server tracks the number of clicks involve with an advertising campaign, which is understood to be the click-through rate (how often are elements of the campaign being clicked). To summarize, the third model that is trained in Gierach is then a model that is used to generate parameters for an attribution server. Therefore, the third model is a click-through rate model as it generated parameters directly for click-through rate prediction.) inputting data of the target object into the CTR prediction model or the CVR prediction model to obtain a prediction result of the target object; (Gierach teaches figure 1, which teaches a request 146 as an input into the system which the attribution server 100 (which is the server that runs the third model) receives and outputs an attribution score (prediction result) as taught by paragraph 0053 for the target object (“attribution scores associated with the P publishing channels “[0053].) and determining a recommendation result of the target object based on the prediction result. (Gierach teaches the outputted attribution scores being sent to the marketer [Figure 7] which are scores tied to the marketing material and therefore constitute a recommendation based on what the scores are.) It would have been obvious to one of ordinary skill in the art to modify the method of claim 8, as disclosed by Zhang to include obtaining, based on the optimized architecture parameters and based on the first model or the second model, a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold and training the third model using a training sample of a target object to obtain a click-through rate (CTR) prediction model or a conversion rate (CVR) prediction model, as discosed by Gierach. One of ordinary skill in the art would be motivated to make this modification to enable the improvement of removing features which do not match a certain threshold to maintain and/or enhance quality of the system and obtain a trainable third model which may be further optimized by the system (Gierach, Par. [0032], “In addition, the second step may also include classifying the second order cross term with the two independent variables X_i and X_j as unimportant if a monotonic non-decreasing function of the fraction B/A is less than a threshold ‘TH4’. A second order cross term being classified as unimportant by either the first step and/or the second step may be identified as insignificant.”) Regarding Claim 9, Zhang as modified by Gierach teaches all the limitations of Claim 8, and Zhang further teaches: The method of claim 8, wherein the first optimization allows the optimized architecture parameters to be sparse. (Zhang teaches: “In order to be able to cross and combine dense numerical features with sparse category features, dense numerical features can also be compressed into low-dimensional spaces by embedding techniques” [Page 3] which teaches a step in the optimization of the architecture parameters which allows the data to be handled when sparse (xi which generates ai which is used then used to generate the optimized architecture parameters as shown in claim 8.) Regarding Claim 11, Zhang as modified by Gierach teaches all the limitations of Claim 8, but Zhang does not further teach: The method of claim 8, wherein a value of at least one of the first architecture parameters is equal to zero after completing the first optimization. Zhang teaches sparsifying input features (See Pg. 3 Section 3) but does not distinctly describe the values of the vector V (vector of architecture parameters) used in equation 12 to calculate the optimized parameters. However, Gierach does teach wherein a value of at least one of the first architecture parameters is equal to zero after completing the first optimization (“The first order model parameter 424 may be a scalar which may be different for different first order cross terms. It may be positive, negative or a zero.” This teaches the weight parameter related to the importance of the feature, to being a value of zero (Paragraph 0019 teaches the first order model parameter being used as a weight parameter). Which teaches the parameter value to be zero before optimization and having it unchanged after optimization) It would have been obvious to one of ordinary skill in the art to modify the method of claim 8, as disclosed by Zhang in view of Gierach to include wherein a value of at least one of the first architecture parameters is equal to zero after completing the first optimization, as disclosed by Gierach. One of ordinary skill in the art would be motivated to make this modification to enable the improvement of labeling the feature item parameters as insignificant (Gierach, Par. [0219], “The first order model parameter 424 may be set to zero for the insignificant first order cross terms.” in which insignificant terms have a weight of zero) Regarding Claim 13, Zhang as modified by Gierach teaches all the limitations of Claim 8, and Zhang further teaches: The method of claim 8, further comprising performing a second optimization on model parameters in the second model, wherein the second optimization comprises scalarization processing on the model parameters. (Zhang teaches a further step of optimization within the Normalization layer on Page 4. Also as taught by Zhang: “Batch normalization [19] is an adaptive reparameterization method that is mainly used to solve the problem that the gradient of the deep neural network disappears during the training process. Batch normalization can transform the data into a statistical distribution with a mean of 0 variance of 1”. [Page 4] The transformation of the data into a distribution is understood to be an scalarization process and as taught in the quotation above, the gradient of the deep neural network is understood to be model parameters of the second model) Regarding Claim 14, Zhang as modified by Gierach teaches all the limitations of Claim 13, and Zhang further teaches: The method of claim 13, wherein the second optimization further comprises batch normalization (BN) processing on the model parameters, (As taught by Zhang: “Batch normalization [19] is an adaptive reparameterization method that is mainly used to solve the problem that the gradient of the deep neural network disappears during the training process” [Page 4] which teaches using batch normalization on the parameters within a deep neural network.) and wherein performing the first optimization and performing the second optimization comprises simultaneously performing the first optimization and the second optimization using the same training data. (Zhang teaches layer “BI-INTERACTION ATTENTION LAYER” and layer “NORMALIZATION LAYER” on Page 4 which are performed right after each other and involve the same training data input which is input in the input layer as seen on page 3 of Zhang.) Regarding Claim 15, Zhang teaches: add first architecture parameter to first feature interaction items in a first model to obtain a second model, wherein the first model is a factorization machine (FM)-based model, and wherein each of the first architecture parameters is a trainable value representing an importance of a corresponding feature interaction item; (Zhang teaches first architecture parameters Vi (Page 4 Section 2) which are added to interaction items (vector x on Page 3). Both the first model and the second model are understood to be factorization machines: “The FNFM model uses the idea of factorization machine to learn the expression of second-order interactive features (reads on second architecture parameters) in the form of hidden vector products.” [Page 4], in which the second model is the ”learned” version of the first model with the expression of second-order interaction features. The embedding vector is understood to be related to the importance of the feature interaction item, as taught in “Space Optimized Mode” on page 2, where each vi is tied to each feature item and is an embedding value for it. Further, the architecture parameters are trainable, as it is stated on Pg. 4 Section C. that the FNFM model uses the idea of factorization machines in order to learn the interactive features) perform a first optimization on second architecture parameters in the second model to obtain optimized architecture parameters; (Zhang teaches equation 12, which teaches optimizing architecture parameters by using an equation involving the summation of all aij (second architecture parameters) to generate optimized architecture parameters fbi(Vx).) Zhang does not distinctly disclose: An electronic device comprising: a memory configured to store instructions; and a processor coupled to the memory and configured to execute the instructions to cause the electronic device to: […] However, Gierach teaches: An electronic device comprising: a memory configured to store instructions; and a processor coupled to the memory and configured to execute the instructions to cause the electronic device to: […] (Gierach teaches Figure 7 which show system (100) comprising the Memory (104) and Processor (102) for the system that runs processes for models) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the device of claim 15, as disclosed by Zhang to include where the device is an electronic device comprising: a memory configured to store instructions; and a processor coupled to the memory and configured to execute the instructions to cause the electronic device to: […], as disclosed by Gierach. One of ordinary skill in the art would have been motivated to make this modification to enable the use of an electronic device, comprising a memory and processor, which may efficiently process data and facilitate efficient data communications (Gierach, Par. [0009], “Disclosed are a method, a device and/or a system of non-converting publisher attribution weighting and analytics server. In one aspect, a method, a device and/or a system of a non-converting publisher attribution weighting and analytics server includes determining ‘P’ number of publishing channels for advertisements in a first marketing campaign for a set of purchasable items using a processor and a memory communicatively coupled with the processor. Further, the method the device and/or the system monitors a marketing effectiveness of the P publishing channels in generating converted users each with a desirable action and/or a purchase from the set of purchasable items in the first marketing campaign. The marketing effectiveness of the P publishing channels is analyzed using the processor and the memory based on a set of marketing data from a data collection server in a cloud.”) Zhang does not distinctly disclose: and obtain, based on the optimized architecture parameters and based on the first model or the second model, a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold. Gierach also teaches: and obtain, based on the optimized architecture parameters and based on the first model or the second model, a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold. (Gierach teaches: “The modified second K-th order attribution model (reads on third model) is effectively the K-th order attribution model (reads on first model) with the set of insignificant second order cross terms (reads on feature interaction item deletion) removed” [0019]. Furthermore, it teaches doing so based on the optimized architecture parameters: “M independent variables which include at least zero-th order terms, first order linear terms and second order cross terms” and “ M independent variables weighted by one of the M first order model parameters (reads on optimized architecture parameters by weighing parameters), which are identical to the corresponding terms in the K-th order attribution mode.” [0019]. The weights combined with the variables are understood to be the optimized architecture parameters as the weights of interaction items provide the “importance of a corresponding feature interaction item” and the M independent variables are the feature interaction items. These are considered optimized as they are generated by a calculation by the model for the purpose of using it with the model (optimized to work with the model). The weights generate the linear combination which is used to generate the second k-th order attribution model and removes the insignificant second order cross terms (feature items) as taught by paragraph 0019. “In addition, the second step may also include classifying the second order cross term with the two independent variables X_i and X_j as unimportant if a monotonic non-decreasing function of the fraction B/A is less than a threshold ‘TH4’.” [0032] As explained, the feature interaction item that is deleted is the second order cross terms which is deleted by the fractional comparison B/A which is formed by taking the ratio of two values created by the architecture parameter and is less than a ratio TH4) It would have been obvious to one of ordinary skill in the art to modify the device of claim 15, as disclosed by Zhang in view of Gierach to also include obtaining, based on the optimized architecture parameters and based on the first model or the second model, a third model by deleting a second feature interaction item corresponding to one of the optimized architecture parameters whose value is less than a threshold, as disclosed by Gierach. One of ordinary skill in the art would have been motivated to make this modification to enable the improvement of removing features which do not match a certain threshold, in order to maintain and/or enhance quality of the system. (Gierach, Par. [0032], “In addition, the second step may also include classifying the second order cross term with the two independent variables X_i and X_j as unimportant if a monotonic non-decreasing function of the fraction B/A is less than a threshold ‘TH4’. A second order cross term being classified as unimportant by either the first step and/or the second step may be identified as insignificant.”) Regarding Claim 17, Zhang as modified by Gierach teaches all the limitations of Claim 15, but Zhang does not further teach: The electronic device of claim 15, wherein a value of at least one of the first architecture parameters is equal to zero after completing the first optimization. Zhang teaches sparsifying input features (See Pg. 3 Section 3) but does not distinctly describe the values of the vector V (vector of architecture parameters) used in equation 12 to calculate the optimized parameters. However, Gierach does teach wherein a value of at least one of the first architecture parameters is equal to zero after completing the first optimization (“The first order model parameter 424 may be a scalar which may be different for different first order cross terms. It may be positive, negative or a zero.” This teaches the weight parameter related to the importance of the feature, to being a value of zero (Paragraph 0019 teaches the first order model parameter being used as a weight parameter). Which teaches the parameter value to be zero before optimization and having it unchanged after optimization.) It would have been obvious to one of ordinary skill in the art to modify the device of claim 15, as disclosed by Zhang in view of Gierach to include wherein a value of at least one of the first architecture parameters is equal to zero after completing the first optimization, as disclosed by Gierach. One of ordinary skill in the art would be motivated to make this modification to enable the improvement of labeling the feature item parameters as insignificant (Gierach, Par. [0219], “The first order model parameter 424 may be set to zero for the insignificant first order cross terms.” in which insignificant terms have a weight of zero) Regarding Claim 19, Zhang as modified by Gierach teaches all the limitations of Claim 15, and Zhang further teaches: The electronic device of claim 15, wherein the processor is further configured to execute the instructions to cause the electronic device to perform a second optimization on model parameters in the second model, wherein the second optimization comprises scalarization processing on the model parameters. (Zhang teaches a further step of optimization within the Normalization layer on Page 4. Also as taught by Zhang: “Batch normalization [19] is an adaptive reparameterization method that is mainly used to solve the problem that the gradient of the deep neural network disappears during the training process. Batch normalization can transform the data into a statistical distribution with a mean of 0 variance of 1”. [Page 4] The transformation of the data into a distribution is understood to be an scalarization process and as taught in the quotation above, the gradient of the deep neural network is understood to be model parameters of the second model) Regarding Claim 20, Zhang as modified by Gierach teaches all the limitations of Claim 19, and further teaches: The electronic device according to claim 19, wherein the first optimization allows the optimized architecture parameters to be sparse, (Zhang teaches: “In order to be able to cross and combine dense numerical features with sparse category features, dense numerical features can also be compressed into low-dimensional spaces by embedding techniques” [Page 3] which teaches a step in the optimization of the architecture parameters which allows the data to be handled when sparse (xi which generates ai which is used then used to generate the optimized architecture parameters as shown in claim 1.) wherein the second optimization comprises batch normalization (BN) processing on the model parameters, (As taught by Zhang: “Batch normalization [19] is an adaptive reparameterization method that is mainly used to solve the problem that the gradient of the deep neural network disappears during the training process” [Page 4] which teaches using batch normalization on the parameters within a deep neural network.) and wherein the processor is further configured to execute the instructions to cause the electronic device to: simultaneously perform the first optimization and the second optimization using the same training data; (Zhang teaches layer “BI-INTERACTION ATTENTION LAYER” and layer “NORMALIZATION LAYER” on Page 4 which are performed right after each other and involve the same training data input which is input in the input layer as seen on page 3 of Zhang.) and train the third model to obtain a click-through rate (CTR) prediction model or a conversion rate (CVR) prediction model. (Gierach teaches the model being used as click-through rate prediction model. Figures 16A teaches the generation of the third model (Second K-th order attribution model). Figured 16B which directly connects to the creation of the model includes step 1634 of the system generating model parameters the parameters are then used with the attribution server 100 as taught by paragraph 0110. Finally Paragraph 0159-160 teach that the attribution server tracks the number of clicks involve with an advertising campaign, which is understood to be the click-through rate (how often are elements of the campaign being clicked). To summarize, the third model that is trained in Gierach is then a model that is used to generate parameters for an attribution server. Therefore, the third model is a click-through rate model as it generated parameters directly for click-through rate prediction.) The reasons of obviousness have been noted in the rejection of Claim 15 above and applicable herein. 9. Claims 5, 12, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (hereinafter Zhang) (“Field-Aware Neural Factorization Machine for Click-Through Rate Prediction”), in view of Gierach et al. (hereinafter Gierach) (US PG-PUB 20180075482), further in view of Chao et al. (hereinafter Chao) (“A Generalization of Regularized Dual Averaging and Its Dynamics”). Regarding Claim 5, Zhang as modified by Gierach teaches all the limitations of Claim 4, and Zhang further teaches: The method of claim 4, further comprising performing the first optimization on the second architecture parameters (As described in Claim 1, Zhang teaches optimization which is performed on the second architecture parameters to generate optimized parameters.) However, Zhang as modified by Gierach does not distinctly disclose: using a generalized regularized dual averaging (gRDA) optimizer, wherein the gRDA optimizer allows values of the second architecture parameters to tend to zero during the first optimization However, Chao teaches using a generalized regularized dual averaging (gRDA) optimizer, wherein the gRDA optimizer allows values of the second architecture parameters to tend to zero during the first optimization (Chao, Pg. 1, “To de velop statistical inference in this case, we propose a class of gener alized regularized dual averaging (gRDA) algorithms with constant step size, which improves RDA (Xiao, 2010; Flammarion and Bach, 2017).”, Chao teaches using an gRDA optimizer that as epochs passes start to bring values toward zero (non-zeros variables in the graph are decreasing to zero as epochs (time) passes.) The gRDA formula is taught on page 3 with the illustrated usage of the system taught on page 4, which as described earlier brings the inputs to tend to zero) PNG media_image1.png 304 390 media_image1.png Greyscale It would have been obvious to one of ordinary skill in the art to modify the method of claim 4, as disclosed by Zhang in view of Gierach to instead include using a generalized regularized dual averaging (gRDA) optimizer, wherein the gRDA optimizer allows values of the second architecture parameters to tend to zero during the first optimization, as disclosed by Chao. One of ordinary skill in the art would have been motivated to make this modification for the improvement of using an optimizer that has the ability of performing penalization through a tuning function. (Chao, Pg. 3, “This observation motivates us to design a class of new algorithms, named as generalized RDA (gRDA), that can adjust the level of penalization with time through a tuning function g(n, γ)”) Regarding Claim 12, Zhang as modified by Gierach teaches all the limitations of Claim 11, and Zhang further teaches: The method of claim 11, further comprising performing the first optimization on the second architecture parameters (As described in Claim 8, Zhang teaches optimization which is performed on the second architecture parameters to generate optimized parameters.) However, Zhang as modified by Gierach does not distinctly disclose: using a generalized regularized dual averaging (gRDA) optimizer, wherein the gRDA optimizer allows values of the second architecture parameters to tend to zero during the first optimization However, Chao teaches using a generalized regularized dual averaging (gRDA) optimizer, wherein the gRDA optimizer allows values of the second architecture parameters to tend to zero during the first optimization (Chao, Pg. 1, “To de velop statistical inference in this case, we propose a class of gener alized regularized dual averaging (gRDA) algorithms with constant step size, which improves RDA (Xiao, 2010; Flammarion and Bach, 2017).”, Chao teaches using an gRDA optimizer that as epochs passes start to bring values toward zero (non-zeros variables in the graph are decreasing to zero as epochs (time) passes.) The gRDA formula is taught on page 3 with the illustrated usage of the system taught on page 4, which as described earlier brings the inputs to tend to zero) PNG media_image1.png 304 390 media_image1.png Greyscale It would have been obvious to one of ordinary skill in the art to modify the method of claim 11, as disclosed by Zhang in view of Gierach to instead include using a generalized regularized dual averaging (gRDA) optimizer, wherein the gRDA optimizer allows values of the second architecture parameters to tend to zero during the first optimization, as disclosed by Chao. One of ordinary skill in the art would have been motivated to make this modification for the improvement of using an optimizer that has the ability of performing penalization through a tuning function. (Chao, Pg. 3, “This observation motivates us to design a class of new algorithms, named as generalized RDA (gRDA), that can adjust the level of penalization with time through a tuning function g(n, γ)”) Regarding Claim 18, Zhang as modified by Gierach teaches all the limitations of Claim 17, and Zhang further teaches: The electronic device of claim 17, wherein the processor is further configured to execute the instructions to cause the electronic device to perform the first optimization on the second architecture parameters (As described in Claim 15, Zhang teaches optimization which is performed on the second architecture parameters to generate optimized parameters.) However, Zhang as modified by Gierach does not distinctly disclose: using a generalized regularized dual averaging (gRDA) optimizer, wherein the gRDA optimizer allows values of the second architecture parameters to tend to zero during the first optimization However, Chao teaches using a generalized regularized dual averaging (gRDA) optimizer, wherein the gRDA optimizer allows values of the second architecture parameters to tend to zero during the first optimization (Chao, Pg. 1, “To de velop statistical inference in this case, we propose a class of gener alized regularized dual averaging (gRDA) algorithms with constant step size, which improves RDA (Xiao, 2010; Flammarion and Bach, 2017).”, Chao teaches using an gRDA optimizer that as epochs passes start to bring values toward zero (non-zeros variables in the graph are decreasing to zero as epochs (time) passes.) The gRDA formula is taught on page 3 with the illustrated usage of the system taught on page 4, which as described earlier brings the inputs to tend to zero) PNG media_image1.png 304 390 media_image1.png Greyscale It would have been obvious to one of ordinary skill in the art to modify the device of claim 17, as disclosed by Zhang in view of Gierach to instead include using a generalized regularized dual averaging (gRDA) optimizer, wherein the gRDA optimizer allows values of the second architecture parameters to tend to zero during the first optimization, as disclosed by Chao. One of ordinary skill in the art would have been motivated to make this modification for the improvement of using an optimizer that has the ability of performing penalization through a tuning function. (Chao, Pg. 3, “This observation motivates us to design a class of new algorithms, named as generalized RDA (gRDA), that can adjust the level of penalization with time through a tuning function g(n, γ)”) 10. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (hereinafter Zhang) (“Field-Aware Neural Factorization Machine for Click-Through Rate Prediction”), in view of Gierach et al. (hereinafter Gierach) (US PG-PUB 20180075482), further in view of Louizos et al. (hereinafter Louizos) (“Learning Sparse Neural Networks Through L0 Regularizations”). Regarding Claim 10, Zhang as modified by Gierach teaches all the limitations of Claim 9. Zhang in view of Gierach does not explicitly disclose comparing the one of the optimized architecture parameters to the threshold using a selection gate mechanism. However, Louizos teaches comparing the one of the optimized architecture parameters to the threshold using a selection gate mechanism (Louizos, Pg. 1, Abstract, “We propose a solution through the inclusion of a collection of non-negative stochastic gates, which collectively determine which weights to set to zero. We show that, somewhat surprisingly, for certain distributions over the gates, the expected L0 regularized objective is differentiable with respect to the distribution parameters. We further propose the hard concrete distribution for the gates, which is obtained by “stretching” a binary concrete distribution and then transforming its samples with a hard-sigmoid. The parameters of the distribution over the gates can then be jointly optimized with the original network parameters.” & Pg. 2 “where the zj correspond to binary “gates” that denote whether a parameter is present and the L0 norm corresponds to the amount of gates being “on”. By letting q(zj|πj) = Bern(πj) be a Bernoulli distribution over each gate zj we can reformulate the minimization of Eq. 1 as penalizing the number of parameters being used, on average, as follows […]”, therefore, one or more parameters are compared to the threshold using a selection gate mechanism, which selects which parameters to set to zero). It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 9, as disclosed by Zhang in view of Gierach to include comparing the one of the optimized architecture parameters to the threshold using a selection gate mechanism, as disclosed by Louizos. One of ordinary skill in the art would have been motivated to make this modification to enable the use of gating which provides a straightforward and efficient learning of model structures and allows for conditional computation, hence reducing resource consumption (Louizos, Pg. 1, Abstract, “We further propose the hard concrete distribution for the gates, which is obtained by “stretching” a binary concrete distribution and then transforming its samples with a hard-sigmoid. The parameters of the distribution over the gates can then be jointly optimized with the original network parameters. As a result our method allows for straightforward and efficient learning of model structures with stochastic gradient descent and allows for conditional computation in a principled way. We perform various experiments to demonstrate the effectiveness of the resulting approach and regularizer.”). 11. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (hereinafter Zhang) (“Field-Aware Neural Factorization Machine for Click-Through Rate Prediction”), in view of Gierach et al. (hereinafter Gierach) (US PG-PUB 20180075482), further in view of Deng et al. (hereinafter Deng) (“A Sparse Deep Factorization Machine for Efficient CTR Prediction”). Regarding Claim 16, Zhang as modified by Gierach teaches all the limitations of Claim 15. Zhang in view of Gierach does not explicitly disclose wherein the FM-based model comprises a DeepFM model. However, Deng teaches wherein the FM-based model comprises a DeepFM model (Deng, Pg. 2, Section 3, “Our model is a direct improvement of DeepFM [12], where the latter has the following formulation: ϕDeepFM(w,v,e)=ϕDeep(w,e)+ϕFM(v,e).”, thus, the FM-based model comprises a DeepFM model) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the device of claim 15, as disclosed by Zhang in view of Gierach to include wherein the FM-based model comprises a DeepFM model, as disclosed by Deng. One of ordinary skill in the art would have been motivated to make this modification to enable the use of DeepFM which may efficiently model low-order and high-order feature interactions through the use of a factorization machine component (Deng, Pg. 1, “The embedding-based neural networks provide a more powerful non-linear modeling by using DNNs. Wide & Deep [4] proposed to train a joint network that combines a linear model and a DNN model to learn both low-order and high-order feature interactions. However, the cross features in the linear model still require expertise feature engineering and cannot be easily adapted to new datasets. DeepFM [12] handled this issue by modeling low-order feature interactions through the FM component instead of the linear model.”). Conclusion 12. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. 13. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Devika S Maharaj whose telephone number is (571)272-0829. The examiner can normally be reached Monday - Thursday 8:30am - 5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DEVIKA S MAHARAJ/Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Sep 20, 2022
Application Filed
Oct 10, 2022
Response after Non-Final Action
Aug 04, 2025
Non-Final Rejection mailed — §103
Oct 31, 2025
Response Filed
Jul 30, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694269
SELECTIVE REPORTING OF MACHINE LEARNING PARAMETERS FOR FEDERATED LEARNING
4y 0m to grant Granted Jul 28, 2026
Patent 12682205
DIFFERENTIAL EQUATIONS NETWORK
7y 9m to grant Granted Jul 14, 2026
Patent 12682215
FLEXIBLE MACHINE LEARNING
4y 1m to grant Granted Jul 14, 2026
Patent 12675689
MULTI-DOMAIN FEATURE ENHANCEMENT FOR TRANSFER LEARNING (FTL)
4y 4m to grant Granted Jul 07, 2026
Patent 12657441
SPIKING NEURAL NETWORK DEVICE THAT UPDATES SYNAPIC WEIGHT BASED ON OUTPUT FREQUENCY AND LEARNING METHOD OF SPIKING NEURAL NETWORK DEVICE
5y 9m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
56%
Grant Probability
65%
With Interview (+9.3%)
4y 7m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 86 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month