Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim interpretation regarding to claim 19 is withdrawn.
Claim rejections related to 112(a) and 112(b) regarding to claims 5-6, 14-15 are withdrawn.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20, 22 are rejected under 35 U.S.C. 101
because the claimed invention is directed to an abstract idea without significantly
more.
When considering subject matter eligibility under 35 U.S.C. 101, it must be
determined whether the claim is directed to one of the four statutory categories of
invention, i.e., process, machine, manufacture, or composition of matter (Step 1). If the
claim does fall within one of the statutory categories, the second step in the analysis is
to determine whether the claim is directed to a judicial exception (Step 2A). The Step 2A
analysis is broken into two prongs. In the first prong (Step 2A, Prong 1), it is determined
whether or not the claims recite a judicial exception (e.g., mathematical concepts,
mental processes, certain methods of organizing human activity). If it is determined in
Step 2A, Prong 1 that the claims recite a judicial exception, the analysis proceeds to the
second prong (Step 2A, Prong 2), where it is determined whether or not the claims
integrate the judicial exception into a practical application. If it is determined at step 2A,
Prong 2 that the claims do not integrate the judicial exception into a practical
application, the analysis proceeds to determining whether the claim is a patent-eligible
application of the exception (Step 2B). If an abstract idea is present in the claim, any
element or combination of elements in the claim must be sufficient to ensure that the
claim integrates the judicial exception into a practical application, or else amounts to
significantly more than the abstract idea itself. Applicant is advised to consult the 2019
PEG for more details of the analysis.
Step 1
According to the first part of the analysis, in the instant case, claims 1-9, 22, 10-18, 19, 20 are directed to a method, a node, a node and a computer program product of building a ML model. Thus, each of the claims falls within one of the four statutory categories (i.e. process, machine, manufacture, or composition of matter). Step 2A,
Step 2A, Prong 1
Following the determination of whether or not the claims fall within one of the four
categories (Step 1), it must be determined if the claims recite a judicial exception (e.g.
mathematical concepts, mental processes, certain methods of organizing human
activity) (Step 2A, Prong 1). In this case, the claims are determined to recite a judicial
exception as explained below.
Regarding Claims 1, 10 and 19, 20 these claims recite
training a ML model using a set of input data, wherein the ML model comprises a plurality of layers and each layer comprises a plurality of filters, and wherein the set of input data comprises class labels;
obtaining a set of output data from training the ML model, wherein the set of output data includes class probabilities values;
determining, for each layer in the ML model, by using the class labels and the class probabilities values, a working value for each filter in the layer;
determining, for each layer in the ML model, a dominant filter, wherein the dominant filter is determined based on whether the working value for the filter exceeds a threshold; and
building a subset ML model based on each dominant filter for each layer, wherein the subset ML model is a subset of the ML model.
The claims recite a mental process. As set forth in MPEP 2106.04(a)(2)(III)(C), “Claims can recite a mental process even if they are claimed as being performed on a computer”. These are recited at a high level such that they could be performed as a human user performing these functions, simply using a computer as a tool-see spec, [0051-0054], Fig. 5, 6. Thus, the claim recites abstract ideas.
Step 2A, Prong 2
Following the determination that the claims recite a judicial exception, it must be
determined if the claims recite additional elements that integrate the exception into a
practical application of the exception (Step 2A, Prong 2). In this case, after considering
all claim elements individually and as an ordered combination, it is determined that the
claims do not include additional elements that integrate the exception into a practical
application of the exception as explained below.
In Prong Two, a claim is evaluated as a whole to determine whether the recited judicial exception is integrated into a practical application of that exception. A claim is not “directed to” a judicial exception, and thus is patent eligible, if the claim as a whole integrates the recited judicial exception into a practical application of that exception. A claim that integrates a judicial exception into a practical application will apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception, such that the claim is more than a drafting effort designed to monopolize the judicial exception. MPEP 2106.04(d). The claims recite an abstract idea and further the claims as a whole does not integrate the recited judicial exception into a practical application of the exception. A claim that integrates a judicial exception into a practical application will apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception, such that the claim is more than a drafting effort designed to monopolize the judicial exception. MPEP 2106.04(d).
Regarding Claims 1, 10, 19, 20 these claims
This limitation recites using one or more neural networks as a tool to perform an
abstract idea, which is not indicative of integration into a practical application. MPEP 2106.05(f).)
This limitation is understood to be generic computer equipment and mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.0S(f))
MPEP § 2106.05(f): Mere Instructions to Apply an Exception. Do the additional element(s) amount to merely the words “apply it” (or an equivalent)
or are mere instructions to implement an abstract idea or other exception on a computer? (Yes)
Step 2B
Based on the determination in Step 2A of the analysis that the claims are
directed to a judicial exception, it must be determined if the claims contain any element
or combination of elements sufficient to ensure that the claim amounts to significantly
more than the judicial exception (Step 2B). In this case, after considering all claim
elements individually and as an ordered combination, it is determined that the claims do
not include additional elements that are sufficient to amount to significantly more than
the judicial exception for the same reasons given above in the Step 2A, Prong 2
analysis. Furthermore, each additional element identified above as being insignificant
extra-solution activity is also well-known, routine, conventional as described below.
Claims 1, 10, 19 and 20: The claims do not include additional elements, alone or in combination, that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements amount to no more than generic computing components and field of use/technological environment which do not amount to significantly more than the abstract idea. The underlying concept merely receives information, analyzes it, and store the results of the analysis – this concept is not meaningfully different than concepts found by the courts to be abstract (see Electric Power Group, collecting information, analyzing it, and displaying certain results of the collection and analysis; see Cybersource, obtaining and comparing intangible data; see Digitech, organizing information through mathematical correlations; see Grams, diagnosing an abnormal condition by performing clinical tests and thinking about the results; see Cyberfone, using categories to organize store and transmit information; see Smartgene, comparing new and stored information and using rules to identify options). The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements when considered both individually and as a combination do not amount to significantly more than the abstract idea. For example, claim 1 recites the additional elements of “training…”, “obtaining…” “determining…”, “determining…”, “building…”
These elements are recited at a high level of generality and are well-understood, routine, and conventional activities in the computer art. Generic computers performing generic computer functions, without an inventive concept, do not amount to significantly more than the abstract idea. Looking at the elements as a combination does not add anything more than the elements analyzed individually. Therefore, these claims do not amount to significantly more than the abstract idea itself.
Step 2A/2B Prong 2 Dependent Claims
Regarding to claim 2, 11
Claim 2, 11 merely recite other additional elements that define storing the data which performing generic functions that when looking at the elements as a combination does not add anything more than the elements analyzed individually. Therefore, these claims also do not amount to significantly more than the abstract idea itself. These claims are not patent eligible.
Regarding to claim 3-4, 7, 12-13, 16
Claim 3-4, 7, 12-13, 16 merely recite other additional elements that define the model which performing generic functions that when looking at the elements as a combination does not add anything more than the elements analyzed individually. Therefore, these claims also do not amount to significantly more than the abstract idea itself. These claims are not patent eligible.
Regarding to claim 5-6, 14-15
Claim 5-6, 14-15 merely recite other additional elements that define the value for the filter and define the filter which performing generic functions that when looking at the elements as a combination does not add anything more than the elements analyzed individually. Therefore, these claims also do not amount to significantly more than the abstract idea itself. These claims are not patent eligible.
Regarding to claim 8-9, 17-18
Claim 8-9, 17-18 merely recite other additional elements that define using the model which performing generic functions that when looking at the elements as a combination does not add anything more than the elements analyzed individually. Therefore, these claims also do not amount to significantly more than the abstract idea itself. These claims are not patent eligible.
Regarding to claim 22
Claim 22 merely recite other additional elements that define the working value for each filter in the layer which performing generic functions that when looking at the elements as a combination does not add anything more than the elements analyzed individually. Therefore, these claims also do not amount to significantly more than the abstract idea itself. These claims are not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claims 1-7, 10-16, 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (Liu) CN 111626330 in view of Li et al. (Li) US 2019/0340510
In regard to claim 1, Liu disclose A computer-implemented method for building a machine learning (ML) model, the method comprising: (abstract, “(1) training YOLOv3 model to generate reference model based on training image data set, using the backbone network Darknet-53 of YOLOv3 to extract features of the image, deep feature through the upper sampling and shallow feature tensor splicing to generate multi-scale feature map;” generate reference model)
training a ML model using a set of input data, wherein the ML model comprises a plurality of layers and each layer comprises prunings, and wherein the set of input data comprises class labels; ((1) training YOLOv3 model to generate reference model based on training image data set, using the backbone network Darknet-53 of YOLOv3 to extract features of the image, deep feature through the upper sampling and shallow feature tensor splicing to generate multi-scale feature map;” (2) performing feature compression along the spatial dimension of the feature map in the step (1); compressing each two-dimensional feature channel into a real number with a global receptive field; the output dimension and the input feature channel number are matched; generating weight for each feature channel through the gating mechanism of the circular neural network, then weighting the weight to the previous characteristic; finishing the re-calibration of the original characteristic on the channel dimension;
As a preference, the step (2) comprises:
(2.1) the step (1) to generate the multi-scale feature map for adaptive sampling, expanding the characteristic diagram of W* H;
(2.2) performing characteristic compression along the space dimension, compressing each two-dimensional characteristic channel into a real number with a global receptive field; the output dimension and the input characteristic channel number are matched; the specific operation is as follows:
wherein W and H are respectively characteristic graph width and height; xc (i, j) represents the appointed element coordinate in the c-th layer channel is (i, j); zc represents the output of the c-th layer channel after being compressed; it is a scalar quantity;..”
“As a preference, in the step (5) introducing softmax function with temperature parameter and knowledge distillation algorithm, the basic model as the teacher network, pruning the model as the student network for migration learning;
The softmax function is defined as:
wherein, zi is the output of neural network i type target detection, .Math. jexp (zj/T) represents the sum of all kinds of output, the ratio of the two is qi, representing the probability value of the ith target, T is temperature parameter;
The loss of the teacher's bounded regression is defined as:
wherein m is the edge distance, yregm represents the real label, Rs is the regression output of the YOLOv3 network after pruning; Rt is the prediction of the initial network, v and [AltContent: rect]is a super-parameter, Ls is a binary cross entropy loss, Lregm is the total regression loss, Lhint is an indication learning, accelerating distillation by indicating learning, using the middle of the teacher as prompting learning to help the training process and improving the distillation effect of the student, using the L2 distance between the characteristic vector V and Z:
wherein Z represents the middle layer selected as the prompt in the teacher network; V represents the output of the guide layer in the student network.” training the ML model using a set of input data, the ML model includes a plurality of layers and each layer includes pruning functions, and the set of input data has labels)
obtaining a set of output data from training the ML model, wherein the set of output data comprises class probabilities values; (“As a preference, in the step (5) introducing softmax function with temperature parameter and knowledge distillation algorithm, the basic model as the teacher network, pruning the model as the student network for migration learning;
The softmax function is defined as:
wherein, zi is the output of neural network i type target detection, .Math. jexp (zj/T) represents the sum of all kinds of output, the ratio of the two is qi, representing the probability value of the ith target, T is temperature parameter;
The loss of the teacher's bounded regression is defined as:
wherein m is the edge distance, yregm represents the real label, Rs is the regression output of the YOLOv3 network after pruning; Rt is the prediction of the initial network, v and is a super-parameter, Ls is a binary cross entropy loss, Lregm is the total regression loss, Lhint is an indication learning, accelerating distillation by indicating learning, using the middle of the teacher as prompting learning to help the training process and improving the distillation effect of the student, using the L2 distance between the characteristic vector V and Z:
wherein Z represents the middle layer selected as the prompt in the teacher network; V represents the output of the guide layer in the student network.”
“step A, as shown in FIG. 2 and FIG. 3, training YOLOv3 model based on training image data set, generating YOLOv3 reference model, using the YOLOv3 backbone network Darknet-53 extraction image characteristic, the deep characteristic is spliced by the upper sampling and the superficial characteristic tensor to generate the multi-scale characteristic diagram; The method specifically comprises:
step A1, using the cross entropy loss function as the optimization target of the model training; calculating the loss function gradient by reverse propagation BP algorithm and updating the model parameter; global loss is Ltotal= ρ Lclass + τ Lreg
wherein, the ρ and τ are hyper-parameter; Lclass is classified loss, expressed as:
wherein D is the training image data set; pc (d) represents the prediction probability of the data set image is classified as c, is the data centralized image is classified as 0-1 binary distribution of c; C is the category number;” generate output from the ML model training and the output have the prediction probabilities of the data set image classified)
determining, for each layer in the ML model, by using the class labels and the class probabilities values, a working value for each filter in the layer; ((4) introducing the γ coefficient of the BN layer in the backbone network into the pruning target function for combined training; normalizing and ordering the training after-γ coefficient; according to the trimming threshold, removing the channel from the model lower than the threshold value of the γ from the model; pruning the YOLOv3 model;
(5) using the model of pruning in the step (4) as the student model; using the benchmark model as the teacher network for knowledge distillation; using the soft label generated by the teacher model to guide the student model to train; and using the indication learning to accelerate the distillation speed;
(6) inputting the image to be detected to the student model trained in the step (5) to detect the target.”
“step D, introducing the γ layer in the backbone network into the pruning target function for combined training; normalizing and ordering the training back γ coefficient; according to the trimming threshold, removing the channel with γ the threshold value according to the channel from the model; pruning the YOLOv3 model; The method specifically comprises:
step D1, introducing the γ layer in the backbone network into the pruning target function for joint training, the conversion function of the BN layer is as follows:
in the formula, zin, zout is input and output of BN μB, is the average value and variance of the input, belongs to a correction parameter close to 0, preventing the sub-is 0, γ is scale factor (scale factor) and β (offset), can linearly convert the output of the BN layer into any scale, recovering the characteristic distribution of the original input, then representing the contribution value of each convolution layer for the input characteristic; measuring the importance of the corresponding convolution layer; therefore, γ as pruning parameter;
the pruning target function is adjusted as:
wherein, Ws is the weight capable of training, xs, ys represents the input and output of the training, n is a hyper-parameter, is Γ set of the coefficient γ the backbone network, f (.) is the loss function of YOLOv3, g (γ) is a penalty function of guiding sparse, wherein g (γ) = |γ |, namely L1 is regularization;
step D2, before training, the γ coefficient presents a positive distribution; after training, the γ coefficient is approaching to 0; normalizing and ordering the training after-γ coefficient; according to the trimming threshold, removing the channel from the model lower than the threshold value of the γ from the model; cutting the branch of the channel without adding operation to the backbone network;” Identifying trimming threshold for each layer in the ML model, by using the labels and the class probabilities values)
wherein the dominant filter is determined based on whether the working value for the filter exceeds a threshold; (“a backbone network compression module, for introducing the γ layer in the backbone network into the pruning target function for combined training; normalizing and ordering the trained γ coefficient; according to the trimming threshold, removing the channel from the model lower than the threshold value of the γ from the model; pruning the YOLOv3 model; and taking the pruning model as the student model; using the reference model as the teacher network for knowledge distillation; teaching the soft label generated by the teacher model to guide the student model to train; and using the indication learning to accelerate the distillation speed;” here it disclose the trimming function trigger condition of the filter which is the trimming threshold, removing the channel from the model lower than the threshold value of the γ) and
building a subset ML model based on each dominant filter for each layer, wherein the subset ML model is a subset of the ML model. (“(1) training YOLOv3 model to generate reference model based on training image data set, using the backbone network Darknet-53 of YOLOv3 to extract features of the image, deep feature through the upper sampling and shallow feature tensor splicing to generate multi-scale feature map;” generate a student model by trimming the teacher model based on the filter for each layer, the student model is a subset of the teacher model)
But Liu failed to explicitly disclose “each layer includes a plurality of filters;
determining, for each layer in the ML model, a dominant filter;”
Li disclose each layer includes a plurality of filters; ([0010] [0027]-[0034] each layer has filters)
determining, for each layer in the ML model, a dominant filter; ([0010][0027]-[0034] each layer has filters and a filter is identified based on a condition)
It would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made to incorporate Li’s ML model training into Liu’s invention as they are related to the same field endeavor of model training and learning. The motivation to combine these arts, as proposed above, at least because Li’s ML model with filters in each layer would help to provide more model training control into Liu’s system. Therefore it would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made that providing filters in each layer would help to improve model training efficiency.
In regard to claim 2, Liu and Li disclose The method according to claim 1,
But Liu fail to explicitly disclose “further comprising: storing the subset ML model in a database.”
Li disclose further comprising: storing the subset ML model in a database. ([0009]-[0018] [0020]-[0024] model is stored at a storage)
It would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made to incorporate Li’s ML model training into Liu’s invention as they are related to the same field endeavor of model training and learning. The motivation to combine these arts, as proposed above, at least because Li’s ML model with data storage would help to provide more model storing into Liu’s system. Therefore it would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made that providing storage to store the model would help to improve model training efficiency.
In regard to claim 3, Liu and Li disclose The method according to claim 1,
Liu disclose wherein the ML model is a teacher model and the subset ML model is a subset teacher model. (“As a preference, in the step (5) introducing softmax function with temperature parameter and knowledge distillation algorithm, the basic model as the teacher network, pruning the model as the student network for migration learning;” a teacher model and a student model which is a subset of the teacher model)
In regard to claim 4, Liu and Li disclose The method according to claim 1,
Liu disclose wherein the ML model and the subset ML model are one of: a neural network, a convolutional neural network (CNN), and a artificial neural network (ANN). “generating weight for each characteristic channel through the gating mechanism of the circulating neural network;” a neural network)
In regard to claim 5, Liu and Li disclose The method according to claim 1,
Liu disclose wherein the working value for each filter in the layer is determined according to:
PNG
media_image1.png
82
581
media_image1.png
Greyscale
where: i is an index ranging from 1 to N, wherein N is the number of filters in the layer; a represents a set of coefficients; a; represents the coefficient for the i-th filter; fi represents the output of each filter based on training using the set of input data; and y represents the class probabilities values from the obtained set of output data. (“step D1, introducing the γ layer in the backbone network into the pruning target function for joint training, the conversion function of the BN layer is as follows:
in the formula, zin, zout is input and output of BN μB, is the average value and variance of the input, belongs to a correction parameter close to 0, preventing the sub-is 0, γ is scale factor (scale factor) and β (offset), can linearly convert the output of the BN layer into any scale, recovering the characteristic distribution of the original input, then representing the contribution value of each convolution layer for the input characteristic; measuring the importance of the corresponding convolution layer; therefore, γ as pruning parameter;
the pruning target function is adjusted as:
wherein, Ws is the weight capable of training, xs, ys represents the input and output of the training, n is a hyper-parameter, is Γ set of the coefficient γ the backbone network, f (.) is the loss function of YOLOv3, g (γ) is a penalty function of guiding sparse, wherein g (γ) = |γ |, namely L1 is regularization;” it disclose a pruning target function with a equation. The scope of the claim cannot be determined since there are undefined variables)
In regard to claim 6, Liu and Li disclose The method according to claim 1,
Liu disclose wherein the dominant filter for each layer is determined according to:
PNG
media_image2.png
74
378
media_image2.png
Greyscale
where: i is an index ranging from 1 to N, wherein N is the number of filters in the layer; a represents a set of coefficients; a; represents the coefficient for the i-th filter; fi represents the output of each filter based on training using the set of input data; y represents the class probabilities values from the obtained set of output data; r represents a regularization parameter; and ||a|| represents a regularization term. (“step D1, introducing the γ layer in the backbone network into the pruning target function for joint training, the conversion function of the BN layer is as follows:
in the formula, zin, zout is input and output of BN μB, is the average value and variance of the input, belongs to a correction parameter close to 0, preventing the sub-is 0, γ is scale factor (scale factor) and β (offset), can linearly convert the output of the BN layer into any scale, recovering the characteristic distribution of the original input, then representing the contribution value of each convolution layer for the input characteristic; measuring the importance of the corresponding convolution layer; therefore, γ as pruning parameter;
the pruning target function is adjusted as:
wherein, Ws is the weight capable of training, xs, ys represents the input and output of the training, n is a hyper-parameter, is Γ set of the coefficient γ the backbone network, f (.) is the loss function of YOLOv3, g (γ) is a penalty function of guiding sparse, wherein g (γ) = |γ |, namely L1 is regularization;” it disclose a pruning target function with a equation. The scope of the claim cannot be determined since there are undefined variables)
In regard to claim 7, Liu and Li disclose The method according to claim 3,
Liu disclose wherein the subset teacher model is used as a student ML model. (“As a preference, in the step (5) introducing softmax function with temperature parameter and knowledge distillation algorithm, the basic model as the teacher network, pruning the model as the student network for migration learning;” the student model is a subset of the teacher model)
In regard to claims 10-16, claims 10-16 are node claims corresponding to the method claims 1-7 above and, therefore, are rejected for the same reasons set forth in the rejections of claims 1-7.
In regard to claim 19, claim 19 is a node claim corresponding to the method claim 1 above and, therefore, is rejected for the same reasons set forth in the rejections of claim 1.
In regard to claim 20, claim 20 is a computer program product claim corresponding to the method claim 1 above and, therefore, is rejected for the same reasons set forth in the rejections of claim 1.
Claims 8-9, 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (Liu) CN 111626330 and Li et al. (Li) US 2019/0340510 as applied to claim 1, further in view of Doraiswami et al. (Doraiswami) US 9582618
In regard to claim 8, Liu and Li disclose The method according to claim 1,
But Liu and Li fail to explicitly disclose “further comprising: using the subset ML model to detect faults in one or more network nodes in a network.”
Doraiswami disclose further comprising: using the subset ML model to detect faults in one or more network nodes in a network. (col. 4, line 28-col. 5, line 46, using the model to identify faults in the actuator or sensors in a network)
It would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made to incorporate Doraiswami’s ML model training into Li and Liu’s invention as they are related to the same field endeavor of model training and learning. The motivation to combine these arts, as proposed above, at least because Doraiswami’s ML to detect faulty component in the system would help to provide more model applications into Li and Liu’s system. Therefore it would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made that providing more model applications in detecting the faulty component in a system would help to expand the model usage.
In regard to claim 9, Liu and Li disclose The method according to claim 1,
But Liu and Li fail to explicitly disclose “further comprising: using the subset ML model to detect faults in one or more wireless sensor devices in a network.”
Doraiswami disclose further comprising: using the subset ML model to detect faults in one or more wireless sensor devices in a network (col. 4, line 28-col. 5, line 46, col. 20, line 13-24, using the model to identify faults in the actuator or sensors in a wireless network)
It would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made to incorporate Doraiswami’s ML model training into Li and Liu’s invention as they are related to the same field endeavor of model training and learning. The motivation to combine these arts, as proposed above, at least because Doraiswami’s ML to detect faulty component in the system would help to provide more model applications into Li and Liu’s system. Therefore it would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made that providing more model applications in detecting the faulty component in a system would help to expand the model usage.
In regard to claims 17-18, claims 17-18 are node claims corresponding to the method claims 8-9 above and, therefore, are rejected for the same reasons set forth in the rejections of claims 8-9.
Claims 22 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (Liu) CN 111626330 and Li et al. (Li) US 2019/0340510 as applied to claim 1, further in view of Song et al. (Song) CN110852426A
In regard to claim 9, Liu and Li disclose The method according to claim 1,
But Liu and Li fail to explicitly disclose “wherein determining the working value for each filter in the layer comprises minimizing a difference between a weighted combination of filter outputs and the class probabilities values.”
Song disclose wherein determining the working value for each filter in the layer comprises minimizing a difference between a weighted combination of filter outputs and the class probabilities values. (“takes a model with huge computation amount as a teacher model, performs pooling operation on likelihood estimation probability values of all teacher models in a teacher model group, and generalizes estimation results of different teacher models, so that the classification probability of data is more accurate, and the comprehension capability of the data is further improved; comparing likelihood estimation probability after pooling of a teacher model group with likelihood estimation probability of a student model to obtain a difference value between the likelihood estimation probability and the teacher model, updating the student model according to the difference value to obtain the student model with the likelihood estimation probability value closest to the likelihood estimation probability value after pooling of the teacher model group, taking a feature extractor and a feature encoder of the obtained student model as a student pre-training model,” “Performing pooling operation on the likelihood estimation probability values output by the teacher model group, and outputting the pooled likelihood estimation probability values, wherein the pooling operation comprises averaging operation and weighting averaging operation; the averaging operation comprises: averaging likelihood estimation probability values corresponding to all teacher models output by the teacher model group; the weighted averaging operation comprises: weighting likelihood estimation probability values corresponding to all teacher models output by the teacher model group, then averaging, and smoothing errors caused by a single teacher model by calculating the average value or weighted average value of all the teacher model output probabilities, thereby improving the prediction accuracy; measuring a difference value between the likelihood estimation probability value of the teacher model group after pooling and the likelihood estimation probability value of the student model; updating the parameters of the student model by adopting gradient descent algorithm calculation, so that the likelihood estimation probability value of the student model iterates to the likelihood estimation probability value of the teacher model group after pooling, and finally obtaining the student model of which the likelihood estimation probability value is closest to the likelihood estimation probability value of the teacher model group after pooling; taking the feature extractor and the feature encoder of the obtained student model as a student pre-training model;” minimizing the difference between a weighted likelihood estimation probability values of the teach model group and the estimated probably values of the student model. Note: please further define the working value, and please use functional language to describe the invention since non-functional language has not much patent weight, call to discuss if necessary.)
It would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made to incorporate Song’s ML model training into Li and Liu’s invention as they are related to the same field endeavor of model training and learning. The motivation to combine these arts, as proposed above, at least because Song’s ML to determine the difference between the models would help to provide more model learning and training method into Li and Liu’s system. Therefore it would have been obvious to one having ordinary skill in the art before the effective filing data of the claimed invention was made that providing more model learning and training method would help to expand the model usage.
Response to Arguments
Applicant's arguments filed on 6/1/2026 with regard to claims 1-20 have been fully considered but they are not persuasive.
With respect to rejection of 35 USC § 101, please see the rejection above for detail.
With respect to claim 1, the applicant argues Liu fail to explicitly disclose “determining, for each layer in the ML model, by using the class labels and the class probabilities values, a working value for each filter in the layer;", The examiner respectfully disagrees. Liu disclose ((4) introducing the γ coefficient of the BN layer in the backbone network into the pruning target function for combined training; normalizing and ordering the training after-γ coefficient; according to the trimming threshold, removing the channel from the model lower than the threshold value of the γ from the model; pruning the YOLOv3 model;
(5) using the model of pruning in the step (4) as the student model; using the benchmark model as the teacher network for knowledge distillation; using the soft label generated by the teacher model to guide the student model to train; and using the indication learning to accelerate the distillation speed;
(6) inputting the image to be detected to the student model trained in the step (5) to detect the target.”
“step D, introducing the γ layer in the backbone network into the pruning target function for combined training; normalizing and ordering the training back γ coefficient; according to the trimming threshold, removing the channel with γ the threshold value according to the channel from the model; pruning the YOLOv3 model; The method specifically comprises:
step D1, introducing the γ layer in the backbone network into the pruning target function for joint training, the conversion function of the BN layer is as follows:
in the formula, zin, zout is input and output of BN μB, is the average value and variance of the input, belongs to a correction parameter close to 0, preventing the sub-is 0, γ is scale factor (scale factor) and β (offset), can linearly convert the output of the BN layer into any scale, recovering the characteristic distribution of the original input, then representing the contribution value of each convolution layer for the input characteristic; measuring the importance of the corresponding convolution layer; therefore, γ as pruning parameter;
the pruning target function is adjusted as:
wherein, Ws is the weight capable of training, xs, ys represents the input and output of the training, n is a hyper-parameter, is Γ set of the coefficient γ the backbone network, f (.) is the loss function of YOLOv3, g (γ) is a penalty function of guiding sparse, wherein g (γ) = |γ |, namely L1 is regularization;
step D2, before training, the γ coefficient presents a positive distribution; after training, the γ coefficient is approaching to 0; normalizing and ordering the training after-γ coefficient; according to the trimming threshold, removing the channel from the model lower than the threshold value of the γ from the model; cutting the branch of the channel without adding operation to the backbone network;” Liu discloses identifying trimming threshold for each layer in the ML model, by using the labels and the class probabilities values, the value is determined for each layer. The examiner would like to suggest the applicant to further define the claim limitations using the functional language since non functional language has not much patent weight, please further specify the differences between the claim limitation with the prior arts since in response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., a pre-filer working value) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). Therefore, the applicant’s argument is not persuasive.
Applicant’s argument that independent claims 10, 19, 20, with similar elements as claim 1, and thus are allowable, is not persuasive, since claim 1 has been shown to be rejected 1.
Applicant's arguments that remaining dependent claims of 2-9 of claim 1 and 11-18 of claim 10 are allowable since they directly or indirectly dependent upon one of the independent claims 1, 10 not persuasive, since the independent claim 1, 10 have been shown/explained to be rejected.
Conclusion
The prior art made of record and not relied upon is considered pertinent to Applicant's disclosure.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XUYANG XIA whose telephone number is (571)270-3045. The examiner can normally be reached Monday-Friday 8am-4pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached at 571-272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
XUYANG XIA
Primary Examiner
Art Unit 2143
/XUYANG XIA/Primary Examiner, Art Unit 2143