Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/11/2026 has been entered.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 4, 8, 9, 11, 15, 16, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Mathews (US 20210097176), in view of Zamora Esquivel (US 20210209473 A1, hereafter referred to as Zamora), Demaj (US 20200175373 A1), and Lee (US 20210303972 A1).
Regarding claim 1, Mathews discloses “A method for detecting software attack, comprising: receiving, from a first machine learning model…” (Paragraph 0026; The DL model is the first machine learning model)
“… one or more neurons of the first machine learning model” (Paragraph 0024; shows that the neural network is made up of neurons)
“… is obtained when the first machine learning model processes a production data sample to generate a prediction outcome” (Paragraph 0026; This denotes the DL model producing output based on production data sample inputs, which include adversarial or drift data samples)
“by using a plurality of layers including at least one first layer and at least one remaining layer,” (Paragraph 0027, 0055, 0056; The DL training server 102 used by the adversarial attack detector 110 uses a plurality of layers. The example model modifier 206 selects a second layer to determine if this layer and subsequent layers can be replaced, but also lets the previous layer (a first layer) remain in the model)
“wherein the production data sample is a software code and the prediction outcome indicates whether the software code has risk of malware…;” (Paragraph 0025, 0026; DL model 108 obtains the production data sample as an executable (software code) from the DL training server 102 and processes the production data sample to determine whether the executable has a risk of malware)
“using, a second machine learning model to process” (Paragraph 0019; This denotes the Adversarial attack detector performing classifications (i.e. a second machine learning model)) “to generate a distribution assessment” (Paragraph 0015; This denotes looking for concept drift, which looks for changes in a distribution).
“determining, based on the distribution assessment, whether the production data sample is an adversarial data sample or a drift data sample” (Paragraph 0026; The adversarial attack detector determines whether the sample is evidence of drift data or adversarial data.)
“and in response to detecting an attack by the production data sample, triggering an incident response” (Paragraph 0022, 0026; After the adversarial attack detector processes the production sample, which is the results of the DL model 108, triggers an incident response)
Mathews fails to explicitly disclose, “pre-activation data, wherein the pre-activation data comprises pre-activation”, “the pre-activation data”, and “wherein the pre-activation information of the one or more neurons is pre-activation information of neurons of the at least one first layer in the first machine learning model”.
Zamora discloses pre-activation data, wherein the pre-activation data comprises pre-activation”, “the pre-activation data”, and “wherein the pre-activation information of the one or more neurons is pre-activation information of neurons of the at least one first layer in the first machine learning model,” (Paragraph 0060, 0063; Zamora discloses pre-activation tensors that are associated with a layer and neurons in the layer)
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having Mathews and Zamora, to modify Mathews by specifying the use of pre-activation data for improved machine learning model training as well as specifying that neurons are present in at least one first layer of the machine learning model. One would be motivated to use pre-activation data to prepare it for use in an activation function, see e.g., paragraph 0061, where Zamora describes that pre-activation tensors are further processed to form spike train activation tensors, which are a vector of activations. One would be also motivated to use pre-activation information from the neurons from at least one first layer to use process all neurons in the layer by the same activation function, see e.g., paragraph 0061, 0063, where Zamora describes that each neuron associated with a pre-activation tensor is processed by the same activation function.
Matthews-Zamora fails to explicitly disclose, “an importance level associated with each layer of the first machine learning model and a configured threshold importance level”
“the at least one first layer is selected among the plurality of layers for pre-activation information analysis based on an importance level”
“wherein the at least one first layer is selected for pre-activation information analysis based on the importance level of the at least one first layer being above the threshold importance level, and the at least one remaining layer is not selected for pre-activation information analysis based on the importance level of the at least one remaining layer being below the threshold importance level”.
Demaj discloses “an importance level associated with each layer of the first machine learning model and a configured threshold importance level” (Paragraph 0089, 0091; importance of a layer is computed for each layer).
“the at least one first layer is selected among the plurality of layers for pre-activation information analysis based on an importance level” (Paragraph 0089, 0091; Demaj discloses selecting at least one layer from a plurality of layers that have scores (importance levels) that are above a threshold score.)
“wherein the at least one first layer is selected for pre-activation information analysis based on the importance level of the at least one first layer being above the threshold importance level, and the at least one remaining layer is not selected for pre-activation information analysis based on the importance level of the at least one remaining layer being below the threshold importance level” (Paragraph 0089, 0091; Demaj discloses selecting layers that have scores (importance levels) that are above a threshold score and ignores selecting layers that have a score below a threshold).
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having Mathews and Demaj to modify Mathews to only select layers that are above the threshold. One would be motivated to do so in order to obtain a group of layers that meet a desired threshold so the group can be used for a desired purpose, see e.g., paragraph 0089 and 0091, where Demaj teaches selecting layers if their computed score is above a threshold for the purpose of reducing memory size of layers, and ignores other layers because their importance scores indicate that memory reduction will not prove useful on those layers.
Mathews-Zamora-Demaj fails to explicitly disclose, “wherein the importance level is determined based on a value of prediction accuracy by using pre-activation output of the respective layer”.
Lee discloses “wherein the importance level is determined based on a value of prediction accuracy by using pre-activation output of the respective layer,” (Paragraph 0097; Lee discloses determining layers that have low importance according to their output data and accuracy. Neural networks are designed to output predictions, so the output of a layer is a prediction and Lee uses the accuracy of the prediction to determine importance level).
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having Mathews and Lee to modify Mathews to determine the importance level based on prediction accuracy from the respective layer’s output. One would be motivated to do so to identify a group of layers based on their importance level, see e.g., paragraph 0097, where Lee teaches that determining the importance level of each layer indicates which layers have relatively low importance.
Regarding claim 2, Mathews discloses “in response to determining that the production data sample is the adversarial data sample,” (Paragraph 0025; the executable is the production data sample and is determined to be either adversarial or drift) “determining whether an attack has been detected based on a configured policy” (Paragraph 0026; based off of the results of the adversarial attack detector, an attack has been detected if an adversarial data sample has been determined) “and storing the adversarial data sample” (Paragraph 0064; The adversarial data sample, which is an input file, is transmitted to a DL training server for storage (518 of FIG. 5)) “for a retraining of the first machine learning model” (Paragraph 0078; The example report from the results of the detector are used to retrain the first machine learning model).
Regarding claim 4, Matthews discloses “in response to determining that the production data sample is the drift data sample,” (Paragraph 0025; the executable is the production data sample and is determined to be either adversarial or drift) “determining whether to retrain the first machine learning model based on a configured policy” (Paragraph 0064; A report is transmitted to the first machine learning model and determines that the model should be retrained based on the information in the report) “and storing the drift data sample” (Paragraph 0064; The drift data sample, which is an input file, is transmitted to a DL training server for storage (518 of FIG. 5)).
Regarding claims 8 and 15, these claims are similar in scope to claim 1.
Regarding claims 9 and 16, these claims are similar in scope to claim 2.
Regarding claims 11 and 18, these claims are similar in scope to claim 4.
Claim Rejections - 35 USC § 103
Claims 7, 14, 21 are rejected under 35 U.S.C. 103 as being unpatentable over Mathews (US 20210097176), in view of Zamora (US 20210209473 A1), Demaj (US 20200175373 A1), and Lee (US 20210303972 A1), and further in view of Dawkins (Paul's Online Notes, hereafter referred to as Dawkins) and Hunter (US 20230004800 A1, hereafter referred to as Hunter).
Regarding claim 7, Mathews discloses “one or more neurons” (Paragraph 0024; This shows the neural network is made up of neurons).
Mathews fails to explicitly disclose, “the pre-activation data is a vector that includes the flattened pre-activation tensors”.
Dawkins discloses “the pre-activation data is a vector” (Section 11.1, 11.2; Vectors are useful for storing things and are resizable).
Dawkins fails to explicitly disclose, “flattened pre-activation tensors”.
However, Hunter discloses, “flattened pre-activation tensors” (Paragraph 0164; the process of combining tensors to associate them with vectors is done by flattening them into a one-dimensional column, which generates a vector).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify how Mathews manipulates neurons with pre-activation data by formatting pre-activation data as a vector to offer the advantages listed by Dawkins which include the ability to store information as well as being able to resize the vector, which is taught in “Vectors are used to represent quantities that have both a magnitude and a direction.” (Dawkins, Page 1, Paragraph 1) and “we can see that if c is positive all scalar multiplication will do is stretch (if c > 1) or shrink (if c < 1) the original vector, but it won’t change the direction” (Dawkins, Page 8, Paragraph 2). Additionally, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dawkins with the addition of flattened pre-activation tensors that is taught by Hunter, in order to compress the tensors that are from the neurons, into vectors. This transforms them into a format that is usable by the second machine learning model, and is taught as follows, “The combined dense tensors are associated with state vectors that include sparse tensor identifiers for the active weight values. The collection of the combined dense tensors and the state vectors may be denoted as augmented weight tensors (AWT). The complementary sparse tensors 810 are combined into a smaller number (L) of dense complementary sparse filter blocks (CSFBs) 820. The dense CSFBs 820 are examples of combined dense tensors. Each of these dense CSFBs is flattened into a one-dimensional column. The collection of the one-dimensional columns are concatenated horizontally into an AWT 830 that has K ports.” (Hunter, Paragraph 164). Thus; one of ordinary skill in the arts would be motivated to combine the references since it provides a scalable method of storing pre-activation data that is created from flattening tensors into a format that can be accepted by the second machine learning model.
Regarding claim 14, this claim is similar in scope to claim 7.
Response to Arguments
Applicant's arguments regarding the 35 USC 103 rejection are moot in view of the new grounds of rejection necessitated by applicant's amendments.
The rejection of Claim 1 under 35 U.S.C. 103 has been maintained. Similarly, the rejection of Claims 8 and 15 under 35 U.S.C. 103 have been maintained.
The rejection of Claims 2, 4, and 7 under 35 U.S.C. 103, which depend directly from Claim 1 have been maintained.
The rejection of Claims 9, 11, and 14 under 35 U.S.C. 103, which depend directly from Claim 8 have been maintained.
The rejection of Claims 16, 18, and 21 under 35 U.S.C. 103, which depend directly from Claim 15 have been maintained.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID KIM whose telephone number is (571)272-4331. The examiner can normally be reached 7:30 AM - 4:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Ell can be reached at (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/D.K./Examiner, Art Unit 2141
/MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141