Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Detailed Action
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submissions filed on 05/26/2026 and 6/17/2026 have been entered.
Claims 1, 18-20, and 26 have been amended
Claim 26 is added new
Claims 1-26 are pending
Priority
This application claims no priority. Therefore, the effective filing date of this application is 07/12/2022.
Response to Arguments
Applicant’s arguments filed on 06/17/2026 have been fully considered.
With respect to the USC 112(b) rejection for claims 1, 18, 19, and 20. The rejection has been overcome due to Applicant omitting the term “looser”. However, the USC 112(b) rejection for claims 20, 24, and 25 has not been overcome.
With respect to the USC 103 rejection for independent claim 1. Examiner is no longer relying on MUTHAIA H-NIEMELA-MA to teach the features of claim 1. Examiner is now using MUTHAIA H in view of SINGH to teach claims 1-3, 6-8, and 10-24. Furthermore, MUTHAIA H is no longer being relied on to teach pre-filtering or a pre-filter model.
Additional arguments are moot in view of new grounds of rejection necessitated by the claim amendments.
Claim Objections
Claims 1 and 18 recite of the limitation “based at least in part on querying a pre-filter model based at least in part on a first set of features for traffic”. Examiner suggests amending this to “based at least in part on querying a pre-filter model that is based at least in part on a first set of features for traffic”. Appropriate correction is required.
Claim 19 recites of the limitation “based at least in part querying a pre-filter model based at least in part on a first set of features”. Examiner suggests amending this to “based at least in part on querying a pre-filter model that is based at least in part on a first set of features for traffic”. Appropriate correction is required.
Claims 12 and 14 recite of the same limitation. Examiner suggests amending one of the claims to make it distinct and different. Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is:
“… the detection model is configured to” in claim 16
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
See specification para. 0105, 0118, and 0161 for functional support
See specification para. 0141 and 0104 for hardware support
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claim 7 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Claim 7 recites of “wherein the pre-filter model is a machine learning model.”. However, claim 1 already recites “wherein the pre-filter model is a machine learning model”. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claim 21 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Claim 21 recites of “wherein determining the first set of features comprises converting the second set of features to obtain the first set of features.”. However, claim 1 already recites “wherein the first set of features is derived from the detection model by converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features”. The recitation of converting the second subset of features is already recited in claim 1, and claim 21 does not recite of any further limiting features. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 3-7, 20, 24, and 25 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 3 and 4 recite the limitation “the machine learning model”. However, claim 1 recites “the pre-filter model is a machine learning model” and claim 2 recites “the detection model is a machine learning model”. It is unclear what machine learning model is being referred to in claims 3 and 4. For the purpose of examination Examiner is interpreting this limitation as “the detection model”. Appropriate correction is required.
Claim 5 depends on claim 4. Therefore, it also inherits the rejection.
Claim 6 recites the limitation “a pre-filter model”. However, claim 1 already recites of “a pre-filter model”. It is unclear if the pre-filter model of claim 1 and 6 are different or the same. For the purpose of examination Examiner is interpreting them as the same. The limitation is interpreted as “the pre-filter model”. Appropriate correction is required.
Claim 7 depends on claim 6. Therefore, it also inherits the rejection.
Claim 20 recites the limitation " pre-filtering the network traffic". There is insufficient antecedent basis for this limitation in the claim. Examiner is interpreting this limitation as “pre-filtering network traffic”. Appropriate correction is required.
Claim 20 recites the limitation " the pre-filter model". There is insufficient antecedent basis for this limitation in the claim. Examiner is interpreting this limitation as “the pre-filtering model”. Appropriate correction is required.
Claim 24 recites the limitation " the pre-filter machine-learning model". However, independent claim 1 recites of a “a pre-filter model”. Examiner is interpreting this limitation as “the pre-filter model”. Appropriate correction is required.
Claim 24 recites the limitation " the trained detection machine-learning model". There is insufficient antecedent basis for this limitation in the claim. Examiner is interpreting this limitation as “the detection model”. Furthermore, there is no basis for “re-training” the detection model. Independent claim 1 does not recite of training the detection model. For the purpose of examination examiner is interpreting “re-training” as “training”. Appropriate correction is required.
Claim 25 recites the limitation " the detection machine-learning model" and “the pre-filter machine-learning model”. There is insufficient antecedent basis for this limitation in the claim. Examiner is interpreting this limitation as “the detection model” and “the pre-filter model”. Appropriate correction is required.
Claim 26 recites the limitation “a subset of the second set of features”. However, claim 1 already recites of “a subset of the second set of features”. It is unclear if the subset recited in claim 26 is the same or different as the one recited in claim 1. For the purpose of examination, Examiner is interpreting this limitation as the same. The limitation is interpreted as “the subset of the second set of features”. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-22 and 24-26 are rejected under 35 U.S.C. 101 because they directed to an abstract idea.
Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites a judicial exception (an abstract idea) that is not integrated into a practical application.
Step 1: Statutory Category
Claim 1 satisfies the statutory category requirement because it is directed to a system, comprising: one or more processors under 35 U.S.C. 101(a).
Step 2A, Prong 1 – Judicial Exception (Abstract Area)
The claim recites of system, comprising: one or more processors configured to: obtain network traffic; pre-filter the network traffic based at least in part on querying a pre-filter model based at least in part on a first set of features for traffic reduction, wherein the pre-filter model is a machine learning model; and use a detection model in connection with determining whether a filtered network traffic comprises malicious traffic, the detection model being based at least in part on a second set of features for malware detection, wherein the first set of features is derived from the detection model by converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, and wherein the pre-filter model is obtained by training or re- training using the first set of features derived from the detection model; and a memory coupled to the one or more processors and configured to provide the one or more processors with instructions.
The limitation of obtain network traffic; pre-filter the network traffic based at least in part on querying a pre-filter model based at least in part on a first set of features for traffic reduction, wherein the pre-filter model is a machine learning model, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually obtain network traffic and pre-filter the network traffic using a pre-filtering model.
The limitation of and use a detection model in connection with determining whether a filtered network traffic comprises malicious traffic, the detection model being based at least in part on a second set of features for malware detection, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually use a detection model to determine if network traffic comprises malicious traffic.
The limitation of wherein the first set of features is derived from the detection model by converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually determine a first set of features by converting at least a subset of the second set of features into corresponding broader, more sensitive versions.
The limitation of and wherein the pre-filter model is obtained by training or re- training using the first set of features derived from the detection model; and a memory coupled to the one or more processors and configured to provide the one or more processors with instructions, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually obtain a pre-filter model by training or re- training using the first set of features derived from the detection model.
Step 2A, Prong 2 – Integration into a practical Application
This judicial exception is not integrated into a practical application. The claim recites of obtaining network traffic, pre-filtering the network traffic based on a pre-filter model, and use a detection model in connection with determining whether a filtered network traffic comprises malicious traffic. The claim further recites of converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, and wherein the pre-filter model is obtained by training or re- training using the first set of features derived from the detection model. However, merely performing conversion and determining whether a filtered network traffic comprises malicious traffic does not impose any meaningful limits on practicing the abstract idea. Furthermore, a machine learning model is a mathematical algorithm to perform data manipulation and data classification which can be done by a user. Simply making a determination does not overcome the abstract idea.
Step 2B- “Significantly More” (Inventive concept)
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. In particular, the claim only recites one additional element of “one or more processors” recited at a high-level of generality (i.e., as a generic processor implementing the system) such that it amounts no more than mere instructions to apply the exception using a generic processor. Mere instructions to apply an exception using a generic processor cannot provide an inventive concept. The claim is not patent eligible.
Claim 2 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the detection model is a machine learning model. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine the detection model is a machine learning model.
Claim 3 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the machine learning model is trained using a set of one or more feature vectors. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually implement a machine learning model train it using a set of one or more feature vectors. A machine learning model is a mathematical algorithm to perform data manipulation and data classification which can be done by a user.
Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the machine learning model is a tree-based model. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine the machine learning model is a tree-based model.
Claim 5 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the tree-based model is trained using an XGBoost machine learning process. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine the tree-based model is trained using an XGBoost machine learning process.
Claim 6 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein pre-filtering the network traffic comprises using a pre-filter model that is based at least in part on the first set of features. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually pre-filter using a pre-filter model that is based on the first set of features.
Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the pre-filter model is a machine learning model. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine the pre-filter model is a machine learning model.
Claim 8 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the one or more processors are further configured to query the detection model. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually query the detection model.
Claim 9 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the one or more processors are further configured to: in response to determining that the filtered network traffic comprises malicious traffic, update a blacklist of files that are deemed to be malicious, the blacklist of files being updated to include one or more identifiers corresponding to network traffic determined to be malicious. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually update a blacklist. Furthermore, simply updating a blacklist does not integrate into a practical application. The malicious traffic still needs to be mitigated.
Claim 10 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the one or more processors are further configured to: in response to determining that the filtered network traffic comprises malicious traffic, provide an indication that the filtered network traffic comprises malicious traffic. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually provide an indication that the filtered network traffic comprises malicious traffic. Simply providing an indication does not mitigate the malicious traffic.
Claim 11 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein determining whether the filtered network traffic comprises malicious traffic is performed at a security entity. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine whether the filtered network traffic comprises malicious traffic.
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein determining whether the filtered network traffic comprises malicious traffic is performed at a cloud-based security service. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine whether the filtered network traffic comprises malicious traffic.
Claim 13 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein pre-filtering the network traffic based at least in part on the first set of features is performed at a security entity. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually pre-filter based on a first set of features.
Claim 14 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein determining whether the filtered network traffic comprises malicious traffic is performed at a cloud-based security service. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine whether the filtered network traffic comprises malicious traffic.
Claim 15 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein: pre-filtering the network traffic comprises detecting one or more malicious or suspicious samples; and the one or more malicious or suspicious samples are forwarded to the detection model. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually pre-filter malicious or suspicious traffic and send it to detection model.
Claim 16 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the detection model is configured to determine whether a suspicious sample is malicious. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine whether a suspicious sample is malicious.
Claim 17 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the first set of features is distinct from the second set of features. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine the first set of features are distinct from the second set of features.
Claim 18 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites a judicial exception (an abstract idea) that is not integrated into a practical application.
Step 1: Statutory Category
Claim 18 satisfies the statutory category requirement because it is directed to a method under 35 U.S.C. 101(a).
Furthermore, claim 18 recites of features similar to that of claim 1. Therefore, a similar Step 2A, Prong 1, Step 2A, Prong 2, and Step 2B analysis as seen for claim 1 applies.
Claim 19 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites a judicial exception (an abstract idea) that is not integrated into a practical application.
Step 1: Statutory Category
Claim 19 satisfies the statutory category requirement because it is directed to a computer program product embodied in a non-transitory computer readable medium under 35 U.S.C. 101(a).
Furthermore, claim 19 recites of features similar to that of claim 1. Therefore, a similar Step 2A, Prong 1, Step 2A, Prong 2, and Step 2B analysis as seen for claim 1 applies.
Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites a judicial exception (an abstract idea) that is not integrated into a practical application.
Step 1: Statutory Category
Claim 20 satisfies the statutory category requirement because it is directed to a system, comprising: one or more processors under 35 U.S.C. 101(a).
Step 2A, Prong 1 – Judicial Exception (Abstract Area)
The claim recites of determine a first set of features for training a pre-filtering model to detect a malicious or suspicious sample; train the pre-filtering model based at least in part on a first training set and the first set of features; determine a second set of features for training a detection model to detect malicious samples, wherein the first set of features is derived from the detection model by converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, and wherein the pre-filter model is obtained by training or re-training using the first set of features derived from the detection model; train the detection model based at least in part on a second training set and the second set of features, wherein the first set of features used in connection with pre- filtering the network traffic is determined based at least in part on the second set of features; cause the pre-filtering model to pre-filter network traffic, wherein the pre-filter model is a machine learning model; and cause the detection model to detect malicious traffic based on a filtered network traffic output by the pre-filtering model and a memory coupled to the one or more processors and configured to provide the one or more processors with instructions.
The limitation of determine a first set of features for training a pre-filtering model to detect a malicious or suspicious sample, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually determine a first set of features for training a pre-filtering model.
The limitation of train the pre-filtering model based at least in part on a first training set and the first set of features, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually train the pre-filtering model based at least in part on a first training set and the first set of features.
The limitation of determine a second set of features for training a detection model to detect malicious samples, wherein the first set of features is derived from the detection model by converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, and wherein the pre-filter model is obtained by training or re-training using the first set of features derived from the detection model, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually determine a first set of features and train the pre-filter model based on converting at least a subset of the second set of features into corresponding broader, more sensitive versions.
The limitation of train the detection model based at least in part on a second training set and the second set of features, wherein the first set of features used in connection with pre- filtering the network traffic is determined based at least in part on the second set of features, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually train a detection model based at least in part on a second training set and the second set of features.
The limitation of cause the pre-filtering model to pre-filter network traffic, wherein the pre-filter model is a machine learning model, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually use a pre-filter model to pre-filter network traffic.
The limitation of and cause the detection model to detect malicious traffic based on a filtered network traffic output by the pre-filtering model and a memory coupled to the one or more processors and configured to provide the one or more processors with instructions, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can be performed in the mind. A user can manually cause a detection model to detect malicious traffic based on a filtered network traffic.
Step 2A, Prong 2 – Integration into a practical Application
This judicial exception is not integrated into a practical application. The claim recites of obtaining network traffic, pre-filtering the network traffic based on a pre-filter model, and use a detection model in connection with determining whether a filtered network traffic comprises malicious traffic. The claim further recites of converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, and wherein the pre-filter model is obtained by training or re- training using the first set of features derived from the detection model. However, merely performing conversion and determining whether a filtered network traffic comprises malicious traffic does not impose any meaningful limits on practicing the abstract idea. Furthermore, a machine learning model is a mathematical algorithm to perform data manipulation and data classification which can be done by a user. Simply making a determination does not overcome the abstract idea.
Step 2B- “Significantly More” (Inventive concept)
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. In particular, the claim only recites one additional element of “one or more processors” recited at a high-level of generality (i.e., as a generic processor implementing the system) such that it amounts no more than mere instructions to apply the exception using a generic processor. Mere instructions to apply an exception using a generic processor cannot provide an inventive concept. The claim is not patent eligible.
Claim 21 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein determining the first set of features comprises converting the second set of features to obtain the first set of features. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine the first set of features by converting the second set of features.
Claim 22 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein every network traffic sample classified by the detection model is necessarily classified as malicious or suspicious by the pre-filter model. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually determine every network traffic sample classified by the detection model is necessarily classified as malicious or suspicious.
Claim 24 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the pre-filter machine- learning model is obtained by re-training the trained detection machine-learning model using the first set of features. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually obtain the pre-filter model by re-training the trained detection machine-learning model.
Claim 25 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the converting comprises automatically converting one or more regular expressions used by the detection machine- learning model into corresponding broader regular expressions for use by the pre-filter machine- learning model. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually convert one or more regular expressions used by the detection machine- learning model into corresponding broader regular expressions.
Claim 26 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This claim recites of wherein the converting of at least a subset of the second set of features into the first set of features comprises automatically converting one or more regular expressions used by the detection model into corresponding broader regular expressions for use by the pre-filter model, and wherein the pre-filter model is obtained by re- training the detection model using the first set of features derived through the automatic conversion. Therefore, the limitations of this claim, as drafted, is a process that, under its broadest reasonable interpretation, covers steps that can also be performed in the mind. A user can manually convert one or more regular expressions used by the detection machine- learning model into corresponding broader regular expressions and re-train the detection model using the first set of features.
The dependent claims 2-17, 21, 22, and 24-26 are directed to abstract ideas and do not include additional elements that are sufficient to amount to significantly more than the judicial exception. This judicial exception is not integrated into a practical application. Therefore, the claims are not patent eligible.
Claim Rejections: 103 Rejections
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 6-8, and 10-24 are rejected under 35 U.S.C. 103 as being unpatentable over MUTHAIA H (US-8314718-B2), in view of SINGH (US-20190294792-A1), hereinafter MUTHAIA H-SINGH.
Regarding claim 1, MUTHAIA H teaches “A system, comprising: one or more processors configured to: obtain network traffic; ([MUTHAIA H, col 4 lines 36-39] “ASIC's as described earlier is a dedicated processor that is designed specific for an application. Hence, efficient implementations can be designed at a system level which frees up the main processor from security related processing.”) ([MUTHAIA H, Col 3 lines 60-63] “For a receiver of a respective host vehicle, the receiver has to authenticate a large number of data packets”) ([MUTHAIA H, col. 6 lines 3-9] “Characteristic data contained in the data packet that may be used in determining whether the data packet is beneficial or not beneficial to the host vehicle includes, but is not limited to, global positioning data … signatures, signal quality”) ([MUTHAIA H, col. 1 lines 20-23] “A substantial cost is involved in incorporating security protection regards to V2V applications. The cost incurred is that of the computational power required to process security”) ([MUTHAIA H, col. 1 lines 60-66] “An advantage of an embodiment of the invention provides for a reduced number of data packets that are provided to a security layer in response to filtering data packets to remove any that are from a vehicle determined not to be within the same road of travel as the host vehicle or from vehicles where malicious node tampering is has been detected.”) pre-filter the network traffic based at least in part on querying a pre-filter [model] based at least in part on a first set of features for traffic reduction, … ; ([Muthaiah, Col. 5 lines 8-22] “The filter 42 uses characteristic data either provided directly in the data packet or is derived from the data of the data packet. The characteristic data is compared to predetermined parameters as determined by the host vehicle. … Examples of data collected by the vehicle interface device that is used for determining the predetermined parameters may include, but is not limited to, GPS data, speed, velocity, acceleration, and steering angle data, which assist in determining a position or trajectory of the host vehicle relative to the remote vehicles. It should be understood that the filter 42 functions as a pre-security processing routine to filter and discard unwanted V2V communication messages. ”) ([Muthaiah, col. 6 lines 3-9] “Characteristic data contained in the data packet that may be used in determining whether the data packet is beneficial or not beneficial to the host vehicle includes, but is not limited to, global positioning data … signatures, signal quality”) … and a memory coupled to the one or more processors and configured to provide the one or more processors with instructions. ([MUTHAIA H, col 6 lines 59-62] “The client device 1 may also be provided with a computer readable medium in the form of a memory 7 on which a computer program 8 in the form of computer readable code is stored.”)
However, MUTHAIA H does not teach “… wherein the pre-filter model is a machine learning model … and use a detection model in connection with determining whether a filtered network traffic comprises malicious traffic, the detection model being based at least in part on a second set of features for malware detection, wherein the first set of features is derived from the detection model by converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, and wherein the pre-filter model is obtained by training or re- training using the first set of features derived from the detection model;”.
In analogous teaching SINGH teaches “…. wherein the pre-filter model is a machine learning model ([SINGH, para. 0013] “The foregoing problems of conventional malware systems are addressed by providing an architecture for performing filtering at a low compute cost, while simultaneously achieving high accuracy (e.g., consistent with a heavy-weight deep learning model), by employing a shallow model to filter data for a deep model.”) ([SINGH, para. 0016] “shallow model 116, can filter or detect that the instruction set may be potentially malicious or benign according to a quickly executing machine learning model based on one or more parameters in a shallow model parameter set (step 220).”) …. and use a detection model in connection with determining whether a filtered network traffic comprises malicious traffic, ([Singh, para. 0015] “System 100 illustrates the malware inference architecture where a shallow model for filtering potential malware is decoupled from a deep model that provides a final classification of the potential threat. System 100 includes cloud system network … if a given instruction set is determined to be of likely relevance to deep model 118, it is passed from shallow model 116 on endpoint 114 to a more robust deep model 118 on cloud system 112, where it can be processed more fully. Thus, while deep model 118 may take more time and more compute resources to execute, deep model 118 only analyzes what shallow model 116 classifies as relevant”) ([Singh, para. 0022] “After receiving the filtered instruction sets from endpoint 114, cloud system 112 can then analyze the instruction set using deep model 118 to verify if the instruction set comprises malicious code (step 240). “) the detection model being based at least in part on a second set of features for malware detection, ([Singh, para. 0024] “the disclosed technology involves training deep model 118 as well as shallow model 116 or in conjunction with shallow model 116. In some aspects, rather than training the shallow model 116 with only ground truth training data, training is performed on outputs of the deep model 118, which helps the shallow model 116 learn appropriate data representations. Deep model 118, for example, can be trained with a high number of parameters that can extract the best possible relevant information from data (e.g., malware, network traffic, etc.) with high accuracy. Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”) wherein the first set of features is derived from the detection model by converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, ([Singh, para. 0024] “Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”) ([Singh, para. 0025] “To mitigate the potential for accuracy loss, the training paradigm of shallow model 116 can act as a good filter, rather than as an explicit classifier. For example, the training paradigm can enable shallow model 116 to make a simplified, binary determination of whether an instruction set is potentially malicious or not”) ([Singh, para. 0017] “in some embodiments the shallow model 116 can be a set of rules that identify different groups and/or behaviors relevant to known malware types. For example, the set of rules can be related to parameters that recognize malware signatures or monitor program execution events exhibiting malware behavior, such as parameters including, but not limited to: APIs called, instructions executed, IP addresses accessed, etc.”) ([Singh, para. 0018] “ For example, one or more parameters of shallow model 116 can determine a particular instruction set to be potentially malware based on the shallow model 116 determining that the instruction set has a probability of being malicious at, or over, 65% for one or more shallow model parameters.”) ([Singh, para. 0021] “Deep model 118 can be based on a set of parameters that are different from the set of parameters used for shallow model 116. For example, in some embodiments deep model 118 can have a larger set of parameters than shallow model 116; however, in other embodiments the number of parameters may be the same or lower, but different than shallow model 116. Regardless of the number of parameters, the rules and/or parameters of deep model 118 can provide a more thorough pass on differentiating between malicious and benign instruction sets at endpoint 114 (e.g., rules and/or parameters related to APIs called, instructions executed, IP addresses accessed, etc.).”) and wherein the pre-filter model is obtained by training or re- training using the first set of features derived from the detection model; ([Singh, para. 0023] “Shallow model 116 can be further refined at endpoint 114 by modifying, based on threshold values of one or more deep models parameters (optimized for malicious instruction detection), one or more corresponding parameters in shallow model 116. For example, deep model 118 can send modified parameters to shallow model 116 for adoption on the next instruction set(s).”) ([Singh, para. 0024] “Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”)
Thus, given the teaching of Singh, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to combine the teaching of obtaining pre-filtering model by Singh into the teaching of a system determining a first set of features for training a pre-filtering model to detect a malicious or suspicious samples by MUTHAIA H. One of ordinary skill in the art would have been motivated to do so because Singh recognizes the need to efficiently detect malware ([Singh, para. 0012] “Aspects of the disclosed technology address the need for providing fast, lightweight malware detection models that can be deployed on an endpoint while not sacrificing the accuracy of the malware detection.”)
Regarding claim 2, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein the detection model is a machine learning model.” ([SINGH, para. 0021] “Those instruction sets that are determined to be potentially malicious by shallow model 116 can be sent to cloud system 112 (step 230) (e.g., sent to deep model 118) for further analysis and/or classification. Deep model 118 is deployed with high accuracy on cloud 112 to differentiate between true positive threats captured by shallow model 116 from true negatives. Deep model 118, for example, can be one or more machine learning techniques based on a deep network, which may execute more slowly than shallow model 116.”).
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 3, MUTHAIA H-SINGH teach all limitations of claim 2. SINGH further teaches “wherein the machine learning model is trained using a set of one or more feature vectors.” ([Singh, para. 0021] “Deep model 118 can be based on a set of parameters that are different from the set of parameters used for shallow model 116. For example, in some embodiments deep model 118 can have a larger set of parameters than shallow model 116; however, in other embodiments the number of parameters may be the same or lower, but different than shallow model 116. Regardless of the number of parameters, the rules and/or parameters of deep model 118 can provide a more thorough pass on differentiating between malicious and benign instruction sets at endpoint 114 (e.g., rules and/or parameters related to APIs called, instructions executed, IP addresses accessed, etc.).”).
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 6, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein pre-filtering the network traffic comprises using a pre-filter model that is based at least in part on the first set of features.” ([SINGH, para. 0016] “Services on endpoint 114, such as shallow model 116, can filter or detect that the instruction set may be potentially malicious or benign according to a quickly executing machine learning model based on one or more parameters in a shallow model parameter set (step 220). In other words, shallow model 116 can determine if a particular instruction set is relevant for further processing on deep model 118.”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 7, MUTHAIA H-SINGH teach all limitations of claim 6. SINGH further teaches “wherein the pre-filter model is a machine learning model.” ([SINGH, para. 0013] “The foregoing problems of conventional malware systems are addressed by providing an architecture for performing filtering at a low compute cost, while simultaneously achieving high accuracy (e.g., consistent with a heavy-weight deep learning model), by employing a shallow model to filter data for a deep model.”) ([SINGH, para. 0016] “shallow model 116, can filter or detect that the instruction set may be potentially malicious or benign according to a quickly executing machine learning model based on one or more parameters in a shallow model parameter set (step 220).”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 8, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein the one or more processors are further configured to query the detection model.” ([SINGH, para. 0031] “Example system 300 includes at least one processing unit (CPU or processor)”) ([SINGH, para. 0021] “Those instruction sets that are determined to be potentially malicious by shallow model 116 can be sent to cloud system 112 (step 230) (e.g., sent to deep model 118) for further analysis and/or classification. ”) ([SINGH, para. 0022] “After receiving the filtered instruction sets from endpoint 114, cloud system 112 can then analyze the instruction set using deep model 118 to verify if the instruction set comprises malicious code (step 240). While deep model 118 may execute more slowly than shallow model 116, the aggregate computational time is decreased since the instruction sets sent to deep model 118 has been filtered or otherwise reduced in size.”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 10, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein the one or more processors are further configured to: in response to determining that the filtered network traffic comprises malicious traffic, provide an indication that the filtered network traffic comprises malicious traffic.” ([SINGH, para. 0031] “For example, if the deep model 118 determines that the IP address constitutes a valid IP address after all based on its different parameter set, the method ends. But if the deep model 118 confirms that the IP address should not be accessed, the deep model 118 can flag, classify, or otherwise notify cloud system 112 that the instruction set constitutes malware.”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 11, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein determining whether the filtered network traffic comprises malicious traffic is performed at a security entity.” ([SINGH, para. 0022] “After receiving the filtered instruction sets from endpoint 114, cloud system 112 can then analyze the instruction set using deep model 118 to verify if the instruction set comprises malicious code (step 240). While deep model 118 may execute more slowly than shallow model 116, the aggregate computational time is decreased since the instruction sets sent to deep model 118 has been filtered or otherwise reduced in size.”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 12, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein determining whether the filtered network traffic comprises malicious traffic is performed at a cloud-based security service.” ([SINGH, para. 0022] “After receiving the filtered instruction sets from endpoint 114, cloud system 112 can then analyze the instruction set using deep model 118 to verify if the instruction set comprises malicious code (step 240). While deep model 118 may execute more slowly than shallow model 116, the aggregate computational time is decreased since the instruction sets sent to deep model 118 has been filtered or otherwise reduced in size.”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 13, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein pre-filtering the network traffic based at least in part on the first set of features is performed at a security entity.” ([SINGH, para. 0015] “System 100 includes cloud system network (cloud system 112) that includes one or more devices (not shown) in communication with one or more endpoints (e.g., endpoints 114, 122, 124) in a network. Each endpoint, such as endpoint 114, can be configured to execute a shallow model 116 that analyzes instruction sets to identify and classify instruction sets that could potentially constitute malware and would therefore be good candidates for analysis by a more robust malware detection model, such as deep model 118. That is, if a given instruction set is determined to be of likely relevance to deep model 118, it is passed from shallow model 116 on endpoint 114 to a more robust deep model 118 on cloud system 112”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 14, MUTHAIA H-SINGH teach all limitations of claim 13. SINGH further teaches “wherein determining whether the filtered network traffic comprises malicious traffic is performed at a cloud-based security service.” ([SINGH, para. 0022] “After receiving the filtered instruction sets from endpoint 114, cloud system 112 can then analyze the instruction set using deep model 118 to verify if the instruction set comprises malicious code (step 240). While deep model 118 may execute more slowly than shallow model 116, the aggregate computational time is decreased since the instruction sets sent to deep model 118 has been filtered or otherwise reduced in size.”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 15, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein: pre-filtering the network traffic comprises detecting one or more malicious or suspicious samples; ([SINGH, para. 0016] “Malware detection begins when an instruction set is received or detected at an endpoint, such as endpoint 114 (step 210). Services on endpoint 114, such as shallow model 116, can filter or detect that the instruction set may be potentially malicious or benign according to a quickly executing machine learning model based on one or more parameters in a shallow model parameter set (step 220).”) and the one or more malicious or suspicious samples are forwarded to the detection model. ([SINGH, para. 0021] “Those instruction sets that are determined to be potentially malicious by shallow model 116 can be sent to cloud system 112 (step 230) (e.g., sent to deep model 118) for further analysis and/or classification. “)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 16, MUTHAIA H-SINGH teach all limitations of claim 15. SINGH further teaches “wherein the detection model is configured to determine whether a suspicious sample is malicious.” ([SINGH, para. 0022] “After receiving the filtered instruction sets from endpoint 114, cloud system 112 can then analyze the instruction set using deep model 118 to verify if the instruction set comprises malicious code (step 240). While deep model 118 may execute more slowly than shallow model 116, the aggregate computational time is decreased since the instruction sets sent to deep model 118 has been filtered or otherwise reduced in size.”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 17, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein the first set of features is distinct from the second set of features.” ([SINGH, para. 0021] “Deep model 118 can be based on a set of parameters that are different from the set of parameters used for shallow model 116. For example, in some embodiments deep model 118 can have a larger set of parameters than shallow model 116”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 18, claim 18 is the method claim of claim 1 and is rejected by the same reasons and motivation by which claim 1 is rejected.
Regarding claim 19, claim 19 is the non-transitory computer readable medium claim of claim 1 and is rejected by the same reasons and motivation by which claim 1 is rejected.
Regarding claim 20, MUTHAIA H teaches “A system, comprising: one or more processors configured to: ([MUTHAIA H , col 4 lines 36-39] “ASIC's as described earlier is a dedicated processor that is designed specific for an application. Hence, efficient implementations can be designed at a system level which frees up the main processor from security related processing.”) … detect malicious traffic based on a filtered network traffic ([MUTHAIA H, Col 3 lines 60-63] “For a receiver of a respective host vehicle, the receiver has to authenticate a large number of data packets”) ([MUTHAIA H, col. 1 lines 60-66] “An advantage of an embodiment of the invention provides for a reduced number of data packets that are provided to a security layer in response to filtering data packets to remove any that are from a vehicle determined not to be within the same road of travel as the host vehicle or from vehicles where malicious node tampering is has been detected.”) ([Muthaiah, Col. 5 lines 8-22] “The filter 42 uses characteristic data either provided directly in the data packet or is derived from the data of the data packet. The characteristic data is compared to predetermined parameters as determined by the host vehicle. … Examples of data collected by the vehicle interface device that is used for determining the predetermined parameters may include, but is not limited to, GPS data, speed, velocity, acceleration, and steering angle data, which assist in determining a position or trajectory of the host vehicle relative to the remote vehicles. It should be understood that the filter 42 functions as a pre-security processing routine to filter and discard unwanted V2V communication messages. ”)
However, MUTHAIA H does not teach “… determine a first set of features for training a pre-filtering model to detect a malicious or suspicious sample; train the pre-filtering model based at least in part on a first training set and the first set of features; determine a second set of features for training a detection model to detect malicious samples, wherein the first set of features is derived from the detection model by converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, and wherein the pre-filter model is obtained by training or re-training using the first set of features derived from the detection model; train the detection model based at least in part on a second training set and the second set of features, wherein the first set of features used in connection with pre- filtering the network traffic is determined based at least in part on the second set of features; cause the pre-filtering model to pre-filter network traffic, wherein the pre-filter model is a machine learning model; and cause the detection model to detect malicious traffic based on a filtered network traffic output by the pre-filtering model and a memory coupled to the one or more processors and configured to provide the one or more processors with instructions.”
In analogous teaching SINGH teaches “… determine a first set of features for training a pre-filtering model to detect a malicious or suspicious sample; ([SINGH, para. 0017] “For example, in some embodiments the shallow model 116 can be a set of rules that identify different groups and/or behaviors relevant to known malware types. For example, the set of rules can be related to parameters that recognize malware signatures or monitor program execution events exhibiting malware behavior, such as parameters including, but not limited to: APIs called, instructions executed, IP addresses accessed, etc.”) train the pre-filtering model based at least in part on a first training set and the first set of features; ([Singh, para. 0024] “the disclosed technology involves training deep model 118 as well as shallow model 116 or in conjunction with shallow model 116. In some aspects, rather than training the shallow model 116 with only ground truth training data, training is performed on outputs of the deep model 118, which helps the shallow model 116 learn appropriate data representations. Deep model 118, for example, can be trained with a high number of parameters that can extract the best possible relevant information from data (e.g., malware, network traffic, etc.) with high accuracy. Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”) determine a second set of features for training a detection model to detect malicious samples, ([Singh, para. 0021] “Deep model 118 can be based on a set of parameters that are different from the set of parameters used for shallow model 116. For example, in some embodiments deep model 118 can have a larger set of parameters than shallow model 116; however, in other embodiments the number of parameters may be the same or lower, but different than shallow model 116. Regardless of the number of parameters, the rules and/or parameters of deep model 118 can provide a more thorough pass on differentiating between malicious and benign instruction sets at endpoint 114 (e.g., rules and/or parameters related to APIs called, instructions executed, IP addresses accessed, etc.). For example, deep model 118 can be configured to not only verify that an instruction set is malware, but also classify a type of security risk associated with the instruction set based on a parameter set that is different from the shallow model's 116 parameter set. Specifically, deep model 118 can include a greater number of rules and/or parameters than shallow model 116.”) wherein the first set of features is derived from the detection model by converting at least a subset of the second set of features into corresponding broader, more sensitive versions of the subset of the second set of features, ([Singh, para. 0024] “Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”) ([Singh, para. 0025] “To mitigate the potential for accuracy loss, the training paradigm of shallow model 116 can act as a good filter, rather than as an explicit classifier. For example, the training paradigm can enable shallow model 116 to make a simplified, binary determination of whether an instruction set is potentially malicious or not”) ([Singh, para. 0017] “in some embodiments the shallow model 116 can be a set of rules that identify different groups and/or behaviors relevant to known malware types. For example, the set of rules can be related to parameters that recognize malware signatures or monitor program execution events exhibiting malware behavior, such as parameters including, but not limited to: APIs called, instructions executed, IP addresses accessed, etc.”) ([Singh, para. 0021] “Deep model 118 can be based on a set of parameters that are different from the set of parameters used for shallow model 116. For example, in some embodiments deep model 118 can have a larger set of parameters than shallow model 116; however, in other embodiments the number of parameters may be the same or lower, but different than shallow model 116.”) ([Singh, para. 0018] “ For example, one or more parameters of shallow model 116 can determine a particular instruction set to be potentially malware based on the shallow model 116 determining that the instruction set has a probability of being malicious at, or over, 65% for one or more shallow model parameters.”) and wherein the pre-filter model is obtained by training or re-training using the first set of features derived from the detection model; ([Singh, para. 0023] “Shallow model 116 can be further refined at endpoint 114 by modifying, based on threshold values of one or more deep models parameters (optimized for malicious instruction detection), one or more corresponding parameters in shallow model 116. For example, deep model 118 can send modified parameters to shallow model 116 for adoption on the next instruction set(s).”) ([Singh, para. 0024] “Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”) train the detection model based at least in part on a second training set and the second set of features, wherein the first set of features used in connection with pre- filtering the network traffic is determined based at least in part on the second set of features; ([Singh, para. 0024] “ Deep model 118, for example, can be trained with a high number of parameters that can extract the best possible relevant information from data (e.g., malware, network traffic, etc.) with high accuracy. Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”) cause the pre-filtering model to pre-filter [network traffic], wherein the pre-filter model is a machine learning model; ([SINGH, para. 0013] “The foregoing problems of conventional malware systems are addressed by providing an architecture for performing filtering at a low compute cost, while simultaneously achieving high accuracy (e.g., consistent with a heavy-weight deep learning model), by employing a shallow model to filter data for a deep model.”) ([SINGH, para. 0016] “shallow model 116, can filter or detect that the instruction set may be potentially malicious or benign according to a quickly executing machine learning model based on one or more parameters in a shallow model parameter set (step 220).”) and cause the detection model to detect malicious [traffic] based on a filtered [network traffic] output by the pre-filtering model and a memory coupled to the one or more processors and configured to provide the one or more processors with instructions. ([SINGH, para. 0021] “Those instruction sets that are determined to be potentially malicious by shallow model 116 can be sent to cloud system 112 (step 230) (e.g., sent to deep model 118) for further analysis and/or classification.”) ([SINGH, para. 0022] “After receiving the filtered instruction sets from endpoint 114, cloud system 112 can then analyze the instruction set using deep model 118 to verify if the instruction set comprises malicious code (step 240). While deep model 118 may execute more slowly than shallow model 116, the aggregate computational time is decreased since the instruction sets sent to deep model 118 has been filtered or otherwise reduced in size.”) ([SINGH, para. 0030] “Example system 300 includes at least one processing unit (CPU or processor) 310 and connection 305 that couples various system components including system memory 315, such as read only memory (ROM) and random access memory (RAM) to processor 310.”)
SINGH does not teach of network traffic. However, MUTHAIA H does, the same rejection applies.
Thus, given the teaching of Singh, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to combine the teaching of obtaining pre-filtering model by Singh into the teaching of a system determining a first set of features for training a pre-filtering model to detect a malicious or suspicious samples by MUTHAIA H. One of ordinary skill in the art would have been motivated to do so because Singh recognizes the need to efficiently detect malware ([Singh, para. 0012] “Aspects of the disclosed technology address the need for providing fast, lightweight malware detection models that can be deployed on an endpoint while not sacrificing the accuracy of the malware detection.”)
Regarding claim 21, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein determining the first set of features comprises converting the second set of features to obtain the first set of features.” ([Singh, para. 0024] “Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”) ([Singh, para. 0017] “in some embodiments the shallow model 116 can be a set of rules that identify different groups and/or behaviors relevant to known malware types. For example, the set of rules can be related to parameters that recognize malware signatures or monitor program execution events exhibiting malware behavior, such as parameters including, but not limited to: APIs called, instructions executed, IP addresses accessed, etc.”) ([Singh, para. 0018] “ For example, one or more parameters of shallow model 116 can determine a particular instruction set to be potentially malware based on the shallow model 116 determining that the instruction set has a probability of being malicious at, or over, 65% for one or more shallow model parameters.”) ([Singh, para. 0021] “Deep model 118 can be based on a set of parameters that are different from the set of parameters used for shallow model 116. For example, in some embodiments deep model 118 can have a larger set of parameters than shallow model 116; however, in other embodiments the number of parameters may be the same or lower, but different than shallow model 116. Regardless of the number of parameters, the rules and/or parameters of deep model 118 can provide a more thorough pass on differentiating between malicious and benign instruction sets at endpoint 114 (e.g., rules and/or parameters related to APIs called, instructions executed, IP addresses accessed, etc.).”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 22, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein every network traffic sample classified by the detection model is necessarily classified as malicious or suspicious by the pre-filter model” ([SINGH, 0016] “Services on endpoint 114, such as shallow model 116, can filter or detect that the instruction set may be potentially malicious or benign according to a quickly executing machine learning model based on one or more parameters in a shallow model parameter set (step 220). In other words, shallow model 116 can determine if a particular instruction set is relevant for further processing on deep model 118.”) ([SINGH, 0021] “Those instruction sets that are determined to be potentially malicious by shallow model 116 can be sent to cloud system 112 (step 230) (e.g., sent to deep model 118) for further analysis and/or classification.”) ([SINGH, 0022] “After receiving the filtered instruction sets from endpoint 114, cloud system 112 can then analyze the instruction set using deep model 118 to verify if the instruction set comprises malicious code (step 240).”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 23, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein the pre-filter model executes on a security entity disposed on a network edge and the detection model executes in a cloud-based analysis service remote from the security entity.” ([SINGH, 0015, Fig. 1] “System 100 includes cloud system network (cloud system 112) that includes one or more devices (not shown) in communication with one or more endpoints (e.g., endpoints 114, 122, 124) in a network. Each endpoint, such as endpoint 114, can be configured to execute a shallow model 116 that analyzes instruction sets to identify and classify instruction sets that could potentially constitute malware and would therefore be good candidates for analysis by a more robust malware detection model, such as deep model 118. That is, if a given instruction set is determined to be of likely relevance to deep model 118, it is passed from shallow model 116 on endpoint 114 to a more robust deep model 118 on cloud system 112, where it can be processed more fully. “)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Regarding claim 24, MUTHAIA H-SINGH teach all limitations of claim 1. SINGH further teaches “wherein the pre-filter machine-learning model is obtained by re-training the trained detection machine-learning model using the first set of features”. ([Singh, para. 0024] “the disclosed technology involves training deep model 118 as well as shallow model 116 or in conjunction with shallow model 116. In some aspects, rather than training the shallow model 116 with only ground truth training data, training is performed on outputs of the deep model 118, which helps the shallow model 116 learn appropriate data representations. Deep model 118, for example, can be trained with a high number of parameters that can extract the best possible relevant information from data (e.g., malware, network traffic, etc.) with high accuracy. Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”) ([Singh, para. 0025] “To mitigate the potential for accuracy loss, the training paradigm of shallow model 116 can act as a good filter, rather than as an explicit classifier. For example, the training paradigm can enable shallow model 116 to make a simplified, binary determination of whether an instruction set is potentially malicious or not (such as whether the instruction set would be relevant to the slower, but more accurate, deep model 118)”) ([Singh, para. 0023] “Shallow model 116 can be further refined at endpoint 114 by modifying, based on threshold values of one or more deep models parameters (optimized for malicious instruction detection), one or more corresponding parameters in shallow model 116. For example, deep model 118 can send modified parameters to shallow model 116 for adoption on the next instruction set(s).”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over MUTHAIA H-SINGH in view of Dennison (US-9043894-B1), hereinafter MUTHAIA H-SINGH-DENNSION
Regarding claim 4, MUTHAIA H-SINGH teaches all limitations of claim 3. However, Neither MUTHAIA H nor SINGH discloses, but Dennison teaches “wherein the machine learning model is a tree-based model ([Dennison, col 5 lines 16-20] “The score can be based on a Support Vector Machine model, a Neural Network model, a Decision Tree model, a Naive Bayes model, or a Logistic Regression model.).”
Thus, given the teaching of Dennison, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to combine the teaching of tree-based model by Dennison into the teaching of a system determining a first set of features for training a pre-filtering model to detect a malicious or suspicious samples by MUTHAIA H-SINGH. One of ordinary skill in the art would have been motivated to do so because Dennison recognizes the need to improve networks ([DENNISON, col 1 lines 41-51] “At least some of the systems, methods, and media can analyze data, such as URL data items, transmitted by computing systems within a local network in order to identify the infected systems and/or systems that have or are likely to access undesirable online resources, thereby improving functioning of the local network. The disclosed systems, methods, and media also improve functioning of at least one computing system by reducing the data to be analyzed to those data items most likely associated with malicious software, significantly improving processing speed when determining potentially malicious addresses.”)
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over MUTHAIA H-SINGH-DENNISON in view of Chen et al. XGBoost: A Scalable Tree Boosting System, August 13-17, 2016, ACM, 785-794. (henceforth Chen).
As for claim 5, MUTHAIA H-SINGH-DENNISON teach all limitations of claim 4. However, neither Muthaiah nor SINGH nor Dennison discloses, but Chen teaches “wherein the tree-based model is trained using an XGBoost machine learning process ([CHEN, Introduction p 785 ¶ 3], “In this paper, we describe XGBoost, a scalable machine learning system for tree boosting.”).”
Thus, given the teaching of Chen, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to combine the teaching of XGBoost by Chen into the teaching of a system determining a first set of features for training a pre-filtering model to detect a malicious or suspicious samples by MUTHAIA H-SINGH-DENNISON. One of ordinary skill in the art would have been motivated to do so because Chen recognizes the benefits of XGBoost ([CHEN, 6.6 Distributed Experiment] “The baseline systems are only able to handle subset of the data with the given resources. This experiment shows the advantage to bring all the system improvement together and solve a real-world scale problem.”)
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Muthaiah-SINGH and further in view of Xaypanya et al. 9332028 B2, 2013 (hereinafter Xaypanya).
Regarding claim 9, MUTHAIA H-SINGH teach all limitations of claim 1. However, Neither Muthaiah nor SINGH discloses, but Xaypanya teaches “wherein the one or more processors are further configured to: in response to determining that the filtered network traffic comprises malicious traffic, update a blacklist of files ([Xaypanya, col 5 line 39, “updating blacklists”) that are deemed to be malicious (FIG. 3 element 308, “malware,” malware is malicious; vide infra Error! Reference source not found.), the blacklist of files being updated to include one or more identifiers (FIG. 3 element 308, “signatures,” vide infra Error! Reference source not found.) corresponding to network traffic determined to be malicious.”
Thus, given the teaching of Xaypanya, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to combine the teaching of update a blacklist by Xaypanya into the teaching of a system determining a first set of features for training a pre-filtering model to detect a malicious or suspicious samples by MUTHAIA H-SINGH. One of ordinary skill in the art would have been motivated to do so because Xaypanya recognizes the need to protect a computer system ([Xaypanya, col. 1 Lines 41-48] “present disclosure to provide a proactive security mechanism that can be employed to proactively protect a computing system or network. The disclosed proactive security mechanism and the hardware that implements it can actively identify new attacks and forms of malware in a matter of milliseconds and within seconds thereafter produce an effective countermeasure or antidote to combat and/or neutralize the malware.”)
Claims 25 and 26 are rejected under 35 U.S.C. 103 as being unpatentable over MUTHAIA H-SINGH in view of Schmid (US-20180012139-A1)
Regarding claim 25, MUTHAIA H-SINGH teaches all limitations of claim 1. Furthermore, MUTHAIA H-SINGH teaches of “detection machine-learning model” and “pre-filter machine-learning model” as can be seen in the rejection of claim 1. However, MUTHAIA H-SINGH does not teach “wherein the converting comprises automatically converting one or more regular expressions used by the … machine-learning model into corresponding broader regular expressions for use by the … machine-learning model.”
In analogous teaching Schmid teaches “wherein the converting comprises automatically converting one or more regular expressions used by the … machine-learning model into corresponding broader regular expressions for use by the … machine-learning model.” ([Schmid, para. 0034] “The pattern search module 104 can allow regular expressions to be adjusted appropriately in order to obtain desired results. Often there can be a tradeoff between accuracy and coverage for a regular expression. If a regular expression associated with an intent is broader, it can cover or have many matching messages, but may also identify messages that are not highly related to the intent. On the other hand, if a regular expression is narrower, it may identify messages that are highly related to the intent, but may not include all messages that may match the intent. Regular expressions can be edited or modified to achieve the desired balance between accuracy and coverage. For example, if a particular intent classification associated with a message does not seem to accurately reflect the user intent, regular expressions may be changed to achieve higher precision and reduce noise. In some embodiments, adjustments to regular expressions performed by the pattern search module 104 can be based on manual or machine learning techniques.”).
Thus, given the teaching of Schmid, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to combine the teaching of regular expressions by Schmid into the teaching of a system determine a first set of features for training a pre-filtering model to detect a malicious or suspicious sample by MUTHAIA H-SINGH . One of ordinary skill in the art would have been motivated to do so because Schmid recognizes the need to better analyze messages ([Schmid, para. 0025] “An improved approach rooted in computer technology can overcome the foregoing and other disadvantages associated with conventional approaches specifically arising in the realm of computer technology. Based on computer technology, the disclosed technology can classify messages based on potential user intent associated with the messages”) ([Schmid, para. 0034] “regular expressions may be changed to achieve higher precision and reduce noise”)
Regarding claim 26, MUTHAIA H-SINGH teaches all limitations of claim 1. Furthermore, this claim recites of features recited in claims 1 and 25. Therefore, the same rejection and motivation as seen in the rejection of claims 1 and 25 applies. SINGH further teaches “wherein the pre-filter model is obtained by re- training the detection model using the first set of features derived through the automatic conversion.” ([SINGH, para. 0024] “disclosed technology involves training deep model 118 as well as shallow model 116 or in conjunction with shallow model 116. In some aspects, rather than training the shallow model 116 with only ground truth training data, training is performed on outputs of the deep model 118, which helps the shallow model 116 learn appropriate data representations. Deep model 118, for example, can be trained with a high number of parameters that can extract the best possible relevant information from data (e.g., malware, network traffic, etc.) with high accuracy. Once deep model 118 is trained and tested for correctness, the shallow model 116 is trained and deployed with one or more shallow model parameters set by the trained deep model 118. To keep the size of the shallow model 116 minimal, a fixed parameter budget of the shallow model 116 can be set.”) [SINGH, para. 0023] “Shallow model 116 can be further refined at endpoint 114 by modifying, based on threshold values of one or more deep models parameters (optimized for malicious instruction detection), one or more corresponding parameters in shallow model 116. For example, deep model 118 can send modified parameters to shallow model 116 for adoption on the next instruction set(s).”)
The same motivation to modify MUTHAIA H with SINGH as in the rejection of claim 1 applies.
Pertinent Art
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
JIANG (US-20220368703-A1): This prior art teaches of method for detecting security based on machine learning in combination with rule matching is provided, including: establishing a machine learning model; training the machine learning model by using a labeled legal traffic and a labeled malicious traffic; collecting a network traffic; preprocessing the collected network traffic; detecting a malicious traffic from the preprocessed network traffic by using a rule-matching-based method; identifying a malicious traffic from the preprocessed network traffic by using the trained machine learning model, including: extracting a feature of the preprocessed network traffic, and identifying the malicious traffic based on the extracted feature by using the trained machine learning model; and integrating the malicious traffic detected by the rule-matching-based method and the malicious traffic identified by the trained machine learning model.
ANDERSSON (US-11128664-B1): This prior art teaches of an intrusion prevention system includes a machine learning model for inspecting network traffic. The intrusion prevention system receives and scans the network traffic for data that match an anchor pattern. A data stream that follows the data that match the anchor pattern is extracted from the network traffic. Model features of the machine learning model are identified in the data stream. The intrusion prevention system classifies the network traffic based at least on model coefficients of the machine learning model that are identified in the data stream. The intrusion prevention system apples a network policy on the network traffic (e.g., block the network traffic) when the network traffic is classified as malicious.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AFAQ ALI whose telephone number is (571)272-1571. The examiner can normally be reached Mon - Fri 7:30am - 5:30pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ALI SHAYANFAR can be reached at (571) 270-1050. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.A./
08/20/2026
/AFAQ ALI/Examiner, Art Unit 2434
/NOURA ZOUBAIR/Primary Examiner, Art Unit 2434