DETAILED ACTION
This communication is in response to Application No. 18/720,203 filed on June 14th, 2024 in which claims 1-4, 6-8, 10-11, and 13-15 are presented for examination.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant's claim for foreign priority based on an application filed in the French Republic on 12/15/2021 and a Patent Cooperation Treaty filing on 12/08/2022. Acknowledgment is also made of receipt of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file.
Information Disclosure Statement
The information disclosure statements submitted on 06/14/2024 and 09/13/2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements were considered by the examiner.
Specification
The contents of the specification are sufficient for examination purposes.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claim 3 is rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Specifically, Claim 3 recites the term “previous” (ln. 5), which is a relative term that renders the claim indefinite. The term “previous” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree of the term, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Specifically, the term “previous” lies on a continuum with the term ongoing or one of its synonyms, such that it is not clear what should be considered the delineation point between these terms. As a result, a person of ordinary skill in the art would not be reasonably appraised on what qualifies as “a previous quality value” (ln. 5). Therefore, the claim is rejected. The claim should be amended to provide a standard for ascertaining the requisite degree for the term “previous”.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-4, 6-8, 10-11 and 13-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract ideas without significantly more.
Regarding Claim 1:
Step 1: Claim 1 is a process claim. Therefore, claims 1-4 are directed to a statutory category of eligible subject matter.
Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. Whereas if a claim limitation, under its broadest reasonable interpretation, covers mathematical relationships, mathematical formulas or equations, or mathematical calculations, then it falls within the “Mathematical Concepts” grouping of abstract ideas. Here, steps of the claimed subject matter are mental processes and mathematical concepts. Specifically, the claim recites
“comprising: selecting an action aimed at modifying the content of a malware . . . wherein selecting an action . . . are iterated as long as a stopping criterion is not reached” (mental process – amounts to exercising judgement to form an opinion on an action, with reference to known or observed information, which can occur iteratively with reference to new information, and which may be aided by pen and paper);
“the reward being defined as:
PNG
media_image1.png
42
220
media_image1.png
Greyscale
with R a positive real value, T a threshold value . . . p(t + 1) a detectability score . . . and p(t) the detectability score . . . in a previous iteration” (mathematical concept – amounts to mathematical calculations performed based on a mathematical formula; mental process – alternatively, amounts to exercising judgment to generate a value, with reference to known or observed information, which may be aided by pen and paper); and
“determining, . . . based on the obtained rewards, a function which associates with each state at least one action to be executed, so as to maximize a sum of the obtained rewards” (mental process – amounts to exercising judgement to form an opinion on a function, with reference to known or observed information, which may be aided by pen and paper).
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
The claim recites the additional elements:
“A training method for training . . . to improve performance of anti-malware software . . . representative of a probability that the malware modified by application of the selected action is considered benign by the anti-malware software . . . specific to the anti-malware software . . . representative of a probability that the malware modified by application of the selected action is considered malicious by the anti-malware software . . . by application of an action selected . . . representative of the malware modified by application of the selected action” (amounts to merely reciting a particular technological environment or field of use, which does not impose any meaningful limits on practicing the abstract idea);
“an autonomous agent implementing a reinforcement learning algorithm . . . the method being implemented by the autonomous agent . . . implementing said anti- malware software . . . by the reinforcement learning algorithm” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea); and
“transmitting the selected action to an environment . . . receiving, from the environment, a reward . . . obtaining a state . . . receiving a reward and obtaining a state” (amounts to insignificant extra-solution because receiving and providing data amounts to the transmission of data, which is incidental to the claimed subject matter).
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
The claim recites the additional element:
“A training method for training . . . to improve performance of anti-malware software . . . representative of a probability that the malware modified by application of the selected action is considered benign by the anti-malware software . . . specific to the anti-malware software . . . representative of a probability that the malware modified by application of the selected action is considered malicious by the anti-malware software . . . by application of an action selected . . . representative of the malware modified by application of the selected action” (merely reciting a particular technological environment or field of use does not provide an inventive concept);
“an autonomous agent implementing a reinforcement learning algorithm . . . the method being implemented by the autonomous agent . . . implementing said anti- malware software . . . by the reinforcement learning algorithm” (mere instructions to apply the exception using generic computer components cannot provide an inventive concept); and
“transmitting the selected action to an environment . . . receiving, from the environment, a reward . . . obtaining a state . . . receiving a reward and obtaining a state” (transmission of data, such as through a network, see buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014), or by accessing information in memory, see Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93, is well‐understood, routine, and conventional; which is recited here with a high level of generality, and remains insignificant extra-solution activity even upon reconsideration).
For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2-4. The additional limitations of the dependent claims are addressed below.
Regarding Claim 2:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. Here, the claim recites additional elements that are mental processes. Specifically, the claim recites:
“determining the state based on the selected action and on the obtained malware” (mental process – amounts to exercising judgement to form an opinion on a state, with reference to known or observed information, which may be aided by pen and paper).
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
The claim recites the additional elements:
“wherein obtaining a state comprises either receiving the state from the environment; or obtaining the malware” (amounts to insignificant extra-solution because receiving and providing data amounts to the transmission of data, which is incidental to the claimed subject matter) and
“on which an action can be applied” (amounts to merely reciting a particular technological environment or field of use, which does not impose any meaningful limits on practicing the abstract idea).
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
The claim recites the additional element:
“wherein obtaining a state comprises either receiving the state from the environment; or obtaining the malware” (transmission of data, such as through a network, see buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014), or by accessing information in memory, see Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93, is well‐understood, routine, and conventional; which is recited here with a high level of generality, and remains insignificant extra-solution activity even upon reconsideration) and
“on which an action can be applied” (merely reciting a particular technological environment or field of use does not provide an inventive concept).
Accordingly, Claim 2 is rejected as being directed to an abstract idea without significantly more.
Regarding Claim 3:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 3 depends on. Here, the claim recites additional elements that are mental processes and mathematical concepts. Specifically, the claim recites:
“the determination of the function comprises a determination, for each state-action pair (s, a), of a value QN(s, a) such that:
PNG
media_image2.png
32
402
media_image2.png
Greyscale
with α ⊂ [0,1] a learning rate, Q(s, a) a previous quality value, r(t + 1) a reward, y ⊂ [0,1] a refresh rate, s(t + 1) a next state and a(t + 1) an action that can be executed from the state s(t + 1)” (mathematical concept – amounts to mathematical calculations performed based on a mathematical formula; mental process – alternatively, amounts to exercising judgment to generate a value, with reference to known or observed information, which may be aided by pen and paper).
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
The claim recites the additional elements:
“wherein the reinforcement learning algorithm is a "Q-learning" algorithm . . . so as to determine an optimal Q-function” (amounts to merely reciting a particular technological environment or field of use, which does not impose any meaningful limits on practicing the abstract idea).
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
The claim recites the additional element:
“wherein the reinforcement learning algorithm is a "Q-learning" algorithm . . . so as to determine an optimal Q-function” (merely reciting a particular technological environment or field of use does not provide an inventive concept).
Accordingly, Claim 3 is rejected as being directed to an abstract idea without significantly more.
Regarding Claim 4:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on. Here, the claim recites additional elements that are mental processes. Specifically, the claim recites:
“wherein the action is selected from a set of actions consisting of: modifying a value of a field of a header of the malware; adding to the content of the malware a sequence of characters extracted from a benign file; adding to the content of the malware determined characters or instructions; adding to the content of the malware a library extracted from a benign file; renaming a section of the content of the malware; removing a debugger mode from the content of the malware; modifying a timestamp of the content of the malware; modifying a hash value calculated for an optional header of the content of the malware; and, decompressing an executable version of the malware” (mental process – amounts to exercising judgment to form an opinion on which action, from a set of known or observed actions, to select, which may be aided by pen and paper).
Step 2A Prong 2 & Step 2B: There are no elements left for consideration of implementation within a practical application or for consideration of significantly more.
Accordingly, Claim 4 is rejected as being directed to an abstract idea without significantly more.
Regarding Claim 6, the claim recites limitations that are all substantially the same as limitations of Claim 1, in the form of a non-transitory computer-readable medium. For substantially the same reasoning, the claim is also directed to performing mental processes and mathematical concepts without integration into a practical application or amounting to either significantly more or an inventive concept.
Accordingly, Claim 6 is rejected under the same rationale.
Regarding Claim 7:
Step 1: Claim 7 is a process claim. Therefore, claims 7-8, 11, and 13 are directed to a statutory category of eligible subject matter.
Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. Whereas if a claim limitation, under its broadest reasonable interpretation, covers mathematical relationships, mathematical formulas or equations, or mathematical calculations, then it falls within the “Mathematical Concepts” grouping of abstract ideas. Here, steps of the claimed subject matter are mental processes and mathematical concepts. Specifically, the claim recites
“An evaluation method for evaluating detectability of a malware . . . the method comprising . . . modifying the content of the malware by application of said action, so as to obtain a modified malware ” (mental process – apart from the “modifying” itself, which may require apply the judicial exception on generic and unspecialized computer components, amounts to exercising judgement to form an opinion on how known or observed information will be modified by a known or observed action, which may be aided by pen and paper);
“analyzing . . . the modified malware” (mental process – amounts to exercising judgement to evaluate known or observed information, which may be aided by pen and paper); and
“the reward being defined as:
PNG
media_image1.png
42
220
media_image1.png
Greyscale
with R a positive real value, T a threshold value . . . p(t + 1) a detectability score . . . and p(t) the detectability score . . . in a previous iteration” (mathematical concept – amounts to mathematical calculations performed based on a mathematical formula; mental process – alternatively, amounts to exercising judgment to generate a value, with reference to known or observed information, which may be aided by pen and paper).
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
The claim recites the additional elements:
“implementing at least one anti-malware software . . . aimed at modifying content of the malware . . . representative of a probability that the malware modified by application of the selected action (a(t)) is considered benign by the anti-malware software . . . specific to the anti-malware software . . . representative of a probability that the malware modified by application of the selected action is considered malicious by the anti-malware software . . . by application of an action selected” (amounts to merely reciting a particular technological environment or field of use, which does not impose any meaningful limits on practicing the abstract idea);
“by an environment . . . implementing a reinforcement learning algorithm . . . modifying . . . by the anti-malware software” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea); and
“receiving, from an autonomous agent . . . an action . . . and transmitting, to the autonomous agent, a reward” (amounts to insignificant extra-solution because receiving and providing data amounts to the transmission of data, which is incidental to the claimed subject matter).
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
The claim recites the additional element:
“implementing at least one anti-malware software . . . aimed at modifying content of the malware . . . representative of a probability that the malware modified by application of the selected action (a(t)) is considered benign by the anti-malware software . . . specific to the anti-malware software . . . representative of a probability that the malware modified by application of the selected action is considered malicious by the anti-malware software . . . by application of an action selected” (merely reciting a particular technological environment or field of use does not provide an inventive concept);
“by an environment . . . implementing a reinforcement learning algorithm . . . modifying . . . by the anti-malware software” (mere instructions to apply the exception using generic computer components cannot provide an inventive concept); and
“receiving, from an autonomous agent . . . an action . . . and transmitting, to the autonomous agent, a reward” (transmission of data, such as through a network, see buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014), or by accessing information in memory, see Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93, is well‐understood, routine, and conventional; which is recited here with a high level of generality, and remains insignificant extra-solution activity even upon reconsideration).
For the reasons above, Claim 7 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 8, 11, and 13. The additional limitations of the dependent claims are addressed below.
Regarding Claim 8:
Step 2A Prong 1: See the rejection of Claim 7 above, which Claim 8 depends on. Here, the claim recites additional elements that are mental processes. Specifically, the claim recites:
“generating an association between the action, and either the score p(t + 1), or the reward r(t + 1)” (mental process – amounts to exercising judgment to form an opinion on an association, with reference to known or observed information, which may be aided by pen and paper).
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
The claim recites the additional elements:
“in an association table” (amounts to merely reciting a particular technological environment or field of use, which does not impose any meaningful limits on practicing the abstract idea).
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
The claim recites the additional element:
“in an association table” (merely reciting a particular technological environment or field of use does not provide an inventive concept).
Accordingly, Claim 8 is rejected as being directed to an abstract idea without significantly more.
Regarding Claim 10, the claim recites limitations that are all substantially the same as limitations of Claim 7, in the form of a non-transitory computer-readable medium. For substantially the same reasoning, the claim is also directed to performing mental processes and mathematical concepts without integration into a practical application or amounting to either significantly more or an inventive concept.
Accordingly, Claim 10 is rejected under the same rationale.
Regarding Claim 11:
Step 2A Prong 1: See the rejection of Claim 7 above, which Claim 11 depends on. Here, the claim recites additional elements that are mental processes. Specifically, the claim recites:
“the method comprising . . . labeling said malwares as malicious” (mental process – amounts to exercising judgement to form an opinion on data, with reference to known or observed information, which may be aided by pen and paper).
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
The claim recites the additional elements:
“obtaining a plurality of modified malwares” (amounts to insignificant extra-solution because receiving and providing data amounts to the transmission of data, which is incidental to the claimed subject matter);
“A method for training anti-malware software implementing a learning algorithm . . . in accordance with a method for evaluating detectability of a malware by an environment implementing at least one anti- malware software according to claim 7, each malware of the plurality having a detectability score (p(t + 1)) representative of a probability that the modified malware is considered malicious by the anti-malware software, the score of each malware from the plurality being less than a defined value . . . with the labeled malwares” (amounts to merely reciting a particular technological environment or field of use, which does not impose any meaningful limits on practicing the abstract idea); and
“training the anti-malware software” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea).
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
The claim recites the additional element:
“obtaining a plurality of modified malwares” (transmission of data, such as through a network, see buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014), or by accessing information in memory, see Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93, is well‐understood, routine, and conventional; which is recited here with a high level of generality, and remains insignificant extra-solution activity even upon reconsideration) and
“A method for training anti-malware software implementing a learning algorithm . . . in accordance with a method for evaluating detectability of a malware by an environment implementing at least one anti- malware software according to claim 7, each malware of the plurality having a detectability score (p(t + 1)) representative of a probability that the modified malware is considered malicious by the anti-malware software, the score of each malware from the plurality being less than a defined value . . . with the labeled malwares” (merely reciting a particular technological environment or field of use does not provide an inventive concept).
“training the anti-malware software” (mere instructions to apply the exception using generic computer components cannot provide an inventive concept).
Accordingly, Claim 11 is rejected as being directed to an abstract idea without significantly more.
Regarding Claim 13, the claim recites limitations that are all substantially the same as limitations of Claim 11, in the form of a non-transitory computer-readable medium. For substantially the same reasoning, the claim is also directed to performing mental processes and mathematical concepts without integration into a practical application or amounting to either significantly more or an inventive concept.
Accordingly, Claim 13 is rejected under the same rationale.
Regarding Claim 14, the claim recites limitations that are all substantially the same as limitations of Claim 1, in the form of a hardware system to implement the agent. For substantially the same reasoning, the claim is also directed to performing mental processes and mathematical concepts without integration into a practical application or amounting to either significantly more or an inventive concept.
Accordingly, Claim 14 is rejected under the same rationale.
Regarding Claim 15, the claim recites limitations that are all substantially the same as limitations of Claim 7, in the form of a hardware system to implement the environment. For substantially the same reasoning, the claim is also directed to performing mental processes and mathematical concepts without integration into a practical application or amounting to either significantly more or an inventive concept.
Accordingly, Claim 15 is rejected under the same rationale.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 6-8, 10-11, and 13-15 are rejected under 35 U.S.C. 103 as being unpatentable over Anderson et al. (hereinafter Anderson) (“Learning to Evade Static PE Machine Learning Malware Models via Reinforcement Learning”) in view of Zeng et al. (hereinafter Zeng) (Pat. App. Pub. No. US 2020/0134887 A1).
Regarding Claim 1, Anderson teaches a training method for training an autonomous agent implementing a reinforcement learning algorithm to improve performance of anti-malware software, the method being implemented by the autonomous agent and comprising (Pg. 5, Col. 1, Para. 3, “we train an ACER agent to learn a policy for our framework depicted in Figure 1” and Pg. 8, Col. 1, Para. 6, “If one assumes that our automated adversary is equally or more capable than attackers in the wild, then this approach represents a genuinely valuable means for generating evasive variants for studying model weaknesses or for adversarial training”, where “our framework” is a training method for training an autonomous agent, “we train an ACER agent to learn a policy” in order to act as “our automated adversary”, which implements a reinforcement learning algorithm, see Pg. 4, Col. 2, Para. 3-4, “We implement our black-box attack using a reinforcement learning approach . . . A reinforcement learning model consists of an agent and an environment that interact for a sequence of turns . . . The agent learns incrementally”, and which is to improve performance of anti-malware software, see Pg. 2, Col. 1, Para. 3, “We test in Section 5 whether adversarially-crafted malicious samples can be used to harden a machine learning model via adversarial training. In particular, by retraining a machine learning model using evasive ransomware variants, the evasion rate of a new ransomware attack drops from 12% to 8%”; see also Pg. 5, Col. 1, Fig. 1, where, as depicted in Fig. 1, the method, the framework for “Markov decision process formulation of the mal ware evasion reinforcement learning problem”, is implemented, in part, by the autonomous agent, “agent”):
selecting an action aimed at modifying the content of a malware (Pg. 4, Col. 2, Para. 4, “For each turn t, an agent may choose an action at ∈ A”; Pg. 5, Col. 2, Para. 7, “the file mutations represent the actions or moves available to the agent within the environment. There are a modest number of modifications that can be made to a PE file”; and Pg. 4, Col. 2, Para. 2, “our approach begins with a pool of malicious PE files and attempts binary manipulations that create an evasive variant. To our knowledge, our approach is the only work to date that produces valid PE malware samples”, where an action is selected, “an agent may choose an action”, which is aimed at modifying the content of a malware, “modifications that can be made to a PE file”, where the “PE file[s]” are “malicious PE files . . . that produces valid PE malware samples”);
transmitting the selected action to an environment implementing said anti-malware software (Pg. 4, Col. 2, Para. 4, “The environment produces a reward rt ∈ R in response to a chosen action” and Pg. 5, Col. 2, Para. 2, “The environment consists of an initial malware sample (one malware sample per “game”), and a customizable anti-malware engine (the attack target)”, where the selected action, “a chosen action”, must be transmitted from the agent and to an “environment” for the “environment” to “produces a reward rt ∈ R in response to a chosen action”, and where the “environment” implements anti-malware, “a customizable anti-malware engine”);
receiving, from the environment, a reward representative of a probability that the malware modified by application of the selected action is considered benign by the anti-malware software (Pg. 4, Col. 2, Para. 4, “The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1. The reward rt and observed state of the environment st+1 are fed back to the agent”; Pg. 5, Col. 1, Para. 3, “The reward function is measured by the anti malware engine, which is converted to a reward: 0 if the modified malware sample is judged to be malicious (no evasion), and R if it is deemed to be benign (evasion). The reward and state are then fed back into the agent”; and Pg. 6, Col. 1, Para. 4, “we attack a gradient boosted decision tree model trained on 100,000 malicious and benign samples, and which achieves an area under the receiver operating characteristic score (ROC-AUC) of 0.993. In our experiments, we set a threshold of 0.9 for the static malware model that approximately corresponds to a 1% false positive rate at a 90% true positive rate”, where “a reward” is received from the environment, “The environment produces a reward . . . [that is] fed back to the agent”, representative of a probability, “we set a threshold of 0.9 for the static malware model that approximately corresponds to a 1% false positive rate at a 90% true positive rate”, that the malware modified by application of the selected action, “produces a reward rt ∈ R in response to a chosen action” that resulted in the “modified malware”, is considered benign by the anti-malware software, “a reward: 0 if the modified malware sample is judged to be malicious (no evasion), and R if it is deemed to be benign (evasion)”),
the reward being defined as:
PNG
media_image3.png
64
226
media_image3.png
Greyscale
with R a positive real value, T a threshold value specific to the anti-malware software (Pg. 5, Col. 1, Para. 3, “The reward function is measured by the anti malware engine, which is converted to a reward: 0 if the modified malware sample is judged to be malicious (no evasion), and R if it is deemed to be benign (evasion). The reward and state are then fed back into the agent”; Pg. 6, Col. 1, Para. 4, “we attack a gradient boosted decision tree model trained on 100,000 malicious and benign samples, and which achieves an area under the receiver operating characteristic score (ROC-AUC) of 0.993. In our experiments, we set a threshold of 0.9 for the static malware model that approximately corresponds to a 1% false positive rate at a 90% true positive rate”; and Pg. 4, Col. 2, Para. 4, “For each turn t . . . The environment produces a reward rt ∈ R in response”, where the “reward” is defined as r(t+1) = R ⊂ R+ if p(t+1)<T, where p(t+1), is the probability that “the modified malware sample” is “malicious” for the current “turn”, such that if < T, “we set a threshold of 0.9 for the static malware model”, then “R if it is deemed to be benign”; and 0 otherwise, “0 if the modified malware sample is judged to be malicious (no evasion)”, and where the “threshold” is a value specific to the anti-malware software, “for the static malware model that approximately corresponds to a 1% false positive rate at a 90% true positive rate”; see also Pg. 6, Col. 2, Para. 3, “rewards of R = 10 or R = 0 are provided for evasion or failed-evasion, respectively”, where “R” is within the set of positive real numbers, “10”),
p(t + 1) a detectability score representative of a probability that the malware modified by application of the selected action is considered malicious by the anti-malware software, and p(t) the detectability score by application of an action selected in a previous iteration (Pg. 5, Col. 1, Para. 3, “The reward function is measured by the anti malware engine, which is converted to a reward: 0 if the modified malware sample is judged to be malicious (no evasion), and R if it is deemed to be benign (evasion). The reward and state are then fed back into the agent”; Pg. 6, Col. 1, Para. 4, “we attack a gradient boosted decision tree model trained on 100,000 malicious and benign samples, and which achieves an area under the receiver operating characteristic score (ROC-AUC) of 0.993. In our experiments, we set a threshold of 0.9 for the static malware model that approximately corresponds to a 1% false positive rate at a 90% true positive rate”; and Pg. 4, Col. 2, Para. 4, “For each turn t . . . The environment produces a reward rt ∈ R in response”, where a detectability score representative of a probability, the value compared to the “threshold of 0.9” which “approximately corresponds to a 1% false positive rate at a 90% true positive rate”, that the malware modified by application of the selected action, “modified malware”, is considered malicious by the anti-malware software, “the modified malware sample is judged to be malicious (no evasion)”, such that p(t+1) is the detectability score for the current iteration, “turn”, whereas p(t) is the detectability score for the previous iteration, “turn”; see also Pg. 6, Col. 2, Para. 1-2, “The agent is allowed to perform up to ten mutations before declaring failure (i.e., ten rounds with exactly one mutation per round) . . . Rounds terminate early if the agent bypasses the malware model prior to the ten allotted mutations (i.e., the agent was bypassed in less than ten mutations). We allow a combined total of 50,000 mutations to train each model”);
obtaining a state representative of the malware modified by application of the selected action (Pg. 4, Col. 2, Para. 4, “The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1. The reward rt and observed state of the environment st+1 are fed back to the agent” and Pg. 5, Col. 1, Para. 3, “our framework [is] depicted in Figure 1 . . . the agent gets an estimate of the environment’s state s ∈ S, represented by a feature vector s of the malware sample . . . the actions space A consists of a set of modifications to the PE file”; Pg. 5, Col. 2, Para. 4, “in order to more concisely represent the current state of the malware sample, the environment emits the state in the form of a feature vector”; and Pg. 6, Col. 1, Para. 5, “the agent’s actions are fully observable by the state representation—that is, the agent can “see” via his feature representation the effect of his actions”, where a state, “observed state”, representative of the malware, “the environment’s state s ∈ S, represented by a feature vector s of the malware sample”, modified by application of the selected action, “the agent’s actions are fully observable by the state representation” where “the actions space A consists of a set of modifications to the PE file”, is obtained by the agent, “the environment emits the state in the form of a feature vector” and the “observed state of the environment st+1 are fed back to the agent”; see also Pg. 5, Col. 1, Fig. 1, where, as depicted in Fig. 1, obtaining a state comprises the “agent” receiving a “state” from the “environment”);
wherein selecting an action, receiving a reward and obtaining a state are iterated as long as a stopping criterion is not reached (Pg. 5, Col. 1, Fig. 1; Pg, 5, Col. 2, Para. 2, “The environment consists of an initial malware sample (one malware sample per “game”), and a customizable anti-malware engine (the attack target)”; and Pg. 6, Col. 2, Para. 2, “The “game” is comprised of several rounds. Each round begins with a known malware sample, which is modified through a series of mutations in the round”, where the above discussed selecting, receiving, and obtaining are iterated through over “several rounds” within a “game”, as long as a stopping criterion is not reached, see Pg. 6, Col. 2, Para. 1-2, “The agent is allowed to perform up to ten mutations before declaring failure (i.e., ten rounds with exactly one mutation per round) . . . Rounds terminate early if the agent bypasses the malware model prior to the ten allotted mutations (i.e., the agent was bypassed in less than ten mutations). We allow a combined total of 50,000 mutations to train each model”); and,
determining, by the reinforcement learning algorithm and based on the obtained rewards, a function which associates with each state at least one action to be executed, so as to maximize a sum of the obtained rewards (Pg. 4, Col. 2, Para. 3-4, “We implement our black-box attack using a reinforcement learning approach . . . For each turn t, an agent may choose an action at ∈ A based on a policy π(a|st) and an observable environmental state vector st. The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1. The reward rt and observed state of the environment st+1 are fed back to the agent to choose a new action based on policy π(a|st+1) . . . The goal of the agent is to derive a policy that maximizes the expected return defined by . . . [EQUATION] . . . This function that estimates the expected utility of taking a given action for a given state is called a Q-function”; Pg. 6, Col. 1, Para. 5, “the agent’s actions are fully observable by the state representation—that is, the agent can “see” via his feature representation the effect of his actions”; and Pg. 5, Col. 1, Para. 3, “The Q-function and action policy determine what action to take. In our framework, the actions space A consists of a set of modifications to the PE file”, where, by the reinforcement learning algorithm, “a reinforcement learning approach”, and based on the obtained rewards, “The reward rt . . . are fed back to the agent”, a function is determined, “defined by . . . [EQUATION] . . . This function that estimates the expected utility of taking a given action for a given state is called a Q-function”, which associates with each state at least one action to be executed, “The Q-function and action policy determine what action to take. In our framework, the actions space A consists of a set of modifications to the PE file” where “the agent’s actions are fully observable by the state representation”, so as to maximize a sum of the obtained rewards, “The goal of the agent is to derive a policy that maximizes the expected return”).
Anderson does not explicitly disclose . . . p(t+1) – p(t) . . . (where, if the otherwise condition is applicable, then the reward value is 0).
However, Zeng teaches . . . [reward =] p(t+1) – p(t) . . . (Para. [0033], “The probability output from the trained discriminators is an image metric. The reward is defined by subtracting the metric between the current acquisition step and the previous step”; see also Para. [0033] – [0034], “In this embodiment, the reward is the negative difference in probability between acquisition steps . . . Other embodiments of the invention may use different metrics. For example, the . . . metric could be any arbitrary metric”, where “the metric”, can be “any arbitrary metric”, including “difference in probability between . . . steps”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the reward defined such that, if p(t+1) < T, then r(t+1) = R ⊂ R+, otherwise r(t+1) = 0, wherein R is a positive real value, T is a threshold value specific to the anti-malware software, p(t + 1) is a detectability score representative of a probability that the malware modified by application of the selected action is considered malicious by the anti-malware software, and p(t) is the detectability score by application of an action selected in a previous iteration of Anderson with the reward = p(t+1) – p(t) of Zeng in order to efficiently generate a reward signal in cases where the modified malware is detected as malicious (compare Anderson, Pg. 5, Col. 1, Para. 3, “The reward function is measured by the anti malware engine, which is converted to a reward: 0 if the modified malware sample is judged to be malicious (no evasion)” with Zeng, Para. [0033] – [0034], “The reward is defined by subtracting the metric between the current acquisition step and the previous step . . . Other embodiments of the invention may use different metrics. For example, the . . . metric could be any arbitrary metric”, where the broad applicability of Zeng’s approach, “the . . . metric could be any arbitrary metric”, will allow it to be efficiently incorporated into Anderson’s approach without computing an additional metric), which will reduce the number of turns without progress (compare Anderson, Pg. 4, Col. 2, Para. 4, “The reward provides the key objective for learning, and notably, may be zero for many turns until a target state is reached through a relatively long series of actions” and Anderson, Pg. 6, Col. 2, Para. 1, “long sequences of moves that finally produce a reward can lead to complications in training the reinforcement learning agent” with Zeng, Para. [0033], “The probability output from the trained discriminators is an image metric. The reward is defined by subtracting the metric between the current acquisition step and the previous step”, where an influential reward value is generated at each step), and may contribute to a reinforcement learning framework that provides nearly optimal results (Zeng, Para. [0033], “The probability output from the trained discriminators is an image metric. The reward is defined by subtracting the metric between the current acquisition step and the previous step” and Zeng, Para. [0050] - [0051], “From the result in TABLE 1 . . .”, where the approach at issue achieved superior performance, compared with other methods, in multiple instances; see generally Zeng, Para. [0052], “The reinforcement learning framework provides nearly optimal results”) thus allowing the framework to harden machine learning models against adversarial evasion attacks against more capable adversaries or adversaries with additional information (Anderson, Pg. 4, Col. 2, Para. 1, “since our ultimate goal is to harden machine learning models against adversarial evasion attacks, one may assume that our automated adversary is at least as capable as the most realistic threat for static Windows PE files, and thus use the generated adversarial samples for model hardening”), while preserving Anderson’s approach for cases where the modified malware is detected as benign (Anderson, Pg. 5, Col. 1, Para. 3, “The reward function is measured by the anti malware engine, which is converted to a reward: . . . R if it is deemed to be benign (evasion). The reward and state are then fed back into the agent”), which already provides sufficient information to the agent, mimics the approaches used by real-world attackers (Anderson, Pg. 4, Col. 2, Para. 1, “we believe this study most closely follows approaches used by real-world adversaries— systematically probing anti-malware engines in an attempt to capture and summarize blind spots”), and preserves the framework’s ability to terminate early (Anderson, Pg. 6, Col. 2, Para. 2, “Rounds terminate early if the agent bypasses the malware model prior to the ten allotted mutations (i.e., the agent was bypassed in less than ten mutations)”).
Regarding Claim 2, Anderson in view of Zeng teach the training method according to claim 1, wherein obtaining a state comprises either receiving the state from the environment; or obtaining the malware on which an action can be applied, and determining the state based on the selected action and on the obtained malware (Anderson, Pg. 5, Col. 1, Fig. 1, where, as depicted in Fig. 1, obtaining a state comprises the “agent” receiving a “state” from the “environment”; see also Anderson, Pg. 5, Col. 1, Para. 3, “our framework [is] depicted in Figure 1 . . . the agent gets an estimate of the environment’s state s ∈ S, represented by a feature vector s of the malware sample”).
Regarding Claim 6, Anderson teaches a non-transitory computer-readable recording medium on which a computer program is recorded comprising instructions which when executed by a processor of an autonomous agent . . . (Pg. 4, Col. 2, Para. 3-4, “METHOD We implement our black-box attack using a reinforcement learning approach . . . For each turn t, an agent may choose an action at ∈ A based on a policy π(a|st) and an observable environmental state vector st. The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1. The reward rt and observed state of the environment st+1 are fed back to the agent to choose a new action based on policy π(a|st+1)” and Pg. 8, Col. 2, Para. 2, “we have provided open source code at https://github. com/endgameinc/gym-malware. We believe that there is significant room for improvement in this approach, and encourage researchers to contribute”, where a person of ordinary skill in the art would understand the “black-box attack using a reinforcement learning approach” to require a computer program comprising instructions, which must be stored, at least temporarily, in a non-transitory computer-readable recording medium in order to be “implement[ed]” or “provided [as] open source code at https://github. com/endgameinc/gym-malware”, and where said implementation, such as “an agent may choose an action at ∈ A based on a policy π(a|st) and an observable environmental state vector st” would require a processor of an “agent” to configure the “agent” to implement the “METHOD”; see also Pg. 5, Col. 1, Fig. 1, where, as depicted in Fig. 1, the method, “Markov decision process formulation of the malware evasion reinforcement learning problem”, is implemented in part by the “agent” and, as a result, must have an associated processor configured to implement the method by executing a computer program recorded on a non-transitory computer-readable medium comprising instructions, see Pg. 5, Col. 1-2, Para. 4-1, “With an aim to engage the broader community, we implement our malware evasion environment as an extensible OpenAI gym [5], which we release at https://github.com/endgameinc/gym-malware”; see also Pg. 5, Col. 1, Para. 3, “we train an ACER agent to learn a policy for our framework depicted in Figure 1” and Pg. 8, Col. 1, Para. 6, “If one assumes that our automated adversary is equally or more capable than attackers in the wild, then this approach represents a genuinely valuable means for generating evasive variants for studying model weaknesses or for adversarial training”, where “our framework” is a method for an autonomous agent, “we train an ACER agent to learn a policy” in order to act as “our automated adversary”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Regarding Claim 7, Anderson teaches an evaluation method for evaluating detectability of a malware by an environment implementing at least one anti-malware software, the method comprising (Pg. 5, Col. 2, Para. 2, “The environment consists of an initial malware sample (one malware sample per “game”), and a customizable anti-malware engine (the attack target)” and Pg. 2, Col. 1, Para. 3, “We test in Section 5 whether adversarially-crafted malicious samples can be used to harden a machine learning model via adversarial training. In particular, by retraining a machine learning model using evasive ransomware variants, the evasion rate of a new ransomware attack drops from 12% to 8%”, where an environment, “environment”, is for an evaluation method for evaluating the malware, see “the evasion rate of a new ransomware attack drops from 12% to 8%”, by anti-malware software, “a customizable anti-malware engine”; see also Pg. 1, Col. 1, Abstract, “an RL agent is equipped with a set of functionality-preserving operations that it may perform on the PE file. Through a series of games played against the anti-malware engine, it learns which sequences of operations are likely to result in evading the detector for any given malware sample” and Pg. 5, Col. 2, Para. 2, “A reward value ∈ {0,R}, where 0 denotes the malware sample was detected by the anti-malware engine and R is the reward given for evading the engine”, where the evaluating is of the detectability of a malware, “where 0 denotes the malware sample was detected by the anti-malware engine”, by anti-malware software, “the engine”; see also Pg. 5, Col. 2, Para. 2, “The environment consists of an initial malware sample (one malware sample per “game”), and a customizable anti-malware engine (the attack target)”, where the “environment” implements at least one anti-malware software, “a customizable anti-malware engine (the attack target)”):
receiving, from an autonomous agent implementing a reinforcement learning algorithm, an action aimed at modifying content of the malware (Pg. 4, Col. 2, Para. 4, “For each turn t, an agent may choose an action at ∈ A . . . The environment produces a reward rt ∈ R in response to a chosen action”; Pg. 5, Col. 2, Para. 7, “the file mutations represent the actions or moves available to the agent within the environment. There are a modest number of modifications that can be made to a PE file”; and Pg. 4, Col. 2, Para. 2, “our approach begins with a pool of malicious PE files and attempts binary manipulations that create an evasive variant. To our knowledge, our approach is the only work to date that produces valid PE malware samples”, where “an agent may choose an action”, which is aimed at modifying the content of a malware, “modifications that can be made to a PE file”, where the “PE file[s]” are “malicious PE files . . . that produces valid PE malware samples”, and is received by the environment, “The environment produces a reward rt ∈ R in response to a chosen action”; see also Pg. 4, Col. 2, Para. 3-4, “We implement our black-box attack using a reinforcement learning approach . . . A reinforcement learning model consists of an agent and an environment that interact for a sequence of turns . . . The agent learns incrementally” and Pg. 8, Col. 1, Para. 6, “If one assumes that our automated adversary is equally or more capable than attackers in the wild, then this approach represents a genuinely valuable means for generating evasive variants for studying model weaknesses or for adversarial training”, where the agent is an autonomous agent, “our automated adversary”, implementing a reinforcement learning algorithm, “A reinforcement learning model consists of an agent”);
modifying the content of the malware by application of said action, so as to obtain a modified malware (Pg. 6, Col. 2, Para. 2, “Each round begins with a known malware sample, which is modified through a series of mutations in the round”; Pg. 5, Col. 2, Para. 7, “the file mutations represent the actions or moves available to the agent within the environment. There are a modest number of modifications that can be made to a PE file that do not break the PE file format and do not alter code execution”; and Pg. 4, Col. 2, Para. 4, “The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1”, where the content of the malware is modified, “There are a modest number of modifications that can be made to a PE file”, by application of said action, “The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1”, so as to obtain a modified malware, “a known malware sample, which is modified through a series of mutations”);
analyzing, by the anti-malware software, the modified malware (Pg. 5, Col. 1, Para. 3, “The reward function is measured by the anti malware engine, which is converted to a reward: 0 if the modified malware sample is judged to be malicious (no evasion), and R if it is deemed to be benign (evasion). The reward and state are then fed back into the agent”, where the modified malware, “the modified malware sample”, is analyzed, “The reward function is measured . . . judged to be malicious . . . deemed to be benign”, by the anti-malware software, “measured by the anti malware engine”);
and transmitting, to the autonomous agent, a reward . . . (Pg. 4, Col. 2, Para. 4, “The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1. The reward rt and observed state of the environment st+1 are fed back to the agent”, where a “reward” is transmit to the autonomous agent, “are fed back to the agent”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Regarding Claim 8, Anderson in view of Zeng teach the evaluation method according to claim 7, further comprising generating an association between the action, and either the score p(t + 1), or the reward r(t + 1) in an association table (Anderson, Pg. 7, Col. 1, Para. 3, “Dominant mutations. We summarize in Table 3 the dominant mutations selected by both the agent and by the random policy for the holdout samples that successfully evaded the machine learning model” and Anderson, Pg. 7, Col. 1, Table 3, “Dominant successful mutations selected by the RL agent and by the random policy. Shown in parentheses is the median number of mutations required for successful evasion”, where an association between the action, “Dominant successful mutations” that comprise one or more actions “the median number of mutations”, and the score p(t + 1) and the reward r(t + 1) is generated, “We summarize . . . Shown in parentheses”, because only “Dominant successful mutations” are included, which associates the action as having “successfully evaded the machine learning model”, thus establishing it has a p(t + 1) below the threshold so as to qualify as a “successful evasion” and a r(t+1) of “10”, see Anderson, Pg. 5, Col. 2, Para. 2, “In our gym and in experiments, we use R = 10”, in an association table, “Table 3”, where “We summarize” associate, “the dominant mutations selected by both the agent and by the random policy for the holdout samples that successfully evaded the machine learning model”; see also Anderson, Pg. 5, Col. 1, Para. 3, “The reward function is measured by the anti malware engine, which is converted to a reward: 0 if the modified malware sample is judged to be malicious (no evasion), and R if it is deemed to be benign (evasion). The reward and state are then fed back into the agent” and Anderson, Pg. 6, Col. 1, Para. 4, “we attack a gradient boosted decision tree model trained on 100,000 malicious and benign samples, and which achieves an area under the receiver operating characteristic score (ROC-AUC) of 0.993. In our experiments, we set a threshold of 0.9 for the static malware model that approximately corresponds to a 1% false positive rate at a 90% true positive rate”).
Regarding Claim 10, Anderson teaches a non-transitory computer-readable recording medium on which a computer program is recorded comprising instructions which when executed by a processor of an environment configure the environment to implement a method . . . (Pg. 4, Col. 2, Para. 3-4, “METHOD We implement our black-box attack using a reinforcement learning approach . . . For each turn t, an agent may choose an action at ∈ A based on a policy π(a|st) and an observable environmental state vector st. The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1. The reward rt and observed state of the environment st+1 are fed back to the agent to choose a new action based on policy π(a|st+1)” and Pg. 8, Col. 2, Para. 2, “we have provided open source code at https://github. com/endgameinc/gym-malware. We believe that there is significant room for improvement in this approach, and encourage researchers to contribute”, where a person of ordinary skill in the art would understand the “black-box attack using a reinforcement learning approach” to require a computer program comprising instructions, which must be stored, at least temporarily, in a non-transitory computer-readable recording medium in order to be “implement[ed]” or “provided [as] open source code at https://github. com/endgameinc/gym-malware”, and where said implementation, such as “environment produces a reward rt ∈ R” would require a processor of an “environment” to configure the “environment” to implement the “METHOD”; see also Pg. 5, Col. 1, Fig. 1, where, as depicted in Fig. 1, the method, “Markov decision process formulation of the mal ware evasion reinforcement learning problem”, is implemented in part by the “environment” and, as a result, must have an associated processor configured to implement the method by executing a computer program recorded on a non-transitory computer-readable medium comprising instructions, see Pg. 5, Col. 1-2, Para. 4-1, “With an aim to engage the broader community, we implement our malware evasion environment as an extensible OpenAI gym [5], which we release at https://github.com/endgameinc/gym-malware”).
The remaining limitations are substantially the same as limitations of Claim 7, therefore it is rejected under the same rationale.
Regarding Claim 11, Anderson in view of Zeng teach a method for training anti-malware software implementing a learning algorithm, the method comprising (Anderson, Pg. 7, Col. 2, Para. 2, “Model hardening. To test whether a model can be hardened by adversarial training, we took the 1543 samples discovered during training of the ransomware dataset and add them to the 100K training set of the original GBDT model and retrained . . . We then used the previously-trained reinforcement learning agent to manipulate the 200 holdout samples (which neither agent or model have seen) and test the evasion efficacy of the retrained model. In this case, the reinforcement learning agent successfully discovered evasive variants for 8% (down from 12%) of the samples, a drop of 33%”, where a method for training a model is disclosed, “add them to the 100K training set of the original GBDT model and retrained”, which implements a learning algorithm, “reinforcement learning”, and where the “original GBT model” is anti-malware software, see Anderson, Pg. 6, Col. 2, Para. 4, “we attack a gradient boosted decision tree model” and Anderson, Pg. 5, Col. 2, Para. 2, “The environment consists of an initial malware sample (one malware sample per “game”), and a customizable anti-malware engine (the attack target)”):
obtaining a plurality of modified malwares (Anderson, Pg. 6, Col. 2, Para. 6, “During training of the reinforcement learning agent, we save malware samples that result in an evasion. The number of evasive variants discovered during training with a fixed budget of 50K mutations are summarized for each category in Table 1”; see also Anderson, Pg. 2, Col. 1, Para. 3, “We test in Section 5 whether adversarially-crafted malicious samples can be used to harden a machine learning model via adversarial training”)
in accordance with a method for evaluating detectability of a malware by an environment implementing at least one anti- malware software (Anderson, Pg. 5, Col. 2, Para. 2, “The environment consists of an initial malware sample (one malware sample per “game”), and a customizable anti-malware engine (the attack target)” and Anderson, Pg. 2, Col. 1, Para. 3, “We test in Section 5 whether adversarially-crafted malicious samples can be used to harden a machine learning model via adversarial training. In particular, by retraining a machine learning model using evasive ransomware variants, the evasion rate of a new ransomware attack drops from 12% to 8%”, where an environment, “environment”, is for an evaluation method for evaluating the malware, see “the evasion rate of a new ransomware attack drops from 12% to 8%”, by anti-malware software, “a customizable anti-malware engine”; see also Anderson, Pg. 1, Col. 1, Abstract, “an RL agent is equipped with a set of functionality-preserving operations that it may perform on the PE file. Through a series of games played against the anti-malware engine, it learns which sequences of operations are likely to result in evading the detector for any given malware sample” and Anderson, Pg. 5, Col. 2, Para. 2, “A reward value ∈ {0,R}, where 0 denotes the malware sample was detected by the anti-malware engine and R is the reward given for evading the engine”, where the evaluating is of the detectability of a malware, “where 0 denotes the malware sample was detected by the anti-malware engine”, by anti-malware software, “the engine”; see also Anderson, Pg. 5, Col. 2, Para. 2, “The environment consists of an initial malware sample (one malware sample per “game”), and a customizable anti-malware engine (the attack target)”, where the “environment” implements at least one anti-malware software, “a customizable anti-malware engine (the attack target)”)
according to claim 7 (see the rejection of claim 7),
each malware of the plurality having a detectability score (p(t + 1)) representative of a probability that the modified malware is considered malicious by the anti-malware software (Anderson, Pg. 5, Col. 1, Para. 3, “The reward function is measured by the anti malware engine, which is converted to a reward: 0 if the modified malware sample is judged to be malicious (no evasion), and R if it is deemed to be benign (evasion). The reward and state are then fed back into the agent”; Anderson, Pg. 6, Col. 1, Para. 4, “we attack a gradient boosted decision tree model trained on 100,000 malicious and benign samples, and which achieves an area under the receiver operating characteristic score (ROC-AUC) of 0.993. In our experiments, we set a threshold of 0.9 for the static malware model that approximately corresponds to a 1% false positive rate at a 90% true positive rate”; and Anderson, Pg. 4, Col. 2, Para. 4, “For each turn t . . . The environment produces a reward rt ∈ R in response”, where a detectability score representative of a probability, the value compared to the “threshold of 0.9” which “approximately corresponds to a 1% false positive rate at a 90% true positive rate”, that the modified malware, “modified malware”, is considered malicious by the anti-malware software, “the modified malware sample is judged to be malicious (no evasion)”, such that p(t+1) is the detectability score for the current iteration, “turn”, which is assigned for each malware of the plurality, see Anderson, Pg. 6, Col. 2, Para. 6, “During training of the reinforcement learning agent, we save malware samples that result in an evasion. The number of evasive variants discovered during training with a fixed budget of 50K mutations are summarized for each category in Table 1”),
the score of each malware from the plurality being less than a defined value (Anderson, Pg. 5, Col. 1, Para. 3, “The reward function is measured by the anti malware engine, which is converted to a reward: 0 if the modified malware sample is judged to be malicious (no evasion), and R if it is deemed to be benign (evasion). The reward and state are then fed back into the agent”; Anderson, Pg. 6, Col. 1, Para. 4, “we attack a gradient boosted decision tree model trained on 100,000 malicious and benign samples, and which achieves an area under the receiver operating characteristic score (ROC-AUC) of 0.993. In our experiments, we set a threshold of 0.9 for the static malware model that approximately corresponds to a 1% false positive rate at a 90% true positive rate”; and Anderson, Pg. 2, Col. 1, Para. 3, “We test in Section 5 whether adversarially-crafted malicious samples can be used to harden a machine learning model via adversarial training. In particular, by retraining a machine learning model using evasive ransomware variants”, where each of the plurality of malware are “adversarially-crafted malicious samples . . . [that are] evasive ransomware variants”, which have scores less than a defined value, where the malicious samples must have a score below the “threshold of 0.9” to be evasive);
labeling said malwares as malicious (Anderson, Pg. 4, Col. 2, Para. 1, “since our ultimate goal is to harden machine learning models against adversarial evasion attacks, one may assume that our automated adversary is at least as capable as the most realistic threat for static Windows PE files, and thus use the generated adversarial samples for model hardening” and Anderson, Pg. 7, Col. 2, Para. 2, “Model hardening. To test whether a model can be hardened by adversarial training, we took the 1543 samples discovered during training of the ransomware dataset and add them to the 100K training set of the original GBDT model and retrained . . . We then used the previously-trained reinforcement learning agent to manipulate the 200 holdout samples (which neither agent or model have seen) and test the evasion efficacy of the retrained model. In this case, the reinforcement learning agent successfully discovered evasive variants for 8% (down from 12%) of the samples, a drop of 33%. It is important to reiterate that the attack efficacy is measured only on malware samples that the machine learning model initially labels as malicious”; and Anderson, Pg. 7, Col. 1, Table 1, where “the 1543 samples” are labeled in table 1 as malicious “dataset” “ransomware” “evasions” “1543” as a perquisite selection for “use the generated adversarial samples for model hardening”; additionally model outputs, which can reasonably be described as labels and for “1543 samples” will include some outputs that a malware sample is malicious, can reasonably be described as labeling said malware as malicious)
and, training the anti-malware software with the labeled malwares (Anderson, Pg. 7, Col. 2, Para. 2, “Model hardening. To test whether a model can be hardened by adversarial training, we took the 1543 samples discovered during training of the ransomware dataset and add them to the 100K training set of the original GBDT model and retrained . . . We then used the previously-trained reinforcement learning agent to manipulate the 200 holdout samples (which neither agent or model have seen) and test the evasion efficacy of the retrained model. In this case, the reinforcement learning agent successfully discovered evasive variants for 8% (down from 12%) of the samples, a drop of 33%. It is important to reiterate that the attack efficacy is measured only on malware samples that the machine learning model initially labels as malicious”, where a method for training a model is disclosed, “add them to the 100K training set of the original GBDT model and retrained”, which, uses the labeled malwares, “we took the 1543 samples discovered during training of the ransomware dataset and add them to the 100K training set of the original GBDT model and retrained”, and where the “original GBT model” is anti-malware software, see Anderson, Pg. 6, Col. 2, Para. 4, “we attack a gradient boosted decision tree model” and Anderson, Pg. 5, Col. 2, Para. 2, “The environment consists of an initial malware sample (one malware sample per “game”), and a customizable anti-malware engine (the attack target)”).
Regarding Claim 13, Anderson teaches a non-transitory computer-readable recording medium on which a computer program is recorded comprising instructions which when executed by a processor of an electronic device configure the electronic device to implement a training method . . . (Pg. 4, Col. 2, Para. 3-4, “METHOD We implement our black-box attack using a reinforcement learning approach . . . For each turn t, an agent may choose an action at ∈ A based on a policy π(a|st) and an observable environmental state vector st. The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1. The reward rt and observed state of the environment st+1 are fed back to the agent to choose a new action based on policy π(a|st+1)” and Pg. 8, Col. 2, Para. 2, “we have provided open source code at https://github. com/endgameinc/gym-malware. We believe that there is significant room for improvement in this approach, and encourage researchers to contribute”, where a person of ordinary skill in the art would understand the “black-box attack using a reinforcement learning approach” to require a computer program comprising instructions, which must be stored, at least temporarily, in a non-transitory computer-readable recording medium in order to be “implement[ed]” or “provided [as] open source code at https://github. com/endgameinc/gym-malware”, and where said implementation, such as “environment produces a reward rt ∈ R” or “an agent may choose an action at ∈ A” would require a processor of an electronic device to configure the “agent” and the “environment” to implement the “METHOD”; see also Pg. 5, Col. 1, Fig. 1, where, as depicted in Fig. 1, the method, “Markov decision process formulation of the mal ware evasion reinforcement learning problem”, is implemented by the device or devices associated with the “agent” and the “environment” which, as a result, must have an associated processor configured to implement the method by executing a computer program recorded on a non-transitory computer-readable medium comprising instructions, see Pg. 5, Col. 1-2, Para. 4-1, “With an aim to engage the broader community, we implement our malware evasion environment as an extensible OpenAI gym [5], which we release at https://github.com/endgameinc/gym-malware”).
The remaining limitations are substantially the same as limitations of Claim 11, therefore it is rejected under the same rationale.
Regarding Claim 14, Anderson teaches . . . the agent comprising: at least one processor; and at least one non-transitory computer readable medium comprising instructions stored thereon which when executed by the at least one processor configure the agent . . . (Pg. 4, Col. 2, Para. 3-4, “METHOD We implement our black-box attack using a reinforcement learning approach . . . For each turn t, an agent may choose an action at ∈ A based on a policy π(a|st) and an observable environmental state vector st. The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1. The reward rt and observed state of the environment st+1 are fed back to the agent to choose a new action based on policy π(a|st+1)” and Pg. 8, Col. 2, Para. 2, “we have provided open source code at https://github. com/endgameinc/gym-malware. We believe that there is significant room for improvement in this approach, and encourage researchers to contribute”, where a person of ordinary skill in the art would understand the “black-box attack using a reinforcement learning approach” to require a computer program comprising instructions, which must be stored, at least temporarily, in a non-transitory computer-readable recording medium in order to be “implement[ed]” or “provided [as] open source code at https://github. com/endgameinc/gym-malware”, and where said implementation, such as “an agent may choose an action at ∈ A based on a policy π(a|st) and an observable environmental state vector st” would require the non-transitory medium to be a comprising component of the agent and for the “agent” to comprise a processor to configure the “agent” to implement the “METHOD”; see also Pg. 5, Col. 1, Fig. 1, where, as depicted in Fig. 1, the method, “Markov decision process formulation of the mal ware evasion reinforcement learning problem”, is implemented in part by the “agent” and, as a result, must have an associated non-transitory computer readable medium and a processor configured to implement the method by executing a computer program recorded on the non-transitory computer-readable medium comprising instructions, see Pg. 5, Col. 1-2, Para. 4-1, “With an aim to engage the broader community, we implement our malware evasion environment as an extensible OpenAI gym [5], which we release at https://github.com/endgameinc/gym-malware”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Regarding Claim 15, Anderson teaches an environment for evaluating the detectability of a malware by anti-malware software (Pg. 5, Col. 2, Para. 2, “The environment consists of an initial malware sample (one malware sample per “game”), and a customizable anti-malware engine (the attack target)” and Pg. 2, Col. 1, Para. 3, “We test in Section 5 whether adversarially-crafted malicious samples can be used to harden a machine learning model via adversarial training. In particular, by retraining a machine learning model using evasive ransomware variants, the evasion rate of a new ransomware attack drops from 12% to 8%”, where the environment, “environment”, is for evaluating the malware, see “the evasion rate of a new ransomware attack drops from 12% to 8%”, by anti-malware software, “a customizable anti-malware engine”; see also Pg. 1, Col. 1, Abstract, “an RL agent is equipped with a set of functionality-preserving operations that it may perform on the PE file. Through a series of games played against the anti-malware engine, it learns which sequences of operations are likely to result in evading the detector for any given malware sample” and Pg. 5, Col. 2, Para. 2, “A reward value ∈ {0,R}, where 0 denotes the malware sample was detected by the anti-malware engine and R is the reward given for evading the engine”, where the evaluating is of the detectability of a malware, “where 0 denotes the malware sample was detected by the anti-malware engine”, by anti-malware software, “the engine”),
the environment comprising: at least one processor; and at least one non-transitory computer readable medium comprising instructions stored thereon which when executed by the at least one processor configure the environment to implement the method . . . (Pg. 4, Col. 2, Para. 3-4, “METHOD We implement our black-box attack using a reinforcement learning approach . . . For each turn t, an agent may choose an action at ∈ A based on a policy π(a|st) and an observable environmental state vector st. The environment produces a reward rt ∈ R in response to a chosen action as well as new environmental state vector st+1. The reward rt and observed state of the environment st+1 are fed back to the agent to choose a new action based on policy π(a|st+1)” and Pg. 8, Col. 2, Para. 2, “we have provided open source code at https://github. com/endgameinc/gym-malware. We believe that there is significant room for improvement in this approach, and encourage researchers to contribute”, where a person of ordinary skill in the art would understand the “black-box attack using a reinforcement learning approach” to require a computer program comprising instructions, which must be stored, at least temporarily, in a non-transitory computer-readable recording medium in order to be “implement[ed]” or “provided [as] open source code at https://github. com/endgameinc/gym-malware”, and where said implementation, such as “The environment produces a reward rt ∈ R” would require the non-transitory medium to be a comprising component of the “environment” and for the “environment” to comprise a processor to configure the “environment” to implement the “METHOD”; see also Pg. 5, Col. 1, Fig. 1, where, as depicted in Fig. 1, the method, “Markov decision process formulation of the malware evasion reinforcement learning problem”, is implemented in part by the “environment” and, as a result, must have an associated non-transitory computer readable medium and a processor configured to implement the method by executing a computer program recorded on the non-transitory computer-readable medium comprising instructions, see Pg. 5, Col. 1-2, Para. 4-1, “With an aim to engage the broader community, we implement our malware evasion environment as an extensible OpenAI gym [5], which we release at https://github.com/endgameinc/gym-malware”).
The remaining limitations are substantially the same as limitations of Claim 7, therefore it is rejected under the same rationale.
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Anderson in view of Zeng and Tostaeva (“Introduction to Q-learning with OpenAI Gym”).
Regarding Claim 3, Anderson in view of Zeng teach the training method according to claim 1, wherein the reinforcement learning algorithm is a "Q-learning" algorithm . . . (Anderson, Pg. 4, Col. 2, Para. 3-4, “We implement our black-box attack using a reinforcement learning Approach . . . A reinforcement learning model consists of an agent and an environment that interact for a sequence of turns (or discrete timesteps) . . . The goal of the agent is to derive a policy that maximizes the expected return defined by Vπ(st) = Eat [Qπ(st,at)|st] with Qπ(st,at) = Est+1:∞,at+1:∞[Rt |st,at] and Rt = i≥0γirt+i
where γ ∈ [0,1] discounts the amount of reward from future actions . . . This function that estimates the expected utility of taking a given action for a given state is called a Q-function”, where the “reinforcement learning” algorithm is a Q-learning algorithm because it utilizes a “Q-function”; see generally Anderson, Pg. 4-5, Col. 2-1, Para. 5-1, “Among the key contributions of the deep reinforcement learning framework was its ability, as in deep learning, for the agent to learn a value function in an end-to-end way: it takes raw pixels as input, and outputs predicted rewards for each action. This learned value function is the basis for so-called deep Q-learning, where the Q-function is learned and refined over hundreds of games”).
Anderson in view of Zeng . . . and the determination of the function comprises a determination, for each state-action pair (s, a), of a value QN(s, a) such that:
PNG
media_image2.png
32
402
media_image2.png
Greyscale
with α ⊂ [0,1] a learning rate, Q(s, a) a previous quality value, r(t + 1) a reward, y ⊂ [0,1] a refresh rate, s(t + 1) a next state and a(t + 1) an action that can be executed from the state s(t + 1), so as to determine an optimal Q-function.
However, Tostaeva teaches . . . [a reinforcement learning algorithm that is a "Q-learning" algorithm] (Pg. 1, Para. 1, “walk you through an example of using Q-learning to solve a reinforcement learning problem”)
and the determination of the function comprises a determination, for each state-action pair (s, a), of a value QN(s, a) such that (Pg. 2, Para. 2, “Q-learning is an off-policy algorithm meant to determine the best action given the current state . . . What Q-learning does is measure how good a state-action combination is in terms of rewards. It does so by keeping track of a Q-table, a reference matrix, or simply a look-up table, that gets updated after each episode with its row corresponding to the state and its column to the action. An episode ends after a set of actions is completed. In the end, the Q-table suggests the optimal policy”, where the determination of the function comprises a determination, for each state-action pair (s, a), “an off-policy algorithm meant to determine the best action given the current state”, of a value QN(s, a), “It does so by keeping track of a Q-table . . . that gets updated after each episode with its row corresponding to the state and its column to the action. An episode ends after a set of actions is completed”, such that “a special mathematical formula” is used, see Pg. 3, Para. 3-4, “Once determined, we take the action and observe the reward it yields . . . Finally, we update our Q-table using a special mathematical formula”):
PNG
media_image2.png
32
402
media_image2.png
Greyscale
with α ⊂ [0,1] a learning rate, Q(s, a) a previous quality value, r(t + 1) a reward, y ⊂ [0,1] a refresh rate, s(t + 1) a next state and a(t + 1) an action that can be executed from the state s(t + 1), so as to determine an optimal Q-function (Pg. 4-5, Para. 3-1, “The special mathematical formula used for the update in Step 5 is known as the Bellman equation, or:
PNG
media_image4.png
40
654
media_image4.png
Greyscale
where alpha is the learning rate and gamma is the discount factor; s, a, r refer to state, action, and reward, respectively . . . Rewards lose their values over time, and the discount factor reflects that . . . We use the equation to determine the value of the maximum reward expected for each state in the Q-table. We do so by taking the received reward r and the future state s’. The first term, Q(s_t, a_t), is the value of the current action in the current state. The second is a bit more complicated: we combine the current reward, r_t, and the discounted value of the future state given the highest-reward yielding action, γ max_a Q(s_{t+1}, a_t). We then take away the current state value (underlying the first term). This second term is multiplied by α, the learning rate, which is simply the importance we put on the future value compared to the present one (hence the (1- α) is multiplied by the first term)”, where α is the learning rate, “alpha is the learning rate”, and α ⊂ [0,1] because “0.7” is a subset in [0,1], see Pg. 8, Pseudocode, “alpha = 0.7 #learning rate”; where Q(s, a), “Q(s_t, a_t)”, is a quality value, “Q(s_t, a_t), is the value of the current action”, which is operated on as a previous quality value, “We then take away the current state value (underlying the first term)”, with respect to the new information “current reward, r_t, and the discounted value of the future state”, see Pg. 10, Para. 3, “Once we choose an action, we carry on with it and measure the associated reward. This is done using the built-in env.step(action) method which makes a one timestep move. It returns the next state, the reward from the previous action”; where r(t + 1) is a reward, “the current reward, r_t” such that the divergence between the label r(t+1) and “r_t” can reasonably be described as a naming convention without a substantive difference; where y, “gamma”, is “the discount factor”, which is within the broadest reasonable interpretation of a refresh rate because it refreshes “Rewards” to reflect “their values over time” and y ⊂ [0,1] because “0.618” is a subset in [0,1], see Pg. 8, Pseudocode, “discount_factor = 0.618”; where s(t + 1) is a next state, “the future state . . . s_{t+1}”; and where a(t + 1) is an action that can be executed from the state s(t + 1), so as to determine an optimal Q-function, “the future state given the highest-reward yielding action, γ max_a Q(s_{t+1}, a_t)”, such that the divergence between the label a(t+1) and “a_t” can reasonably be described as a naming convention without a substantive difference; see also Pg. 2, Para. 4, “In the end, the Q-table suggests the optimal policy”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the reinforcement learning algorithm that is a "Q-learning" algorithm of Anderson in view of Zeng with the reinforcement learning algorithm that is a "Q-learning" algorithm and the determination of the function comprises a determination, for each state-action pair (s, a), of a value QN(s, a) such that:
PNG
media_image5.png
40
503
media_image5.png
Greyscale
with α ⊂ [0,1] a learning rate, Q(s, a) a previous quality value, r(t + 1) a reward, y ⊂ [0,1] a refresh rate, s(t + 1) a next state and a(t + 1) an action that can be executed from the state s(t + 1), so as to determine an optimal Q-function of Tostaeva in order to utilize the reinforcement learning policy to achieve the optimal policy for action selection (compare Anderson, Pg. 4, Col. 2, Para. 3-4, “We implement our black-box attack using a reinforcement learning approach . . . This function that estimates the expected utility of taking a given action for a given state is called a Q-function” with Tostaeva, Pg. 2, Para. 4, “What Q-learning does is measure how good a state-action combination is in terms of rewards. It does so by keeping track of a Q-table . . . In the end, the Q-table suggests the optimal policy” and Tostaeva, Pg. 1, Para. 1, “walk you through an example of using Q-learning to solve a reinforcement learning problem”), through a simple to implement method (Tostaeva, Pg. 5, Para. 2, “We will first briefly describe the OpenAI Gym environment for our problem and then use Python to implement the simple Q-learning algorithm in our environment”) that is compatible with the Anderson’s reduced capacity implementation environment (compare Anderson, Pg. 5, Col. 1, Para. 4, “With an aim to engage the broader community, we implement our malware evasion environment as an extensible OpenAI gym” with Tostaeva, Pg. 5, Para. 3, “To make sure we are all on the same page, an environment in OpenAI gym is basically a test problem — it provides the bare minimum needed to have an agent interacting with a world. The primary purpose is to test the agent and learning algorithms without having to worry about simulating the environment”).
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Anderson in view of Zeng and Boutsikas et al. (hereinafter Boutsikas) (“Evading Malware Classifiers via Monte Carlo Mutant Feature Discovery”).
Regarding Claim 4, Anderson in view of Zeng teach the training method according to claim 1, wherein the action is selected from a set of actions consisting of (Anderson, Pg. 5-6, Col. 2-1, Para. 7-1, “the file mutations represent the actions or moves available to the agent within the environment. There are a modest number of modifications that can be made to a PE file that do not break the PE file format and do not alter code execution . . . Some of these include: . . .”):
modifying a value of a field of a header of the malware (Anderson, Pg. 6, Col. 1, Para. 1, “modifying (breaking) header checksum”, where the “header checksum”, which is a value of a field of a header of the malware, is “modif[ied]”);
adding to the content of the malware a sequence of characters extracted from a benign file (Anderson, Pg. 6, Col. 1, Para. 1-2, “manipulating existing section names creating new (unused) sections . . . when renaming a section, a new section name is drawn uniformly from a list of section names found in benign files”, where a sequence of characters extracted from a benign file, “a list of section names found in benign files”, is added to the content of the malware, “manipulating existing section names creating new (unused) sections”);
adding to the content of the malware determined characters or instructions (Anderson, Pg. 6, Col. 1, Para. 1, “appending bytes to extra space at the end of sections . . . creating a new entry point which immediately jumps to the original entry point”, where determined characters are added to the content of the malware, “appending bytes to extra space at the end of sections”, and instructions are added to the content of the malware, “creating a new entry point which immediately jumps to the original entry point”);
. . .
renaming a section of the content of the malware (Anderson, Pg. 6, Col. 1, Para. 1, “manipulating existing section names”, where a section of the content of the malware, “existing section”, is renamed, “manipulating . . . names”);
. . .
modifying a hash value calculated for an optional header of the content of the malware (Anderson, Pg. 6, Col. 1, Para. 1, “modifying (breaking) header checksum”, where a “checksum”, which is calculated for a header of the content of the malware, “header”, which is within the broadest reasonable interpretation of a hash value because is a value mapped to input data using a function, is “modif[ied]”, which must be optional for “breaking” to “not break the PE file format” and “not alter code execution”, see Anderson, Pg. 5, Col. 2, Para. 7, “There are a modest number of modifications that can be made to a PE file that do not break the PE file format and do not alter code execution”); and,
decompressing an executable version of the malware (Anderson, Pg. 6, Col. 1, Para. 1, “packing or unpacking the file”, where the malware, which is an executable version, see Anderson, Pg. 5, Col. 2, Para. 7, “There are a modest number of modifications that can be made to a PE file that do not break the PE file format and do not alter code execution” and Anderson, Pg. 7, Col. 2, Para. 3, “For a random sampling of ten evasive variants generated from the VirusShare dataset, we discovered that only eight executed properly in a virtual machine”, is decompressed, “unpacking the file”).
Anderson in view of Zeng do not explicitly discuss . . . adding to the content of the malware a library extracted from a benign file . . . removing a debugger mode from the content of the malware; modifying a timestamp of the content of the malware . . . .
However, Boutsikas teaches . . . [selecting an action aimed at modifying the content of a malware, the action selected from a set of actions consisting of] (Pg. 1, Col. 1, Abstract, “In this experiment, a malicious actor trains a surrogate model using the EMBER-2018 dataset to discover binary mutations that cause an instance to be misclassified via a Monte Carlo tree search. Then, mutated malware is sent to the victim model that takes the place of an antivirus API to test whether it can evade detection”; see also Pg. 4, Col. 2, Para. 2, “Here we list the set of possible mutations available to our Monte Carlo implementation, and their corresponding heuristic that limits the changes: . . .”)
adding to the content of the malware a library extracted from a benign file . . . (Pg. 5, Col. 1, Para. 5, “Import Function: This mutation will add a function to the import table of the malware, as well as the matching DLL if it is missing. In order to decide what those functions would be, we made a list of the most common functions in the benign samples that do not appear in malware samples. This gave us 14 candidate functions, so each application of this mutation will select one of those 14 at random”, where a library, “a function . . . as well as the matching DLL if it is missing”, extracted from a benign file, “functions in the benign samples that do not appear in malware samples”, is added to the content of the malware, “This mutation will add a function to the import table of the malware, as well as the matching DLL”)[;]
removing a debugger mode from the content of the malware (Pg. 5, Col. 1, Para. 7, “Remove Debug: This mutation will set the debug flag to false. The mutation can only be applied if the sample’s debug flag is set to true”, where a debugger mode from the content of the malware is removed, “if the sample’s debug flag is set to true” then “This mutation will set the debug flag to false”); [and]
modifying a timestamp of the content of the malware . . . (Pg. 5, Col. 1, Para. 6, “Change Timestamp: This mutation will modify the timestamp of the sample towards a target timestamp, by the given step size”, where a timestamp of the content of the malware, “the timestamp of the sample”, is “modif[ied]”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the selection of an action from a set of actions aimed at modifying the content of a malware of Anderson in view of Zeng with the selecting an action aimed at modifying the content of a malware, the action selected from a set of actions consisting of: adding to the content of the malware a library extracted from a benign file; removing a debugger mode from the content of the malware; and modifying a timestamp of the content of the malware of Boutsikas in order to harden machine learning models against a greater number of adversarial evasion attacks (Anderson, Pg. 4, Col. 2, Para. 1, “since our ultimate goal is to harden machine learning models against adversarial evasion attacks, one may assume that our automated adversary is at least as capable as the most realistic threat for static Windows PE files, and thus use the generated adversarial samples for model hardening”), including modifications that mimic the library compositions of benign software (Boutsikas, Pg. 3, Col. 1, Para. 3, “the average number of library exports and imports is greater for benign software”) and simple modifications that bypass ML technology (Boutsikas, Pg. 1, Col. 2, Para. 2, “Recent work has shown that top AV that utilize some form of ML technology can be bypassed with simple modifications on the malware such as by adding a new section, appending a single byte, removing the debug and certificate values, or renaming a section . . . In our work, we borrow the feature changes that are shown to be effective in prior research, and explore a new mutant malware discovery methodology that is based on Monte Carlo Tree Search”).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Brenon et al. (“Preliminary Study of Adaptive Decision-Making System for Vocal Command in Smart Home”) discloses storing associations between an action and a reward in an association table (Pg. 218-219, Col. 2-1, Para. 5-1).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW BRYCE GOLAN whose telephone number is (571)272-5159. The examiner can normally be reached Monday through Friday, 8:00 AM to 5:00 PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW BRYCE GOLAN/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123