DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant's claim for foreign priority to JP 2024-009177. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Response to Amendment
This office action is in response to the amendment filed on 06/04/2026. Claims 7-20 are new. Claims 1, 4, and 6 are amended. Claims 1-20 are pending. Claims 1, 4, and 6 are independent.
Response to Arguments
Rejections under 35 U.S.C. § 103:
Firstly, see page 7 of Applicant’s Remarks (hereinafter “REMARKS”), filed 06/04/2026, with respect to the comment on “MANTIN in view of MANTIN”, the second MANTIN refers to the same reference as the first MANTIN. The citations refer to different embodiments disclosed in the same reference. The rejection of claims 4-6 has been updated to remove the duplicative statement of the reference, however, the same ground(s) of rejection with regard to the reference is maintained.
Applicant’s arguments, see pages 7-9 of REMARKS, with respect to the rejection of the claims under 35 U.S.C. § 103 have been fully considered and are persuasive. Specifically, that the prior art of record does not disclose all of the limitations of the amended independent claims. Therefore, in view of the amendment, a new ground(s) of rejection is made over MANTIN et al. (US PGPub No. 2025/0111051; hereinafter “MANTIN”) in view of HARUKI et al. (US PGPub No. 2023/0274005; hereinafter “HARUKI”) in view of CLEMENT et al. (US PGPub No. 2024/0386103; hereinafter “CLEMENT”) in view of SALEM et al. (US PGPub No. 2025/0175497; hereinafter “SALEM”).
SALEM teaches general purpose intelligence from other LLMs useful to the current systems and other systems (¶ 0012, ¶ 0024-0025, ¶ 0060).
Therefore, the combination of MANTIN in view of HARUKI in view of CLEMENT in view of SALEM disclose all the limitations of amended independent claim 1. Independent claims 4 and 6 are rejected over MANTIN in view of SALEM.
Regarding applicant’s arguments with respect to the dependent claims, the amendment to the independent claims have necessitated a new ground(s) of rejection with respect to the independent claims from which the dependent claims depend, thereby requiring new grounds of rejection for the dependent claims.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “universal” in independent claims 1, 4, and 6 is a relative term which renders the claim indefinite. The term “universal” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. It’s unclear how intelligence gathered could be universal for all target systems, as one could not test whether gathered general intelligence would have a sufficient expectation of working on all possible systems. Given the terms use in the limitation “general-purpose intelligence universal to all target systems”, it’s also unclear what encompasses “all target systems” that the intelligence would be useable on. Therefore, the claims are rejected. Dependent claims are rejected for failing to remedy the issue.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a diagnostics unit” in claims 1-6, “a monitoring unit” in claims 8-9 and 17-18.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. The diagnostics unit is described as diagnosing inputs and outputs from an LLM in relation to predetermined attacks on that same LLM, and making determinations as to whether the system is subject to an attack from those inputs/outputs (¶ 0006-0010, ¶ 0018-0019, ¶ 0021-0026, ¶ 0036-0043). The monitoring unit is described in the specification as monitoring of the results of the input/output of the LLM, whereby new threats are registered and accumulated as dedicated intelligence and the general-purpose intelligence as a black list, and false-positives are fed back as a whitelist (¶ 0018, ¶ 0026-0027, ¶ 0029-0030, ¶ 0036-0038, ¶ 0047-0048).
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-3, 7-13, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over MANTIN et al. (US PGPub No. 2025/0111051; hereinafter “MANTIN”) in view of HARUKI et al. (US PGPub No. 2023/0274005; hereinafter “HARUKI”) in view of CLEMENT et al. (US PGPub No. 2024/0386103; hereinafter “CLEMENT”) in view of SALEM et al. (US PGPub No. 2025/0175497; hereinafter “SALEM”).
As per claim 1: MANTIN discloses a security countermeasure support system that supports diagnostics and monitoring of security of a target system using a large language model (LLM), the system comprising (a guardian controller that may be provided with a system that includes the control application, the large language model, and a controlled application (e.g., the email program in the example above). The guardian controller includes a machine learning model that monitors the output of the machine learning model. The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017]; The data repository (104) stores large language model data (106). The large language model data (106) is data input to or output from a large language model, such as the large language model (136) defined further below. Thus, the large language model data (106) includes a first input (108) and a first output (110). The large language model data (106) may be text data. Thus, the first input (108) is text input and the first output (110) is text output [MANTIN ¶ 0023]):
a diagnostics unit that diagnoses (a guardian controller that may be provided with a system that includes the control application, the large language model, and a controlled application ( e.g., the email program in the example above). The guardian controller includes a machine learning model that monitors the output of the machine learning model. The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017]), on a basis of a response from the LLM (The server (128) may include a large language model (136) [MANTIN ¶ 043]) to a [predetermined] pseudo-attack on the LLM used in the target system (takes, as input, the prompt injection cyberattack and generates a first output [MANTIN ¶ 0005]; the guardian controller (138) may be programmed to monitor the first output (110) of the large language model (136). The guardian controller (138) may be programmed to determine the probability (122) that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102). The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045]), whether or not the predetermined attack has succeeded (The guardian controller (138) may be programmed to determine the probability (122) that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102). The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045, Examiner’s Note: satisfying the threshold is succeeding as it’s passing the threshold for qualifying as a prompt injection]) [with reference to an attack signature as intelligence], wherein
the intelligence comprises dedicated intelligence unique to the target system (generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: the historical queries relevant to the system are dedicated intelligence]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [¶ 0096]) and [general-purpose intelligence universal to all target systems], and
the [predetermined] attack includes, in a user prompt to be input to the LLM, [information that violates a command in a system prompt input to the LLM] (takes, as input, the prompt injection cyberattack and generates a first output [MANTIN ¶ 0005]; The user query (120) is a query received from one of the user devices (146). In embodiment, the user query (120) may be received from the malicious user device (100), in which case the user query (120) may be referred to as a malicious query. The user query (120) may be received at the server (128) or may be received directly by one of the control application (132) or the large language model (136) [MANTIN ¶ 0030]), [information that violates a command in a system prompt input to the LLM] (In a direct attack, a malicious user modifies a large language model's input in an attempt to overwrite existing system prompts [¶ 0004]).
MANTIN discloses the claimed subject matter as discussed above but does not explicitly disclose with reference to an attack signature accumulated as intelligence. However, MANTIN teaches with reference to an attack signature accumulated as intelligence (generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: known prompt injections are attack signatures accumulated and represented as historical queries]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [¶ 0096]). MANTIN and the instant application are analogous art because they are from the same field of endeavor of LLM security. Therefore, based on MANTIN in view of MANTIN, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of MANTIN to the system of MANTIN in order to monitor the output based on historical examples for improved detection. Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
MANTIN discloses the claimed subject matter as discussed above but does not explicitly disclose predetermined pseudo-attack; predetermined attack. However, HARUKI teaches predetermined pseudo-attack (The verification execution unit attacks a verification environment in which at least one of attack countermeasures indicated by attack countermeasure information is applied to a verification target system by using each of a plurality of attack scenarios, and creates a possible attack scenario list that is a list of attack scenarios in which an attack has succeeded [HARUKI ¶ 0043]; the storage unit 14 stores verification target system definition information 14A, an attack scenario database (DB) 14B [HARUKI ¶ 0052]); predetermined attack (The verification execution unit attacks a verification environment in which at least one of attack countermeasures indicated by attack countermeasure information is applied to a verification target system by using each of a plurality of attack scenarios, and creates a possible attack scenario list that is a list of attack scenarios in which an attack has succeeded [HARUKI ¶ 0043]; the storage unit 14 stores verification target system definition information 14A, an attack scenario database (DB) 14B [HARUKI ¶ 0052]). MANTIN and HARUKI are analogous art because they are from the same field of endeavor of security testing and evaluation. Therefore, based on MANTIN in view of HARUKI, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of HARUKI to the system of MANTIN in order to improve robustness of security through verification of the system against known attack scenarios. Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
MANTIN in view of HARUKI discloses the claimed subject matter as discussed above but does not explicitly disclose information that violates a command in a system prompt input to the LLM. However, CLEMENT teaches information that violates a command in a system prompt input to the LLM (Initially, the large language model is given initial instructions 202 to perform the target task. System prompts (<system> . . . </system>) are created by the developer of the conversational interactions with the model to constrain the model to acting in prescribed ways consistent with the intent and policies of the service hosting the large language model. For example, as shown in FIG. 2, the initial instructions 202 indicate, in part, that “you will always begin every response by repeating the SECRET in the user prompt.” This is the original goal (i.e., intent/policy) of the large language model. When the model generates a response that does not repeat the SECRET then it is assumed that a prompt injection attack has occurred [CLEMENT ¶ 0036]). CLEMENT and the instant application are analogous art because they are from the same field of endeavor of LLM security. Therefore, based on MANTIN in view of HARUKI in view of CLEMENT, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of CLEMENT to the system of MANTIN in view of HARUKI in order to provide another layer of security to prompt injection in the form of a verifiable secret. Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
MANTIN in view of HARUKI in view of CLEMENT discloses the claimed subject matter as discussed above but does not explicitly disclose general-purpose intelligence universal to all target systems. However, SALEM teaches general-purpose intelligence universal to all target systems (utilizing an attack defense system to improve the defense robustness of a targeted large generative model (LGM) by generating a set of variant prompt injection attacks that are successful against the targeted LGM, where the set of variants is based on a prompt injection attack (e.g., jailbreak) against the targeted LGM or another LGM [¶ 0012, ¶ 0024-0025]; In some implementations, an LGM security system monitors prompt injection attacks against different LGMs and when a prompt injection attack is identified at one of the LGMs, the LGM security system provides it to the prompt variation generator 212 of the attack defense system 206 to improve the security robustness of the targeted LGM 240 [¶ 0060]). SALEM and the instant application are analogous art because they are from the same field of endeavor of large language model security. Therefore, based on MANTIN in view of HARUKI in view of CLEMENT in view of SALEM, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of SALEM to the system of MANTIN in view of HARUKI in view of CLEMENT in order to improve the detection and training of the current system through the benefit of attacks detected on other models (¶ 0019, ¶ 0024-0025). Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
As per claim 2: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, MANTIN and CLEMENT disclose wherein the user prompt (takes, as input, the prompt injection cyberattack and generates a first output [MANTIN ¶ 0005]; The user query (120) is a query received from one of the user devices (146). In embodiment, the user query (120) may be received from the malicious user device (100), in which case the user query (120) may be referred to as a malicious query. The user query (120) may be received at the server (128) or may be received directly by one of the control application (132) or the large language model (136) [MANTIN ¶ 0030]) includes a command to output content of the system prompt (In goal hijacking, the inserted text is used to confuse the model or cause it to forget its instructions, allowing the user to ask the model questions which violate the rules of interaction set out in the initial or system prompt… In prompt leaking, the unintended goal is to print out a portion of or the whole original or system prompt which may be used for malicious purposes [CLEMENT ¶ 0004]).
As per claim 3: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, MANTIN and CLEMENT disclose wherein the user prompt (takes, as input, the prompt injection cyberattack and generates a first output [MANTIN ¶ 0005]; The user query (120) is a query received from one of the user devices (146). In embodiment, the user query (120) may be received from the malicious user device (100), in which case the user query (120) may be referred to as a malicious query. The user query (120) may be received at the server (128) or may be received directly by one of the control application (132) or the large language model (136) [MANTIN ¶ 0030]) includes a command to ignore or avoid a restriction related to the command of the system prompt (In goal hijacking, the inserted text is used to confuse the model or cause it to forget its instructions, allowing the user to ask the model questions which violate the rules of interaction set out in the initial or system prompt… In prompt leaking, the unintended goal is to print out a portion of or the whole original or system prompt which may be used for malicious purposes [CLEMENT ¶ 0004]).
As per claim 7: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, MANTIN and SALEM disclose wherein content of the dedicated intelligence is updated on a basis of a result of diagnostics in which the diagnostics unit diagnoses, on a basis of a response from the LLM to a predetermined pseudo- attack on the LLM used in the target system, whether or not the predetermined attack has succeeded (generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: known prompt injections are attack signatures accumulated as historical queries]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [MANTIN ¶ 0096, Examiner’s Note: incorporating the trained model into the detection based on the learned attack signatures]; The attack defense system improves attack detection accuracy by determining sets of variant prompt injection attacks that are successful and effective against a targeted LGM [SALEM ¶ 0019]; By generating a large, reliable set of variant prompt injection attacks that are successful against a targeted LGM, the attack defense system improves attack detection and accuracy. For example, the attack defense system uses a set of successful variant prompt injection attacks to fine-tune the targeted LGM and/or a corresponding classifier. This is greatly advantageous as obtaining accurate training data of a newly discovered prompt injection attack is very difficult [SALEM ¶ 0024]; Furthermore, the attack defense system is flexibly transferable between LGMs. For instance, for a prompt injection attack discovered to be used against a first targeted LGM, the attack defense system quickly, easily, but safely adapts the attack to penetrate other targeted LGMs. For example, the attack defense system automatically generates sets of variant prompt injection attacks that are tailored to successfully attack another targeted LGM. Then, the attack defense system can fortify the other LGM against successful versions of the prompt injection attack. This way, the attack defense system can build up the defense of many different LGMs before threat actors can discover ways to apply a new prompt injection attack to the other LGMs [SALEM ¶ 0025]).
As per claim 8: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, MANTIN and HARUKI disclose further comprising a monitoring unit that monitors an input and output with respect to the LLM used in the target system and detects presence or absence of an attack on the LLM by one or more predetermined methods with reference to the attack signature accumulated as the intelligence (a guardian controller that may be provided with a system that includes the control application, the large language model, and a controlled application (e.g., the email program in the example above). The guardian controller includes a machine learning model that monitors the output of the machine learning model. The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017]; The data repository (104) stores large language model data (106). The large language model data (106) is data input to or output from a large language model, such as the large language model (136) defined further below. Thus, the large language model data (106) includes a first input (108) and a first output (110). The large language model data (106) may be text data. Thus, the first input (108) is text input and the first output (110) is text output [MANTIN ¶ 0023]; generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: known prompt injections are attack signatures accumulated and represented as historical queries]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [MANTIN ¶ 0096]; The verification execution unit attacks a verification environment in which at least one of attack countermeasures indicated by attack countermeasure information is applied to a verification target system by using each of a plurality of attack scenarios, and creates a possible attack scenario list that is a list of attack scenarios in which an attack has succeeded [HARUKI ¶ 0043]; the storage unit 14 stores verification target system definition information 14A, an attack scenario database (DB) 14B [HARUKI ¶ 0052]).
As per claim 9: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 8. Furthermore, MANTIN and SALEM disclose wherein a threat newly detected by the monitoring unit as a result of the monitoring is registered and accumulated in the dedicated intelligence or the general-purpose intelligence, and is fed back so that the intelligence is utilized in both a diagnostic service by a red team and a monitoring service by a blue team (By generating a large, reliable set of variant prompt injection attacks that are successful against a targeted LGM, the attack defense system improves attack detection and accuracy. For example, the attack defense system uses a set of successful variant prompt injection attacks to fine-tune the targeted LGM and/or a corresponding classifier. This is greatly advantageous as obtaining accurate training data of a newly discovered prompt injection attack is very difficult [SALEM ¶ 0024]; To illustrate, the robustness measures 712 include model fine-tuning. For example, the defense robust model 710 generates a training dataset from the variant prompt injection attacks 430 and the prompt variation effectiveness scores 632, which is used to fine-tune the targeted LGM 240 (and/or a corresponding classifier) to detect a broader, more creative, and more complex style of prompt injection attacks. When the attack defense system 206 generates training data sets for various prompt injection attacks, it can significantly fortify the targeted LGM 240 against prompt injection attacks [SALEM ¶ 0148-0149]; The machine learning model (140) of the guardian controller (138) may be a classification machine learning model, which may be either a supervised or an unsupervised machine learning model [MANTIN ¶ 0046]).
As per claim 10: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, SALEM discloses wherein the general-purpose intelligence comprises universal intelligence considered to be usable in all target systems and an attack method whose effectiveness has been confirmed in a target system other than the target system (utilizing an attack defense system to improve the defense robustness of a targeted large generative model (LGM) by generating a set of variant prompt injection attacks that are successful against the targeted LGM, where the set of variants is based on a prompt injection attack (e.g., jailbreak) against the targeted LGM or another LGM [SALEM ¶ 0012, ¶ 0024-0025]; In some implementations, an LGM security system monitors prompt injection attacks against different LGMs and when a prompt injection attack is identified at one of the LGMs, the LGM security system provides it to the prompt variation generator 212 of the attack defense system 206 to improve the security robustness of the targeted LGM 240 [SALEM ¶ 0060]).
As per claim 11: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, MANTIN and CLEMENT disclose wherein the diagnostics unit diagnoses whether or not the attack has succeeded by designating one or more methods from a plurality of methods comprising heuristic scoring (The data repository (104) also may store a probability (122). The probability (122), as used herein, refers to one or more probabilities. Nevertheless, the probability (122) is one or more values output by the machine learning model (140) of the guardian controller (138), both defined further below. The probability (122) represents a likelihood that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102) [MANTIN ¶ 0031]; The data repository (104) also may store a threshold (124). The threshold (124) is a value that represents a point at which the probability (122) is considered high enough that the first output (110) of the large language model (136) will be treated as having been poisoned by the prompt injection cyberattack (102). Thus, the probability (122) may be compared to the threshold (124), and if the threshold (124) is satisfied by the probability (122), then the first output (110) of the large language model (136) is treated as being poisoned by the prompt injection cyberattack (102) [MANTIN ¶ 0032]; the machine learning model of the guardian controller may be replaced by rules or policies that determine the probability [MANTIN ¶ 0073]), LLM scoring, vector scoring, and a canary token (Initially, the large language model is given initial instructions 202 to perform the target task. System prompts (<system> . . . </system>) are created by the developer of the conversational interactions with the model to constrain the model to acting in prescribed ways consistent with the intent and policies of the service hosting the large language model. For example, as shown in FIG. 2, the initial instructions 202 indicate, in part, that “you will always begin every response by repeating the SECRET in the user prompt.” This is the original goal (i.e., intent/policy) of the large language model. When the model generates a response that does not repeat the SECRET then it is assumed that a prompt injection attack has occurred [CLEMENT ¶ 0036]).
As per claim 12: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, MANTIN discloses wherein the diagnostics unit diagnoses whether or not the attack has succeeded by performing scoring regarding a threat (a guardian controller that may be provided with a system that includes the control application, the large language model, and a controlled application ( e.g., the email program in the example above). The guardian controller includes a machine learning model that monitors the output of the machine learning model. The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017]) using one or more methods selected from the group consisting of: heuristic scoring based on an empirical rule (The data repository (104) also may store a probability (122). The probability (122), as used herein, refers to one or more probabilities. Nevertheless, the probability (122) is one or more values output by the machine learning model (140) of the guardian controller (138), both defined further below. The probability (122) represents a likelihood that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102) [MANTIN ¶ 0031]; The data repository (104) also may store a threshold (124). The threshold (124) is a value that represents a point at which the probability (122) is considered high enough that the first output (110) of the large language model (136) will be treated as having been poisoned by the prompt injection cyberattack (102). Thus, the probability (122) may be compared to the threshold (124), and if the threshold (124) is satisfied by the probability (122), then the first output (110) of the large language model (136) is treated as being poisoned by the prompt injection cyberattack (102) [MANTIN ¶ 0032]; the machine learning model of the guardian controller may be replaced by rules or policies that determine the probability [MANTIN ¶ 0073]); LLM scoring in which an inquiry is made to an external or internal LLM; and vector scoring based on a similarity between vectorized text of the user prompt and the response and a signature.
As per claim 13: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, CLEMENT discloses wherein the diagnostics unit diagnoses whether or not the attack has succeeded by instructing the LLM to output a token including a predetermined character string at an end of processing, and checking whether the token is correctly output in the response from the LLM (Initially, the large language model is given initial instructions 202 to perform the target task. System prompts (<system> . . . </system>) are created by the developer of the conversational interactions with the model to constrain the model to acting in prescribed ways consistent with the intent and policies of the service hosting the large language model. For example, as shown in FIG. 2, the initial instructions 202 indicate, in part, that “you will always begin every response by repeating the SECRET in the user prompt.” This is the original goal (i.e., intent/policy) of the large language model. When the model generates a response that does not repeat the SECRET then it is assumed that a prompt injection attack has occurred [CLEMENT ¶ 0036]).
As per claim 15: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, MANTIN discloses wherein upon reception of a result of diagnostics indicating that an adversarial attack is detected, the system performs a countermeasure comprising at least one of outputting a warning (However, if the probability satisfies the threshold, then the guardian controller (432) implements a security scheme on the control application (400). The security scheme may be to alert the user (402) that the content of the summarization may be manipulated by a prompt injection cyberattack [MANTIN ¶ 0119, ¶ 0118]), stopping the processing of the target system, or storing the detected adversarial attack as a log.
Claims 4, 6, 17-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over MANTIN in view of SALEM et al. (US PGPub No. 2025/0175497; hereinafter “SALEM”).
As per claim 4: MANTIN 1 discloses a security countermeasure support system that supports diagnostics and monitoring of security of a target system using a large language model (LLM), the system comprising (a guardian controller that may be provided with a system that includes the control application, the large language model, and a controlled application ( e.g., the email program in the example above). The guardian controller includes a machine learning model that monitors the output of the machine learning model. The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017, Fig. 1A]; The data repository (104) stores large language model data (106). The large language model data (106) is data input to or output from a large language model, such as the large language model (136) defined further below. Thus, the large language model data (106) includes a first input (108) and a first output (110). The large language model data (106) may be text data. Thus, the first input (108) is text input and the first output (110) is text output [MANTIN ¶ 0023]):
a diagnostics unit that obtains an input and output (takes, as input, the prompt injection cyberattack and generates a first output [MANTIN ¶ 0005]; the guardian controller (138) may be programmed to monitor the first output (110) of the large language model (136). The guardian controller (138) may be programmed to determine the probability (122) that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102). The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045, Fig. 1A, Examiner’s Note: server is interpreted as the diagnostic unit]) with respect to the LLM (The server (128) may include a large language model (136) [MANTIN ¶ 043]) used in the target system (a guardian controller that may be provided with a system that includes the control application, the large language model, and a controlled application ( e.g., the email program in the example above). The guardian controller includes a machine learning model that monitors the output of the machine learning model. The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017]), and diagnoses, on a basis of the input and output, presence or absence of an attack on the LLM (The guardian controller (138) may be programmed to determine the probability (122) that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102). The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045, Examiner’s Note: satisfying the threshold is succeeding as it’s passing the threshold for qualifying as a prompt injection]) by one or more predetermined methods (The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045, Fig. 1A, Examiner’s Note: satisfying the threshold is succeeding as it’s passing the threshold for qualifying as a prompt injection]; The data repository (104) also may store a threshold (124). The threshold (124) is a value that represents a point at which the probability (122) is considered high enough that the first output (110) of the large language model (136) will be treated as having been poisoned by the prompt injection cyberattack (102). Thus, the probability (122) may be compared to the threshold (124) [MANTIN ¶ 0032, Examiner’s Note: threshold stored prior to detection in data repository]) [with reference to an attack signature accumulated as intelligence] comprising dedicated intelligence unique to the target system (generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: the historical queries relevant to the system are dedicated intelligence]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [¶ 0096]) and [general-purpose intelligence universal to all target systems].
MANTIN discloses the claimed subject matter as discussed above but does not explicitly disclose with reference to an attack signature accumulated as intelligence. However, MANTIN teaches with reference to an attack signature accumulated as intelligence (generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: known prompt injections are attack signatures accumulated as historical queries]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [¶ 0096, Examiner’s Note: incorporating the trained model into the detection based on the learned attack signatures]). MANTIN and MANTIN are analogous art because they are from the same field of endeavor of LLM security. Therefore, based on MANTIN in view of MANTIN, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of MANTIN to the system of MANTIN in order to monitor the output based on historical examples for improved detection. Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
MANTIN discloses the claimed subject matter as discussed above but does not explicitly disclose general-purpose intelligence universal to all target systems. However, SALEM teaches general-purpose intelligence universal to all target systems (utilizing an attack defense system to improve the defense robustness of a targeted large generative model (LGM) by generating a set of variant prompt injection attacks that are successful against the targeted LGM, where the set of variants is based on a prompt injection attack (e.g., jailbreak) against the targeted LGM or another LGM [¶ 0012, ¶ 0024-0025]; In some implementations, an LGM security system monitors prompt injection attacks against different LGMs and when a prompt injection attack is identified at one of the LGMs, the LGM security system provides it to the prompt variation generator 212 of the attack defense system 206 to improve the security robustness of the targeted LGM 240 [¶ 0060]). SALEM and the instant application are analogous art because they are from the same field of endeavor of large language model security. Therefore, based on MANTIN in view of SALEM, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of SALEM to the system of MANTIN in order to improve the detection and training of the current system through the benefit of attacks detected on other models (¶ 0019, ¶ 0024-0025). Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
As per claim 6: MANTIN discloses a security countermeasure support system that supports diagnostics and monitoring of security of a target system using a large language model (LLM), the system comprising (a guardian controller that may be provided with a system that includes the control application, the large language model, and a controlled application ( e.g., the email program in the example above). The guardian controller includes a machine learning model that monitors the output of the machine learning model. The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017, Fig. 1A]; The data repository (104) stores large language model data (106). The large language model data (106) is data input to or output from a large language model, such as the large language model (136) defined further below. Thus, the large language model data (106) includes a first input (108) and a first output (110). The large language model data (106) may be text data. Thus, the first input (108) is text input and the first output (110) is text output [MANTIN ¶ 0023]):
a diagnostics unit that obtains an input and output with respect to the LLM used in the target system, and diagnoses (takes, as input, the prompt injection cyberattack and generates a first output [MANTIN ¶ 0005]; the guardian controller (138) may be programmed to monitor the first output (110) of the large language model (136). The guardian controller (138) may be programmed to determine the probability (122) that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102). The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045, Fig. 1A, Examiner’s Note: server is interpreted as the diagnostic unit]), on a basis of the input and output, presence or absence of an attack on the LLM (a guardian controller that may be provided with a system that includes the control application, the large language model, and a controlled application ( e.g., the email program in the example above). The guardian controller includes a machine learning model that monitors the output of the machine learning model. The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017]) by one or more predetermined methods (The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045, Fig. 1A, Examiner’s Note: satisfying the threshold is succeeding as it’s passing the threshold for qualifying as a prompt injection]; The data repository (104) also may store a threshold (124). The threshold (124) is a value that represents a point at which the probability (122) is considered high enough that the first output (110) of the large language model (136) will be treated as having been poisoned by the prompt injection cyberattack (102). Thus, the probability (122) may be compared to the threshold (124) [MANTIN ¶ 0032, Examiner’s Note: threshold stored prior to detection in data repository]) [with reference to an attack signature accumulated as intelligence] comprising dedicated intelligence unique to the target system (generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: the historical queries relevant to the system are dedicated intelligence]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [¶ 0096]) and [general-purpose intelligence universal to all target systems], [wherein content of the dedicated intelligence is updated on a basis of a result of diagnostics in which the diagnostics unit diagnoses, on a basis of a response from the LLM to a predetermined pseudo- attack on the LLM used in the target system, whether or not the predetermined attack has succeeded].
MANTIN discloses the claimed subject matter as discussed above but does not explicitly disclose with reference to an attack signature accumulated as intelligence; wherein content of the dedicated intelligence is updated on a basis of a result of diagnostics in which the diagnostics unit diagnoses, on a basis of a response from the LLM to a predetermined pseudo- attack on the LLM used in the target system, whether or not the predetermined attack has succeeded. However, MANTIN teaches with reference to an attack signature accumulated as intelligence (generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: known prompt injections are attack signatures accumulated as historical queries]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [¶ 0096, Examiner’s Note: incorporating the trained model into the detection based on the learned attack signatures]); wherein content of the dedicated intelligence is updated on a basis of a result of diagnostics in which the diagnostics unit diagnoses, on a basis of a response from the LLM to a predetermined pseudo- attack on the LLM used in the target system, whether or not the predetermined attack has succeeded (generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: known prompt injections are attack signatures accumulated as historical queries]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [¶ 0096, Examiner’s Note: incorporating the trained model into the detection based on the learned attack signatures]). MANTIN and MANTIN are analogous art because they are from the same field of endeavor of LLM security. Therefore, based on MANTIN in view of MANTIN, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of MANTIN to the system of MANTIN in order to monitor the output based on historical examples for improved detection. Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
MANTIN discloses the claimed subject matter as discussed above but does not explicitly disclose general-purpose intelligence universal to all target systems. However, SALEM teaches general-purpose intelligence universal to all target systems (utilizing an attack defense system to improve the defense robustness of a targeted large generative model (LGM) by generating a set of variant prompt injection attacks that are successful against the targeted LGM, where the set of variants is based on a prompt injection attack (e.g., jailbreak) against the targeted LGM or another LGM [¶ 0012, ¶ 0024-0025]; In some implementations, an LGM security system monitors prompt injection attacks against different LGMs and when a prompt injection attack is identified at one of the LGMs, the LGM security system provides it to the prompt variation generator 212 of the attack defense system 206 to improve the security robustness of the targeted LGM 240 [¶ 0060]). SALEM and the instant application are analogous art because they are from the same field of endeavor of large language model security. Therefore, based on MANTIN in view of SALEM, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of SALEM to the system of MANTIN in order to improve the detection and training of the current system through the benefit of attacks detected on other models (¶ 0019, ¶ 0024-0025). Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
As per claim 17: MANTIN in view of SALEM teach all the limitations of claim 6. Furthermore, MANTIN discloses further comprising a monitoring unit that monitors an input and output with respect to the LLM used in the target system and detects presence or absence of an attack on the LLM by one or more predetermined methods with reference to the attack signature accumulated as the intelligence (a guardian controller that may be provided with a system that includes the control application, the large language model, and a controlled application (e.g., the email program in the example above). The guardian controller includes a machine learning model that monitors the output of the machine learning model. The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017]; The data repository (104) stores large language model data (106). The large language model data (106) is data input to or output from a large language model, such as the large language model (136) defined further below. Thus, the large language model data (106) includes a first input (108) and a first output (110). The large language model data (106) may be text data. Thus, the first input (108) is text input and the first output (110) is text output [MANTIN ¶ 0023]; generating, by a control application, queries to a large language model, where at least some of the queries include known prompt injection cyberattacks. The queries may be generated by a data scientist manipulating the control application, or some other input scheme, for generating the queries to the large language model. In an embodiment, the generation of the queries may be performed by retrieving historical queries submitted to the large language model, some of which were known to include the prompt injection cyberattacks. The generated queries may be considered training data, in reference to FIG. 1B [MANTIN ¶ 0091, Examiner’s Note: known prompt injections are attack signatures accumulated and represented as historical queries]; The method of FIG. 3 may be varied. For example, the method of FIG. 3 also may include adding the trained machine learning model to a guardian application that monitors the monitored outputs of the large language model prior to the guardian application passing of the monitored outputs to a control application [MANTIN ¶ 0096]).
As per claim 18: MANTIN in view of SALEM teach all the limitations of claim 17. Furthermore, MANTIN and SALEM discloses wherein a threat newly detected by the monitoring unit as a result of the monitoring is registered and accumulated in the dedicated intelligence or the general-purpose intelligence, and is fed back so that the intelligence is utilized in both a diagnostic service by a red team and a monitoring service by a blue team (By generating a large, reliable set of variant prompt injection attacks that are successful against a targeted LGM, the attack defense system improves attack detection and accuracy. For example, the attack defense system uses a set of successful variant prompt injection attacks to fine-tune the targeted LGM and/or a corresponding classifier. This is greatly advantageous as obtaining accurate training data of a newly discovered prompt injection attack is very difficult [SALEM ¶ 0024]; To illustrate, the robustness measures 712 include model fine-tuning. For example, the defense robust model 710 generates a training dataset from the variant prompt injection attacks 430 and the prompt variation effectiveness scores 632, which is used to fine-tune the targeted LGM 240 (and/or a corresponding classifier) to detect a broader, more creative, and more complex style of prompt injection attacks. When the attack defense system 206 generates training data sets for various prompt injection attacks, it can significantly fortify the targeted LGM 240 against prompt injection attacks [SALEM ¶ 0148-0149]; The machine learning model (140) of the guardian controller (138) may be a classification machine learning model, which may be either a supervised or an unsupervised machine learning model [MANTIN ¶ 0046]).
As per claim 20: MANTIN in view of SALEM teach all the limitations of claim 6. Furthermore, MANTIN discloses wherein the diagnostics unit diagnoses whether or not the attack has succeeded by designating one or more methods from a plurality of methods comprising heuristic scoring (The data repository (104) also may store a probability (122). The probability (122), as used herein, refers to one or more probabilities. Nevertheless, the probability (122) is one or more values output by the machine learning model (140) of the guardian controller (138), both defined further below. The probability (122) represents a likelihood that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102) [MANTIN ¶ 0031]; The data repository (104) also may store a threshold (124). The threshold (124) is a value that represents a point at which the probability (122) is considered high enough that the first output (110) of the large language model (136) will be treated as having been poisoned by the prompt injection cyberattack (102). Thus, the probability (122) may be compared to the threshold (124), and if the threshold (124) is satisfied by the probability (122), then the first output (110) of the large language model (136) is treated as being poisoned by the prompt injection cyberattack (102) [MANTIN ¶ 0032]; the machine learning model of the guardian controller may be replaced by rules or policies that determine the probability [MANTIN ¶ 0073]), LLM scoring, vector scoring, and a canary token, and upon reception of a result of diagnostics indicating that an adversarial attack is detected (The output of the machine learning model is one or more probabilities that reflect an assessed likelihood that the output of the large language model is influenced by a prompt injection cyberattack. If the one or more probabilities satisfy a threshold, then a determination is made that the output of the large language model is influenced (i.e., poisoned) by the prompt injection cyberattack [MANTIN ¶ 0017]), the system performs a countermeasure comprising at least one of outputting a warning (However, if the probability satisfies the threshold, then the guardian controller (432) implements a security scheme on the control application (400). The security scheme may be to alert the user (402) that the content of the summarization may be manipulated by a prompt injection cyberattack [MANTIN ¶ 0119, ¶ 0118]), stopping the processing of the target system, or storing the detected adversarial attack as a log.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over MANTIN in view of SALEM in view of CLEMENT.
As per claim 5: MANTIN in view of SALEM teach all the limitations of claim 4. Furthermore, MANTIN and SALEM discloses wherein the predetermined method includes any of (The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045, Fig. 1A, Examiner’s Note: satisfying the threshold is succeeding as it’s passing the threshold for qualifying as a prompt injection]; The data repository (104) also may store a threshold (124). The threshold (124) is a value that represents a point at which the probability (122) is considered high enough that the first output (110) of the large language model (136) will be treated as having been poisoned by the prompt injection cyberattack (102). Thus, the probability (122) may be compared to the threshold (124) [MANTIN ¶ 0032, Examiner’s Note: threshold stored prior to detection in data repository]) scoring based on an empirical rule accumulated in the intelligence based on the input and output, scoring in which the LLM is inquired about whether or not the input and output correspond to the attack (The term "effectiveness score" refers to a prompt variation effectiveness score of a targeted LGM output that corresponds to a variant prompt injection attack. In various implementations, a prompt variation evaluator (e.g., prompt variant evaluation model) determines an effectiveness score for a variant prompt injection attack. The attack defense system may employ various approaches to determine effectiveness scores, as provided below. In addition, the attack defense system may also use effectiveness scores from one iteration of variant prompt injection attacks to generate improved variant prompt injection attack versions in a later iteration (e.g., "new variant prompt injection attacks") [SALEM ¶ 0035]; The attack defense system also determines an effectiveness score for each variant prompt injection attack in the set of variant prompt injection attacks using a prompt variant evaluator. In some instances, the attack defense system provides the set of variant prompt injection attacks and corresponding effectiveness scores with the system-level prompt to the LGM to generate new variant prompt injection attacks in the set of variant prompt injection attacks [¶ 0015]; For example, the prompt variation evaluation model 130 generates prompt variation effectiveness scores 132 for the variant prompt injection attacks 112 based on the targeted LGM outputs 122. In general, the prompt variation evaluation model 130 determines how effective or successful each variant prompt injection attack was at evading the defenses of the targeted LGM 120 [¶ 0041]), scoring based on similarity between vectorized text of the input and output and vectorized text of the attack signature accumulated in the intelligence (In general, the similarity comparison model 620 determines whether and/or to what extent a targeted LGM output is successful by comparing it to known successful and/or unsuccessful targeted LGM outputs. For example, the similarity comparison model 620 allows the attack defense system 206 to compare variants being tested to variants with known outputs within an embedding space [SALEM ¶ 0132]), or [determination on whether or not a predetermined canary token specified in a system prompt is included in an output from the LLM].
MANTIN in view of SALEM discloses the claimed subject matter as discussed above but does not explicitly disclose determination on whether or not a predetermined canary token specified in a system prompt is included in an output from the LLM. However, CLEMENT teaches determination on whether or not a predetermined canary token specified in a system prompt is included in an output from the LLM (A security agent is used sign a user prompt destined to a large language model from a user with a secret in order to prevent a prompt injection attack. The security agent resides on the server hosting the large language model and is isolated from the user application and user device that generates the large language model prompt. The secret is tailored for a specific user identifier and session identifier associated with the user prompt. The large language model is instructed to repeat the secret in each response. The security agent retrieves the response from the large language model and checks for the secret. When the secret is not part of the response, an error message is forwarded to the user application instead of the response [CLEMENT ¶ 0006]). MANTIN in view of SALEM and CLEMENT are analogous art because they are from the same field of endeavor of LLM security. Therefore, based on MANTIN in view of SALEM in view of CLEMENT, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of CLEMENT to the system of MANTIN in view of SALEM in order to provide another layer of LLM prompt protection to help in preventing prompt injection for improved output security. Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over MANTIN in view of HARUKI in view of CLEMENT in view of SALEM in view of Neystadt et al. (US PGPub No. 2025/0133111; hereinafter “Neystadt”).
As per claim 14: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, MANTIN discloses further comprising a monitoring unit that monitors an input and output with respect to the LLM (takes, as input, the prompt injection cyberattack and generates a first output [MANTIN ¶ 0005]; the guardian controller (138) may be programmed to monitor the first output (110) of the large language model (136). The guardian controller (138) may be programmed to determine the probability (122) that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102). The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045, Fig. 1A]), [wherein a false-positive attack, which has been detected as an attack but determined to have no problem as a result of analysis, is fed back to the intelligence as a whitelist] (The step may be to enforce at least one of a whitelist and a blacklist on the first output of the large language model. The enforcement scheme also may be to enforce at least one of the whitelist or the blacklist [MANTIN ¶ 0084]).
MANTIN in view of HARUKI in view of CLEMENT in view of SALEM discloses the claimed subject matter as discussed above but does not explicitly disclose wherein a false-positive attack, which has been detected as an attack but determined to have no problem as a result of analysis, is fed back to the intelligence as a whitelist. However, Neystadt teaches wherein a false-positive attack, which has been detected as an attack but determined to have no problem as a result of analysis, is fed back to the intelligence as a whitelist (The Email Security Agent 112 further performs logging and monitoring (arrow 162), using a Logging and Monitoring Unit 117, of each LLM-based query/response/evaluation, as well as of the WSC determined for each email message and the particular LLM-based responses and their confidence scores; thereby enabling to implement a feedback loop for monitoring and improving the accuracy of detection. For example, drifts from accurate classifications can be used to fine-tune the system/the ML model/the LLM, and/or to temporarily disable blocking or quarantining of emails to prevents “false positive” errors [¶ 0041]; In a demonstrative example, an end-user team-member of the Protected Entity utilizes an electronic device (e.g., desktop computer, laptop computer, smartphone, tablet, smart-watch) equipped with an Email Reader 121 application or module, to read or access incoming email messages (arrow 163). The end-user receives a notification from the Email Security Agent 112 and/or from the Email Server 110 with regard to messages that were quarantined, and may be provided with a mechanism to review or release such messages. The end-user further sees the relevant indicators or warnings or flags that were generated for emails that were not deleted/not quarantined. The end-user may provide feedback via a feedback loop or feedback mechanism (arrow 164), by indicating his feedback back to the Email Security Agent 112 (directly, or via the Email Server 110); with feedback such as, “yes, this email message that was flagged/quarantined as malicious is indeed malicious”, or conversely “no, this email message was incorrectly flagged/quarantined as malicious but is actually legitimate”; and in some embodiments may provide a third feedback of “I am not sure whether or not the classification as malicious is correct”. The user's feedback may be utilized by the system to fine-tune or re-train the ML units/LLM units involved in the evaluation process, to modify weights assigned to particular features or parameters or indicators, to construct or to update a white-list or a black-list of senders, or for other fine-tuning operations [¶ 0042]). Neystadt and the instant application are analogous art because they are from the same field of endeavor of LLM security. Therefore, based on MANTIN in view of HARUKI in view of CLEMENT in view of SALEM in view of Neystadt, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of Neystadt to the system of MANTIN in view of HARUKI in view of CLEMENT in view of SALEM in order improve the model through fine-tuning for improved results. Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over MANTIN in view of HARUKI in view of CLEMENT in view of SALEM in view of FREY et al. (US PGPub No. 2022/0019674; hereinafter “FREY”).
As per claim 16: MANTIN in view of HARUKI in view of CLEMENT in view of SALEM teach all the limitations of claim 1. Furthermore, MANTIN discloses further comprising [a dashboard screen that displays a list of detected attacks and a graph showing a time-series transition of scores of each detection item].
MANTIN in view of HARUKI in view of CLEMENT in view of SALEM discloses the claimed subject matter as discussed above but does not explicitly disclose a dashboard screen that displays a list of detected attacks and a graph showing a time-series transition of scores of each detection item. However, FREY teaches a dashboard screen that displays a list of detected attacks and a graph showing a time-series transition of scores of each detection item (The analytic module 116 can record the “validated analytics” 606, the “analytic gap” analytics 608, the “undetected threats” 610, the “unsuccessful analytics” 610, and the “unvalidated analytics” 612, and provide statistics for these occurrences for a given attack session, group of attack sessions, analytic test session, or group of analytic test sessions. The computer system 102 can also present the statistics, along with other cyberattack data, defense action data, time lapse data, attack-defense time lapse data, etc. via the user interface 200 to a user (see FIG. 6). This presentation can involve a video overlay (see FIGS. 7A-7C) that is a time-lapse video of when attacks and defense actions occurred. The video overlay can include a timeline with points along the timeline identifying attacks (e.g., star icons 614) and defense actions (e.g., circle icons 616). Other shapes and icons can be used. A solid star icon 614 indicates that a defense action occurred in time proximity with it that is within the attack-defense time lapse threshold. A solid circle icon 616 indicates that the defense action occurred in time proximity with an attack that is within the attack-defense time lapse threshold. An open star icon 614 indicates that a defense action did not occur in time proximity with it that is within the attack-defense time lapse threshold. An open circle icon 616 indicates that the defense action did not occurred in time proximity with an attack that is within the attack-defense time lapse threshold. Referring to FIGS. 7B-7C, for example, there is a labelled attack (red star) at the same time as the detection (blue dot), so both appear filled in because they are mapped to the same MITRE ATT&CK technique. There is a second labelled attack (hollow red star) using Windows Management Instrumentation (WMI) event subscriptions for persistence. This one does not have a corresponding analytic, so there is a detection gap. The computer system 102 can then prompt the user to create an analytic for this attack. In this second example, there are two analytics that do not correspond to a labelled attack, so they are marked as hollow blue dots. There is a labeled attack for opening a command prompt that does not have a matching analytic, so it is represented as a hollow red star. There are matching labelled attacks and analytics for using the Background Intelligence Transfer Service (BITS) jobs at 2:06, so they are filled in [¶ 0055]). FREY and the instant application are analogous art because they are from the same field of endeavor of attack visualization. Therefore, based on MANTIN in view of HARUKI in view of CLEMENT in view of SALEM in view of FREY, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of FREY to the system of MANTIN in view of HARUKI in view of CLEMENT in view of SALEM in order to improve data evaluation through visualization. Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over MANTIN in view of SALEM in view of Neystadt.
As per claim 19: MANTIN in view of SALEM teach all the limitations of claim 6. Furthermore, MANTIN discloses wherein [a false-positive attack, which has been detected as an attack but determined to have no problem as a result of analysis, is fed back to the intelligence as a whitelist] (takes, as input, the prompt injection cyberattack and generates a first output [MANTIN ¶ 0005]; the guardian controller (138) may be programmed to monitor the first output (110) of the large language model (136). The guardian controller (138) may be programmed to determine the probability (122) that the first output (110) of the large language model (136) is poisoned by the prompt injection cyberattack (102). The guardian controller (138) may be programmed to determine whether the probability (122) satisfies the threshold (124) [MANTIN ¶ 0045, Fig. 1A]), [wherein a false-positive attack, which has been detected as an attack but determined to have no problem as a result of analysis, is fed back to the intelligence as a whitelist] (The step may be to enforce at least one of a whitelist and a blacklist on the first output of the large language model. The enforcement scheme also may be to enforce at least one of the whitelist or the blacklist [MANTIN ¶ 0084]).
MANTIN in view of SALEM discloses the claimed subject matter as discussed above but does not explicitly disclose a false-positive attack, which has been detected as an attack but determined to have no problem as a result of analysis, is fed back to the intelligence as a whitelist. However, Neystadt teaches a false-positive attack, which has been detected as an attack but determined to have no problem as a result of analysis, is fed back to the intelligence as a whitelist (The Email Security Agent 112 further performs logging and monitoring (arrow 162), using a Logging and Monitoring Unit 117, of each LLM-based query/response/evaluation, as well as of the WSC determined for each email message and the particular LLM-based responses and their confidence scores; thereby enabling to implement a feedback loop for monitoring and improving the accuracy of detection. For example, drifts from accurate classifications can be used to fine-tune the system/the ML model/the LLM, and/or to temporarily disable blocking or quarantining of emails to prevents “false positive” errors [¶ 0041]; In a demonstrative example, an end-user team-member of the Protected Entity utilizes an electronic device (e.g., desktop computer, laptop computer, smartphone, tablet, smart-watch) equipped with an Email Reader 121 application or module, to read or access incoming email messages (arrow 163). The end-user receives a notification from the Email Security Agent 112 and/or from the Email Server 110 with regard to messages that were quarantined, and may be provided with a mechanism to review or release such messages. The end-user further sees the relevant indicators or warnings or flags that were generated for emails that were not deleted/not quarantined. The end-user may provide feedback via a feedback loop or feedback mechanism (arrow 164), by indicating his feedback back to the Email Security Agent 112 (directly, or via the Email Server 110); with feedback such as, “yes, this email message that was flagged/quarantined as malicious is indeed malicious”, or conversely “no, this email message was incorrectly flagged/quarantined as malicious but is actually legitimate”; and in some embodiments may provide a third feedback of “I am not sure whether or not the classification as malicious is correct”. The user's feedback may be utilized by the system to fine-tune or re-train the ML units/LLM units involved in the evaluation process, to modify weights assigned to particular features or parameters or indicators, to construct or to update a white-list or a black-list of senders, or for other fine-tuning operations [¶ 0042]). Neystadt and the instant application are analogous art because they are from the same field of endeavor of LLM security. Therefore, based on MANTIN in view of SALEM in view of Neystadt, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teaching of Neystadt to the system of MANTIN in view of SALEM in order to improve the model through fine-tuning for improved results. Hence, it would have been obvious to combine the references above to obtain the invention as specified in the instant claim.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES P MOLES whose telephone number is (703)756-1043. The examiner can normally be reached M-F 8:00am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jung Kim can be reached at (571) 272-3804. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAMES P MOLES/Examiner, Art Unit 2494
/JUNG W KIM/Supervisory Patent Examiner, Art Unit 2494
1 Although the rejection has been updated it still retains the same ground(s) of rejection with respect to the MANTIN reference. This is applicable to claims 5-6 as well.