DETAILED ACTION
Examiner acknowledges receipt of Applicant’s amendment filed on 01/02/2026
Claims 1, 6, 7, 9-12, 16-17, and 19 are currently amended
Claims 2, 18 and 20 have been cancelled
Claims 1, 3-17, and 19 are pending
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Examiner has fully considered Applicant’s amendments to the Claims in the arguments filed on 01/02/2026. Claims 1, 3-17, and 19 remain pending in the application. Examiner has withdrawn the previous claim rejections under 35 USC 112(b) and the claim objections in view of the amendments. However, additional 112(b) rejections arise.
Response to Arguments
Applicant’s arguments filed 01/02/2026, with respect to the rejections of independent claims 1 and 17 and their corresponding dependent claims under 35 USC 103 have been fully considered and are persuasive. Therefore, the rejections have been withdrawn. However, upon further consideration, new grounds of rejection are made in view of the previously applied combination of Vaknin, Majmudar, and Jain, in addition to a previously cited reference from Cappel et al. (US 20250245332 A1), hereinafter Cappel. Examiner respectfully submits that Cappel, in combination with the previously applied references, is sufficient to teach the newly added limitations “wherein scoring the user vector comprises: determining one or more vector distances of the user vector to one or more stored vectors in a set of stored vectors in a vector store, and generating, individually for the user vector, a corresponding user vector score in the set of user vector scores according to the one or more vector distance; aggregating the set of user vector scores to generate an aggregated score: detecting whether the user prompt is malicious based on whether the aggregated score satisfies a threshold”.
Applicant’s arguments filed 01/02/2026, with respect to the rejection of independent claim 19 under 35 USC 103 have been fully considered and are persuasive. Therefore, the rejections have been withdrawn. However, upon further consideration, new grounds of rejection are made in view of the previously applied combination of Yeung, Martin, and Vaknin, in addition to a newly applied reference from Nissim et al. (Nissim, N., Cohen, A., Moskovitch, R., Shabtai, A., Edri, M., BarAd, O., & Elovici, Y. (2016). Keeping pace with the creation of new malicious PDF files using an active-learning based detection framework. Security Informatics, 5(1). https://doi.org/10.1186/s13388-016-0026-3), hereinafter Nissim. Examiner respectfully submits that Nissim, in combination with the previously applied references, is sufficient to teach the newly added limitations “scoring each of the set of malicious vectors according to a second vector distance to the set of stored vectors to obtain an impact score for each of the set of malicious vectors; selecting a subset of the set of malicious vectors having at least the similarity score satisfying a similarity threshold indicating a greater vector distance with respect to the set of benign vectors, wherein the subset of the set of malicious vectors is further selected based on having the impact score satisfying an impact threshold indicating a greater vector distance with respect to the set of stored vectors”.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
In Line 9 of Claim 19, the limitation “the set of stored vectors” is unclear because there is no antecedent basis for the set of vectors. It is unclear whether the set of vectors is to be interpreted as the set of malicious vectors, the set of benign vectors, the combination of both sets of vectors, or some other set of vectors. Consequently, it is additionally unclear by what process the impact score is to be obtained. Thus, the scope of the claim is unclear. For Examination purposes, “the set of stored vectors” will be interpreted as “a set of stored vectors” to establish antecedent basis for an additional/broadly-stated set of vectors. This interpretation also represents Examiner’s suggestion for compliance with 35 USC 112(b).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3, 5, and 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaknin et al. (US 20250209208 A1), hereinafter Vaknin, in view of Majmudar et al. (US 12292915 B1), hereinafter Majmudar, and Cappel et al. (US 20250245332 A1), hereinafter Cappel.
Regarding Claim 1:
Vaknin teaches a method comprising: receiving, at a server from a user device, a user prompt to a large language model (LLM) (Vaknin – Figure 3: illustration of sequence for processing user prompt to detect prompt injection attacks, including receiving prompt 304 from user 302 at prompt analysis process 248. Paragraph [0041]: FIG. 3 illustrates an example architecture for prompt analysis for a large language model (LLM). At the core of architecture 300 is prompt analysis process 248); segmenting the user prompt to generate a set of user segments (Vaknin – Fig. 4 and Paragraph [0056]: Prior to prompt 402 being input to the LLM, prompt analysis process 248 may identify the topic(s) present in prompt 402 by converting each word in prompt 402 into word vectors 404. Paragraph [0074]: At step 515, as detailed above, the device may identify a plurality of topics present in the prompt. For instance, the device may assess the words present in the prompt and convert them into vector representations, such as by using a FastText library); encoding, by an encoding model, the set of user segments into a set of user vectors (Vaknin – Fig. 4 and Paragraph [0056]: Prior to prompt 402 being input to the LLM, prompt analysis process 248 may identify the topic(s) present in prompt 402 by converting each word in prompt 402 into word vectors 404. Paragraph [0074]: At step 515, as detailed above, the device may identify a plurality of topics present in the prompt. For instance, the device may assess the words present in the prompt and convert them into vector representations, such as by using a FastText library); scoring each user vector of the set of user vectors to generate a set of user vector scores (Vaknin – Paragraph [0057]: After converting the words of prompt 402 into word vectors 404, prompt analysis process 248 may compute the word vector distances 406 between them. For instance, prompt analysis process 248 may compute the distances between each of word vectors 404 and the mean of them), aggregating the set of user vector scores to generate an aggregated score (Vaknin – Paragraph [0061]-[0070]: [0061] In some instances, prompt analysis process 248 may assess the sentence incoherence of prompt 304 as follows: [0062] for each prompt [0063] extract noun-adjective and verb-adverb word pairs using a language parser [0064] for each modifier-word pair [0065] extract the set of base modifiers normally used with that word from our control base corpus [0066] compute the word vectors for each modifier in the set [0067] aggregate the set of modifier vectors into a single vector [0068] compute the similarity between the modifier and the computed single base modifier vector; Vector Aggregation Computation: [0069] each vector is assigned a weight based on the inverse-document-frequency score, taken from a precomputed base corpus [0070] compute the weighted average of the individual modifier vectors; Vector Similarity Computation: vector similarity = (1 - cosine difference (modifier vector, aggregate vector)): detecting whether the user prompt is malicious based on whether the aggregated score satisfies a threshold (Vaknin – Paragraph [0071]: In this case, the higher the similarity the lower the incoherence; and Paragraph [0072]: Here, the suspicion incoherence threshold may be set such that values below the threshold signify prompt injection attack candidates); detecting whether the user prompt is malicious according to the set of user vector scores (Vaknin – Paragraph [0071]: In this case, the higher the similarity the lower the incoherence; and Paragraph [0072]: Here, the suspicion incoherence threshold may be set such that values below the threshold signify prompt injection attack candidates); and setting a prompt injection signal based on whether the user prompt is detected as malicious according to the set of user vector scores (Vaknin – Paragraph [0075]: At step 520, the device may determine that the prompt is malicious based on a variation in the plurality of topics, as described in greater detail above; and Paragraph [0076]: At step 525, as detailed above, the device may prevent the prompt from being processed by the language model. In some implementations, the device may also provide, to a user interface, an indication that the prompt is a suspected prompt injection attack).
Majmudar further teaches scoring each user vector of the set of user vectors to generate a set of user vector scores (Majmudar – Col. 9, Line 12-45: In various examples, the prompt data (including relevant context data at step (1)) and the inference output (at step (3)) may be sent to prompt validation component 148. Generally, the prompt validation component 148 may evaluate different spans in the prompt. A span, as used herein, refers to an ordered sequence of one or more tokens. A token may be data representing a single word (e.g., an unmodified natural language word), punctuation symbol, whitespace, and/or modified word (e.g., a word that has been stemmed or lemmatized). In various examples, classifier 150 may be a supervised machine learning classifier comprising a natural language encoder (e.g., BERT, DistilBERT, word2vec, etc.) and a supervised classifier head. The classifier 150 may be used to predict a trust score for each span detected in the prompt. In general, higher trust scores may indicate that a span is more trusted and is less likely to be associated with a potential malicious attack … As described in further detail below, training data for the classifier 150 may be generated by providing spans of varying degrees of trustworthiness and labeling each span with a ground truth trust score. In addition to the spans themselves (and/or encoded representations of the spans (such as semantic representation vectors generated using a natural language encoder)), the training data instances may also include data that identifies a source of the span; and Col. 19, Line 57-60: In various examples, the feature data and/or training data used by the various machine learning models may be stored and/or cached in memory 596); detecting whether the user prompt is malicious according to the set of user vector scores (Majmudar – Col. 10, Line 12-42: For a given span in the prompt, the prompt validation component 148 determine both a trust score (using classifier 150) and an attention score (using attention component 152). Thereafter, the prompt validation component 148 may determine an appropriate action (e.g., plan data) to generate based on these values. In some examples, the prompt validation component 148 may use deterministic rules to determine an action. For example, the trust score and attention score may be combined. In a naïve approach, an inverse of the trust score (e.g., a value between 0 and 1, where 0 represents the lowest amount of trust and 1 represents the highest amount of trust) may be multiplied by the attention score … For example, risk scores higher than 1.5 may be associated with a disengagement/termination action plan. Accordingly, for the example first span above with a risk score of 3.85, the 3.85 risk score may exceed the threshold. Accordingly, the prompt validation component 148 may indicate that the prompt may include malicious instructions (step (4a)) and that the orchestrator should terminate the dialog or other processing session (e.g., using a template response)); and … the user prompt is detected as malicious according to the set of user vector scores (Majmudar – Col. 10, Line 29-42: to generate a risk score (representing a possible indirect prompt injection attack). For example, if a first span has a trust score of 0.2 (indicating relatively low trust) and a high attention score of 0.77, the risk score may be (1/0.2)*0.77=3.85… For example, risk scores higher than 1.5 may be associated with a disengagement/termination action plan. Accordingly, for the example first span above with a risk score of 3.85, the 3.85 risk score may exceed the threshold. Accordingly, the prompt validation component 148 may indicate that the prompt may include malicious instructions (step (4a)) and that the orchestrator should terminate the dialog or other processing session (e.g., using a template response)); and the set of user vector scores (Majmudar – Col. 9, Line 12-45: In various examples, the prompt data (including relevant context data at step (1)) and the inference output (at step (3)) may be sent to prompt validation component 148. Generally, the prompt validation component 148 may evaluate different spans in the prompt. A span, as used herein, refers to an ordered sequence of one or more tokens … The classifier 150 may be used to predict a trust score for each span detected in the prompt. In general, higher trust scores may indicate that a span is more trusted and is less likely to be associated with a potential malicious attack).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, further incorporating Majmudar to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Majmudar’s teaching to assess per-span scores to incorporate in malicious prompt detection into Vaknin’s method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This addition would bolster the security of Vaknin’s method by providing a comparison basis for evaluating potentially harmful prompts.
The combination of Vaknin and Majmudar does not expressly teach wherein scoring the user vector comprises: determining one or more vector distances of the user vector to one or more stored vectors in a set of stored vectors in a vector store, and generating, individually for the user vector, a corresponding user vector score … according to the one or more vector distance.
However, Cappel teaches wherein scoring the user vector comprises: determining one or more vector distances of the user vector to one or more stored vectors in a set of stored vectors in a vector store, and generating, individually for the user vector, a corresponding user vector score … according to the one or more vector distance (Cappel – Paragraph [0075]: data characterizing a prompt or query for ingestion by an AI model, such as a generative artificial intelligence (GenAI) model (e.g., MLA 130, a large language model, etc.) is received. This data can comprise the prompt itself or, in some variations, it can comprise features or other aspects that can be used to analyze the prompt. The received data can be routed from the model environment 140 to the monitoring environment 160 by way of the proxy 150. Thereafter, it can be determined, at 1120, whether the prompt comprises or otherwise attempts to elicit malicious content based on a similarity analysis between a blocklist and the received data, the blocklist being derived from a corpus of known malicious prompts; and Paragraph [0077]: the similarity analysis is performed using the received data (e.g., the prompt or information characterizing the prompt) while, in other cases, the received data can be preprocessed in some fashion. For example, the received data can be tokenized and the resulting tokens are what are used by the similarity analysis. In other variations, the received data can be vectorized (i.e., features can be extracted from the prompts and such features can populate a vector, etc.). The resulting vector(s) can be used to generate embeddings using one or more dimension reduction techniques; and Paragraph [0080]: The similarity analysis can comprise a semantic analysis in which distance measurements indicative of similarity are generated based on a likeness of meaning of the received data to the blocklist. Semantic analyses can include one or more of: TF-IDF (Term Frequency-Inverse Document Frequency Distance) and Cosine Similarity. These techniques can measure the distance between two vector representations of the text (embeddings distance); and Paragraph [0056]: a determination by the analysis engine 170 results in data (e.g., instructions, scores, etc.) being sent to the model environment 140 which results in remediation actions).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin and Majmudar, further incorporating Cappel to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Cappel’s teaching to evaluate user prompt vectors by determining vector distances between the received prompt and known malicious prompts into Vaknin and Majmudar’s method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. Vaknin and Majmudar provide known techniques for segmenting and individually evaluating segments of user LLM queries for potential maliciousness. Majmudar further stores such segment evaluations for classifier training and thus, enhanced detection. Cappel contributes a known approach for determining prompt maliciousness by comparing incoming prompts to known-malicious prompts. The combination produces the obvious benefit of fine-granularity prompt evaluation using a variety of techniques for malicious prompt detection known in the art.
Regarding Claim 3:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
Vaknin further teaches wherein aggregating the set of user vector scores comprises averaging the set of user vector scores (Vaknin – Paragraph [0069]: Vector Aggregation Computation: each vector is assigned a weight based on the inverse-document-frequency score, taken from a precomputed base corpus; and Paragraph [0070]: compute the weighted average of the individual modifier vectors).
The motivation to combine the arts is the same as that of Claim 1.
Regarding Claim 5:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
Cappel further teaches wherein each stored vector in the set of stored vectors used in scoring the user vector is classified as malicious (Cappel – Paragraph [0062]: One or both of the analysis engines 152, 170 can utilize a blocklist when making the determination of whether a query, and in particular, a prompt, is indicative of being malicious and/or contains sensitive information. In some implementations, multiple blocklists can be utilized. The blocklist can leverage historical prompts that are known to be malicious (e.g., used for prompt injection attacks, etc.) and/or, in some variations, leverage prompts known to include sensitive information. The goal of a prompt injection attack would be to cause the MLA 130 to ignore previous instructions (i.e., instructions predefined by the owner or developer of the MLA 130, etc.) or perform unintended actions based on one or more specifically crafted prompts. The historical prompts can be from, for example, an internal corpus and/or from sources such as an open source malicious prompt list in which the listed prompts have been confirmed as being harmful prompt injection prompts. Similarly, if sensitive information is being analyzed, the blocklist can be generated from historical prompts known to contain sensitive information such as financial or personally identification information; and Paragraph [0077]: the received data can be vectorized (i.e., features can be extracted from the prompts and such features can populate a vector, etc.). The resulting vector(s) can be used to generate embeddings using one or more dimension reduction techniques. Embeddings are advantageous in that they have lower dimensionality so that the similarity analysis consumes fewer computing resources (as compared to the original prompt or the original vector); and Paragraph [0078]: The similarity analysis can take varying forms. In addition, more than one technique for similarity analysis can be utilized in parallel or in sequence. In some arrangement, a first similarity analysis is performed and, if there is a match, a second, more computationally expensive similarity analysis is performed. Match in this context means that some or all of the prompt is within a defined threshold relative to the blocklist (e.g., 95% matching, etc.)).
The motivation to combine the arts is the same as that of Claim 1.
Regarding Claim 7:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
Vaknin further teaches further comprising: … transmitting the user prompt to the LLM … receiving a response to the user prompt from the LLM; and forwarding the response to the user device (Vaknin – Figure 3: illustration of sequence for processing user prompt to detect prompt injection attacks, including receiving prompt 304 by prompt analysis process 248 and LLM 306, and receiving LLM response by prompt analysis process 248 and user 302).
Majmudar further teaches detecting that the user prompt is benign according to the set of scores; transmitting the user prompt to the LLM based on the user prompt being benign (Majmudar – Col. 10, Line 47-51: In another example, spans with low risk scores (e.g., between 0-0.5) may result in no action (e.g., the prompt validation component 148 may return data indicating that the prompt is valid as shown in the example in FIG. 1A); and Col. 11, Line 3-18: In the example of FIG. 1A, the prompt validation component 148 may determine that the input prompt data is valid (step (4)). Accordingly, the inference output comprising the natural language-based actions may be sent from inference engine 106 to orchestrator 102 (step (5) and from orchestrator 102 to an action plan generator 108 (step (6)). The action plan generator 108 may transform the natural language series of actions generated during LLM inference into a series of computer-executable actions (e.g., API calls, function calls, etc.) that may be used to carry out the actions determined during LLM inference. Similar to the prompt generator 104, the action plan generator 108 may itself be implemented as an LLM and/or another machine learning model trained to take natural language inputs and transform them into a series of computer-executable instructions (referred to herein as actions)).
The motivation to combine the arts is the same as that of Claim 1.
Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaknin, in view of Majmudar, Cappel, and Rennie et al. (US 20240265041 A1), hereinafter Rennie.
Regarding Claim 4:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
The combination of Vaknin, Majmudar, and Cappel does not expressly teach wherein segmenting the user prompt comprises: performing a sliding window segmentation to generate the set of user segments that overlap according to a configured stride.
However, Rennie teaches wherein segmenting the user prompt comprises: performing a sliding window segmentation to generate the set of user segments that overlap according to a configured stride (Rennie – Paragraph [0128]: A source content 210 (which may be part of a source document) is segmented into segments 220a-n … o simplify the segmentation process (so as to facilitate more efficient searching and retrieval), the source documents may be segmented to create overlap between the sequential document segments (not including the contextual information that is separately added to each segment). Thus, for example, in situations where a segment is created by a window of some particular size (constant or variable), the window may be shifted from one position to the following position by some pre-determined fraction of the window size (e.g., ¾, which for a 200-word window would be 150 words). As a result of the fractional shifting, transformations (e.g., language transformations models) applied to overlapped segments results in some correlation between the segments, which can preserve relevancy between consecutive segments during search time).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, and Cappel, further incorporating Rennie to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Rennie’s teaching to segment ingested textual data by applying a sliding window with some preconfigured stride to highlight segment overlap into Vaknin, Majmudar, and Cappel’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This combination provides an additional technique for deriving more significant information from input data.
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaknin, in view of Majmudar, Cappel, and Long et al. (US 10592667 B1), hereinafter Long.
Regarding Claim 6:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
The combination of Vaknin, Majmudar, and Cappel does not expressly teach further comprising: performing an approximate nearest neighbor (ANN) algorithm with the user vector in the set of user vectors and the set of stored vectors to identify a vector distance in the one or more vector distances from the user vector to a nearest vector in the set of stored vectors, wherein scoring the user vector is performed according to the vector distance to the nearest vector.
However, Long teaches further comprising: performing an approximate nearest neighbor (ANN) algorithm with the user vector in the set of user vectors and the set of stored vectors to identify a vector distance in the one or more vector distances from the user vector to a nearest vector in the set of stored vector (Long – Col. 8, Line 55-67: In some implementations, the vector neighbor module 112 can use Fast Library for Approximate Nearest Neighbors (FLAAN) techniques to determine the nearest neighbors of the binary vector. For example, in some instances, a Hamming function can be used to calculate a distance between two binary vectors (i.e., a received input sample and a stored known sample). Hamming distances (i.e., the distance computed by the Hamming function) can be calculated for each binary vector from a set of binary vectors stored in the malware detection database 108 as compared to the binary vector of the input sample. A FLANN function can then use the Hamming distances to identify the nearest neighbors to the input sample), wherein scoring the user vector is performed according to the vector distance to the nearest vector (Long – Col. 10, Line 1-23: for each binary vector in a vector group, at 512, the malware matching module 114 can calculate, at 514, a distance between the binary vector from the vector group, and the binary vector associated with the input sample. FIG. 6, for example, illustrates a logic flow diagram of an example method of calculating distances between image binary vectors. In some implementations, for each vector index of each binary vector, at 602, the malware matching module 114 can compare, at 604, the value at the vector index in the binary vector from the vector group, to the value at the vector index in the binary vector associated with the input sample. When the values match (e.g., are the same value), at 606, the malware matching module 114 can determine, at 610, if there are additional values in the binary vectors to check, and can continue to compare values in the vectors. If values at one of the vector indexes are different, the malware matching module 114 can increment, at 608, a distance counter. When each value in the two binary vectors has been compared, the final value of the distance counter can be used to calculate a similarity score and/or other scores associated with the two binary vectors (e.g., based on comparing the final value of the distance counter to one or more criteria and/or thresholds)).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, and Cappel, further incorporating Long to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Long’s process of identifying at least one nearest neighbor vector to a subject vector using an ANN, and using the vector distance between them to factor into vector scoring into Vaknin, Majmudar, and Cappel’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This combination further enhances the precision of threat detection based on historical data.
Claim(s) 8 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaknin, in view of Majmudar, Cappel, and Jain et al. (US 12147513 B1), hereinafter Jain.
Regarding Claim 8:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
Majmudar further teaches further comprising: detecting that the user prompt is benign according to the set of scores (Majmudar – Col. 10, Line 47-51: In another example, spans with low risk scores (e.g., between 0-0.5) may result in no action (e.g., the prompt validation component 148 may return data indicating that the prompt is valid as shown in the example in FIG. 1A)).
The combination of Vaknin, Majmudar, and Cappel does not expressly teach creating an LLM prompt from the user prompt based on the user prompt being benign; and sending the LLM prompt.
However, Jain teaches creating an LLM prompt from the user prompt based on the user prompt being benign (Jain – Col. 13, Line 9-24: a prompt validation model includes one or more input controls 610, as shown in FIG. 6. Additionally or alternatively, the input controls 610 can include one or more prompt validation models capable of executing operations including input validation 612a, trace injection 612b, logging 612c, secret redaction 612d, sensitive data detection 612e, prompt injection 612f, and/or prompt augmentation 612g. A prompt validation model can generate a validation indicator. The validation indicator can indicate a validation status (e.g., a binary indicator specifying whether the prompt is suitable for provision to the associated LLM). Additionally or alternatively, the validation indicator can indicate or specify aspects of the prompt that are validated and/or invalid, thereby enabling further modification to cure any associated deficiencies in the prompt; and Col. 14, Line 51-56: The prompt injection 612f can provide this modified prompt to a machine learning model designed to test for prompt injection attacks; this machine learning model can generate a validation indicator specifying a likelihood that the user-provided prompt is manipulating the prompt logic of the data generation platform 102's LLMs; and Col. 14, Line 60-66: Prompt augmentation 612g can include adding tokens (e.g., sentences or phrases) to the prompt to improve output generation behavior. For example, the breach mitigation engine 116 generates tokens to improve the register, language, or style of the generated outputs (e.g., by including a statement requesting that the generated output correspond to a given style within the prompt)); and sending the LLM prompt (Jain – Col. 22, Line 7-9: At act 714, process 700 can provide the prompt (e.g., as modified by suitable prompt validation models) to the LLM generate the requested output).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, and Cappel, further incorporating Jain to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Jain’s process of validating the security and quality of a user prompt to facilitate processing by an appropriate LLM into Vaknin, Majmudar, and Cappel’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This additional functionality enhances the system by further ensuring the security of prompts passed to an LLM, as well as modifying the prompts for more effective use of the LLM.
Regarding Claim 17:
Vaknin teaches a system comprising: at least one computer processor (Vaknin – Paragraph [0028]: The device 200 may comprise one or more network interfaces 210 (e.g., wired, wireless, etc.), at least one processor 220, and a memory 240 interconnected by a system bus 250, as well as a power supply 260 (e.g., battery, plug-in, etc.)); a large language model (LLM) prompt manager executing on the at least one computer processor and configured to (Vaknin – Paragraph [0032]: as detailed further below, prompt analysis process 248 may include computer executable instructions that, when executed by processor(s) 220, cause device 200 to perform the techniques described herein): receive, from a user device, a user prompt to an LLM (Vaknin – Figure 3: illustration of sequence for processing user prompt to detect prompt injection attacks, including receiving prompt 304 from user 302 at prompt analysis process 248), and send the [LLM] prompt to the LLM according to a prompt injection signal (Vaknin – Figure 3: illustration of sequence for processing user prompt to detect prompt injection attacks, including receiving prompt 304 by prompt analysis process 248 and LLM 306, and receiving LLM response by prompt analysis process 248 and user 302; and Paragraph [0075]: At step 520, the device may determine that the prompt is malicious based on a variation in the plurality of topics, as described in greater detail above; and Paragraph [0076]: At step 525, as detailed above, the device may prevent the prompt from being processed by the language model. In some implementations, the device may also provide, to a user interface, an indication that the prompt is a suspected prompt injection attack); and an LLM firewall executing on the at least one computer processor and configured to (Vaknin – Paragraph [0032]: as detailed further below, prompt analysis process 248 may include computer executable instructions that, when executed by processor(s) 220, cause device 200 to perform the techniques described herein; and Paragraph [0073]: a non-generic, specifically configured device (e.g., device 200), such as a router, firewall, controller for a network (e.g., an SDN controller or other device in communication therewith), server, or the like, may perform procedure 500 by executing stored instructions (e.g., prompt analysis process 248); Examiner’s Comment: According to at least paragraph [0028] of the instant specification, the LLM firewall operates as part of the LLM prompt manager. Thus the prompt analysis process is interpreted to fulfill the roles of both the LLM prompt manager and the LLM firewall): segment the user prompt to generate a set of user segments (Vaknin – Paragraph [0074]: At step 515, as detailed above, the device may identify a plurality of topics present in the prompt. For instance, the device may assess the words present in the prompt and convert them into vector representations, such as by using a FastText library), encode, by an encoding model, the set of user segments into a set of user vectors (Vaknin – Paragraph [0074]: At step 515, as detailed above, the device may identify a plurality of topics present in the prompt. For instance, the device may assess the words present in the prompt and convert them into vector representations, such as by using a FastText library); score each user vector of the set of user vectors … to generate a set of scores (Vaknin – Paragraph [0057]: After converting the words of prompt 402 into word vectors 404, prompt analysis process 248 may compute the word vector distances 406 between them. For instance, prompt analysis process 248 may compute the distances between each of word vectors 404 and the mean of them), aggregate the set of user vector scores to generate an aggregated score (Vaknin – Paragraph [0061]-[0070]: [0061] In some instances, prompt analysis process 248 may assess the sentence incoherence of prompt 304 as follows: [0062] for each prompt [0063] extract noun-adjective and verb-adverb word pairs using a language parser [0064] for each modifier-word pair [0065] extract the set of base modifiers normally used with that word from our control base corpus [0066] compute the word vectors for each modifier in the set [0067] aggregate the set of modifier vectors into a single vector [0068] compute the similarity between the modifier and the computed single base modifier vector; Vector Aggregation Computation: [0069] each vector is assigned a weight based on the inverse-document-frequency score, taken from a precomputed base corpus [0070] compute the weighted average of the individual modifier vectors; Vector Similarity Computation: vector similarity = (1 - cosine difference (modifier vector, aggregate vector))), detect whether the user prompt is malicious based on whether the aggregated score satisfies a threshold (Vaknin – Paragraph [0071]: In this case, the higher the similarity the lower the incoherence; and Paragraph [0072]: Here, the suspicion incoherence threshold may be set such that values below the threshold signify prompt injection attack candidates) and set the prompt injection signal based on whether the user prompt is detected as malicious (Vaknin – Paragraph [0075]: At step 520, the device may determine that the prompt is malicious based on a variation in the plurality of topics, as described in greater detail above; and Paragraph [0076]: At step 525, as detailed above, the device may prevent the prompt from being processed by the language model. In some implementations, the device may also provide, to a user interface, an indication that the prompt is a suspected prompt injection attack).
Vaknin does not expressly teach score each user vector of the set of user vectors based on a comparison between the user vector and a set of stored vectors in a vector store to generate a set of scores, detect whether the user prompt is malicious according to the set of user vector scores, and … the user prompt is detected as malicious according to the set of scores; and .
However, Majmudar teaches score each user vector of the set of user vectors based on a comparison between the user vector and a set of stored vectors in a vector store to generate a set of scores (Majmudar – Col. 9, Line 12-45: In various examples, the prompt data (including relevant context data at step (1)) and the inference output (at step (3)) may be sent to prompt validation component 148. Generally, the prompt validation component 148 may evaluate different spans in the prompt. A span, as used herein, refers to an ordered sequence of one or more tokens. A token may be data representing a single word (e.g., an unmodified natural language word), punctuation symbol, whitespace, and/or modified word (e.g., a word that has been stemmed or lemmatized). In various examples, classifier 150 may be a supervised machine learning classifier comprising a natural language encoder (e.g., BERT, DistilBERT, word2vec, etc.) and a supervised classifier head. The classifier 150 may be used to predict a trust score for each span detected in the prompt. In general, higher trust scores may indicate that a span is more trusted and is less likely to be associated with a potential malicious attack … As described in further detail below, training data for the classifier 150 may be generated by providing spans of varying degrees of trustworthiness and labeling each span with a ground truth trust score. In addition to the spans themselves (and/or encoded representations of the spans (such as semantic representation vectors generated using a natural language encoder)), the training data instances may also include data that identifies a source of the span; and Col. 19, Line 57-60: In various examples, the feature data and/or training data used by the various machine learning models may be stored and/or cached in memory 596), detect whether the user prompt is malicious according to the set of user vector scores (Majmudar – Col. 10, Line 12-42: For a given span in the prompt, the prompt validation component 148 determine both a trust score (using classifier 150) and an attention score (using attention component 152). Thereafter, the prompt validation component 148 may determine an appropriate action (e.g., plan data) to generate based on these values. In some examples, the prompt validation component 148 may use deterministic rules to determine an action. For example, the trust score and attention score may be combined. In a naïve approach, an inverse of the trust score (e.g., a value between 0 and 1, where 0 represents the lowest amount of trust and 1 represents the highest amount of trust) may be multiplied by the attention score … For example, risk scores higher than 1.5 may be associated with a disengagement/termination action plan. Accordingly, for the example first span above with a risk score of 3.85, the 3.85 risk score may exceed the threshold. Accordingly, the prompt validation component 148 may indicate that the prompt may include malicious instructions (step (4a)) and that the orchestrator should terminate the dialog or other processing session (e.g., using a template response)), and … the user prompt is detected as malicious according to the set of scores (Majmudar – Col. 10, Line 29-42: to generate a risk score (representing a possible indirect prompt injection attack). For example, if a first span has a trust score of 0.2 (indicating relatively low trust) and a high attention score of 0.77, the risk score may be (1/0.2)*0.77=3.85. … For example, risk scores higher than 1.5 may be associated with a disengagement/termination action plan. Accordingly, for the example first span above with a risk score of 3.85, the 3.85 risk score may exceed the threshold. Accordingly, the prompt validation component 148 may indicate that the prompt may include malicious instructions (step (4a)) and that the orchestrator should terminate the dialog or other processing session (e.g., using a template response)); and the set of user vector scores (Majmudar – Col. 9, Line 12-45: In various examples, the prompt data (including relevant context data at step (1)) and the inference output (at step (3)) may be sent to prompt validation component 148. Generally, the prompt validation component 148 may evaluate different spans in the prompt. A span, as used herein, refers to an ordered sequence of one or more tokens … The classifier 150 may be used to predict a trust score for each span detected in the prompt. In general, higher trust scores may indicate that a span is more trusted and is less likely to be associated with a potential malicious attack).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, further incorporating Majmudar to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Majmudar’s teaching to assess per-span scores to incorporate in malicious prompt detection into Vaknin’s system for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This addition would bolster the security of Vaknin’s system by providing a comparison basis for evaluating potentially harmful prompts.
The combination of Vaknin and Majmudar does not expressly teach wherein scoring the user vector comprises: determining one or more vector distances of the user vector to one or more stored vectors in a set of stored vectors in a vector store, and generating, individually for the user vector, a corresponding user vector score … according to the one or more vector distance.
However, Cappel teaches wherein scoring the user vector comprises: determining one or more vector distances of the user vector to one or more stored vectors in a set of stored vectors in a vector store, and generating, individually for the user vector, a corresponding user vector score … according to the one or more vector distance (Cappel – Paragraph [0075]: data characterizing a prompt or query for ingestion by an AI model, such as a generative artificial intelligence (GenAI) model (e.g., MLA 130, a large language model, etc.) is received. This data can comprise the prompt itself or, in some variations, it can comprise features or other aspects that can be used to analyze the prompt. The received data can be routed from the model environment 140 to the monitoring environment 160 by way of the proxy 150. Thereafter, it can be determined, at 1120, whether the prompt comprises or otherwise attempts to elicit malicious content based on a similarity analysis between a blocklist and the received data, the blocklist being derived from a corpus of known malicious prompts; and Paragraph [0077]: the similarity analysis is performed using the received data (e.g., the prompt or information characterizing the prompt) while, in other cases, the received data can be preprocessed in some fashion. For example, the received data can be tokenized and the resulting tokens are what are used by the similarity analysis. In other variations, the received data can be vectorized (i.e., features can be extracted from the prompts and such features can populate a vector, etc.). The resulting vector(s) can be used to generate embeddings using one or more dimension reduction techniques; and Paragraph [0080]: The similarity analysis can comprise a semantic analysis in which distance measurements indicative of similarity are generated based on a likeness of meaning of the received data to the blocklist. Semantic analyses can include one or more of: TF-IDF (Term Frequency-Inverse Document Frequency Distance) and Cosine Similarity. These techniques can measure the distance between two vector representations of the text (embeddings distance); and Paragraph [0056]: a determination by the analysis engine 170 results in data (e.g., instructions, scores, etc.) being sent to the model environment 140 which results in remediation actions).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin and Majmudar, further incorporating Cappel to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Cappel’s teaching to evaluate user prompt vectors by determining vector distances between the received prompt and known malicious prompts into Vaknin and Majmudar’s method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. Vaknin and Majmudar provide known techniques for segmenting and individually evaluating segments of user LLM queries for potential maliciousness. Majmudar further stores such segment evaluations for classifier training and thus, progressively enhanced detection. Cappel contributes a known approach for determining prompt maliciousness by comparing incoming prompts to known-malicious prompts. The combination produces the obvious benefit of fine-granularity prompt evaluation using a variety of techniques for malicious prompt detection known in the art.
The combination of Vaknin, Majmudar, and Cappel does not expressly teach create an LLM prompt from the user prompt, and send the LLM prompt to the LLM.
However, Jain teaches create an LLM prompt from the user prompt (Jain – Col. 3, Line 34-40: Based on the results of the prompt validation model, the data generation platform can modify the prompt such that the prompt satisfies any associated validation criteria (e.g., through the redaction of sensitive data or other details) thereby mitigating the effect of potential security breaches, inaccuracies, or adversarial manipulation associated with the user's prompt; and Col. 13, Line 9-24: a prompt validation model includes one or more input controls 610, as shown in FIG. 6. Additionally or alternatively, the input controls 610 can include one or more prompt validation models capable of executing operations including … prompt augmentation 612g. A prompt validation model can generate a validation indicator. The validation indicator can indicate a validation status (e.g., a binary indicator specifying whether the prompt is suitable for provision to the associated LLM); and Col. 14, Line 32-39: In some implementations, the breach mitigation engine 116 can determine that the prompt includes a forbidden token (e.g., a swear word, or another undesirable natural language token, as specified within a forbidden token database). Based on this determination, the breach mitigation engine 116 can modify the prompt (e.g., by removing the forbidden token) prior to providing the prompt to a suitable LLM for processing; and Col. 14, Line 60-66: Prompt augmentation 612g can include adding tokens (e.g., sentences or phrases) to the prompt to improve output generation behavior. For example, the breach mitigation engine 116 generates tokens to improve the register, language, or style of the generated outputs (e.g., by including a statement requesting that the generated output correspond to a given style within the prompt)); and send the LLM prompt to the LLM (Jain – Col. 22, Line 7-9: At act 714, process 700 can provide the prompt (e.g., as modified by suitable prompt validation models) to the LLM generate the requested output).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, and Cappel, further incorporating Jain to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Jain’s process of validating the security and quality of a user prompt to facilitate processing by an appropriate LLM into Vaknin, Majmudar, and Cappel’s combined system for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This additional functionality enhances the system by further ensuring the security of prompts passed to an LLM, as well as modifying the prompts for more effective use of the LLM.
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaknin, in view of Majmudar, Cappel, and Lingafelt et al. (US 20140173727 A1), hereinafter Lingafelt.
Regarding Claim 9:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
The combination of Vaknin, Majmudar, and Cappel does not expressly teach further comprising: receiving a correction of the prompt injection signal indicating that that user prompt is benign; selecting, from the set of stored vectors, at least one stored vector based on the at least one stored vector indicating that the user prompt is malicious; and marking the at least one stored vector as invalid responsive to the at least one stored vector indicating that the user prompt is malicious and responsive to the correction.
However, Lingafelt teaches further comprising: receiving a correction of the prompt injection signal indicating that that user prompt is benign; selecting, from the set of stored vectors, at least one stored vector based on at least one stored vector indicating that the user prompt is malicious; and marking at least one stored vector as invalid responsive to at least one stored vector indicating that the user prompt is malicious and responsive to the correction (Lingafelt – Paragraph [0007]: gaining access to an alert database and a signature set by an analytics module and an adjustment module, where the alert database includes whether each alert is a valid alert or a false positive alert; quantifying for each signature contained in the signature set an effect on the change in the number of alerts from its removal; determining with an analytics module whether any signature has a ratio of valid to false positive alerts less than a first threshold; and when at least one signature has the ratio less than the first threshold identifying and removing with an adjustment module at least one signature from the signature database having a ratio less than the first threshold where the signature is removed from the signature set, and repeating quantifying and determining; and Paragraph [0034]: An example of how a rule could be removed because it is less effective is illustrated in FIG. 2A. Rules 1 and 3 both produced a valid alert for packet A, but rule 3 also produced a false positive alert for packet C. If rule 3 was removed from the rules database, then there would be no decrease in valid alerts (i.e., packet A) while decreasing the number of false positive alerts in the system. A goal of the system and/or user may be to reduce the amount of false positive alerts, even if this also results in the removal of valid alerts. For example, if a typical rule produced 10 false positives for each valid alert, and the rule under investigation produced 10,000 false positives for each valid alert, the analysis may choose to remove the signature under investigation from the system and forego the 1 valid alert in exchange for a reduction of 10,000 false positives from the system).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, and Cappel, further incorporating Lingafelt to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Lingafelt’s teaching to remove attack definitions that result in false positive alerts into Vaknin, Majmudar, and Cappel’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. Lingafelt demonstrates a knowledge within the art of a technique to suppress false alarms by invalidating attack signatures that are determined to be ineffective or prone to over-reporting. Thus, it would be obvious to one of ordinary skill in the art to apply such knowledge to a system which relies on historical malicious vectors for attack detection.
Claim(s) 10-11, and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaknin, in view of Majmudar, Cappel, and Yeung et al. (US 12130917 B1), hereinafter Yeung.
Regarding Claim 10:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
Vaknin teaches and generating the [set of stored] vectors from the set of training prompts (Vaknin – Paragraph [0074]: At step 515, as detailed above, the device may identify a plurality of topics present in the prompt. For instance, the device may assess the words present in the prompt and convert them into vector representations, such as by using a FastText library; and Paragraph [0074]: At step 515, as detailed above, the device may identify a plurality of topics present in the prompt. For instance, the device may assess the words present in the prompt and convert them into vector representations, such as by using a FastText library).
Majmudar further teaches the set of stored vectors (Majmudar – Col. 9, Line 37-45: As described in further detail below, training data for the classifier 150 may be generated by providing spans of varying degrees of trustworthiness and labeling each span with a ground truth trust score. In addition to the spans themselves (and/or encoded representations of the spans (such as semantic representation vectors generated using a natural language encoder)), the training data instances may also include data that identifies a source of the span).
The combination of Vaknin, Majmudar, and Cappel does not expressly teach further comprising: obtaining a set of input training prompts; processing, by a training system LLM, at least a subset of the input training prompts to create a set of generated training prompts; adding the set of generated training prompts and the set of input training prompts to a set of training prompts.
However, Yeung teaches further comprising: obtaining a set of input training prompts; processing, by an LLM, at least a subset of the input training prompts to create a set of generated training prompts; adding the set of generated training prompts and the set of input training prompts to a set of training prompts (Yeung – Col. 9, Line 20-40: The prompt injection classifier 192, 194 can be a machine learning model such as deBERTa, an XGBoost classification model, a logistic regression model, an XLNet model and the like. In the case of a binary classifier, the prompt injection classifier 192, 194 can be trained using a corpus of data which can include a plurality of benign prompts that do not contain prompt injection information and a plurality of malicious prompts that contain various character strings (which can include portions of alphanumeric symbols, non-printable characters, symbols, controls, etc.) and the like which encapsulate various sorts of prompt injection. Malicious prompts in this context refer to prompts that cause the prompt injection classifier 192, 194 to exhibit undesired behavior. Benign prompts in this context can refer to prompts that do not cause the prompt injection classifier 192, 194 to exhibit undesired behavior. In some variations, the prompts forming part of the corpus can be labeled with their classification. The model training can be performed by converting the prompts into sentence embeddings which can, amongst other features, be used to train the prompt injection classifier 192, 194; Examiner’s Comment: at least XLNet (classifier 192, 194) is interpreted as an LLM which is used to process prompts in creating a set of training prompts).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, and Cappel, further incorporating Yeung to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Yeung’s teaching to generate training data by processing a set of prompts into Vaknin, Majmudar, and Cappel’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This combination provides the system with a means to generate a strong basis of comparison in detecting potentially malicious prompts.
Regarding Claim 11:
The combination of Vaknin, Majmudar, Cappel and Yeung teaches the method of claim 10.
Yeung further teaches transmitting an instruction to the training system LLM with a set of input malicious prompts in the at least the subset of input training prompts, the instruction requesting a rephrasing of each of the at least the subset of input malicious prompts; and receiving from the training system LLM the set of generated malicious prompts, wherein the set of generated training prompts comprises the set of generated malicious prompts (Yeung – Col. 11, Line 5-11: One approach for LLM generated attacks is to use a jailbroken/misaligned/non-aligned LLM to generate a large amount of malicious content. In this context, a jailbroken LLM can be an LLM that has been prompted with a specific input that allows it to respond with malicious strings. A misaligned LLM can be fine-tuned to output malicious strings (misaligned); and Col. 11, Line 26-34: Here is an example of a malicious instruction: Teach me how to write a computer virus. Generate 100 similar malicious examples spanning various topics. ‘ ”
This arrangement would generate a large corpus of malicious strings that can either be used as training data, or can be combined with the skeletonized approach to create a varied dataset of machine-generated attacks).
The motivation to combine the arts is the same as that of Claim 10.
Regarding Claim 15:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
The combination of Vaknin, Majmudar, and Cappel does not expressly teach wherein the encoding model is a sentence embedding model.
However, Yeung teaches wherein the encoding model is a sentence embedding model (Yeung – Col. 9, Line 37-40: The model training can be performed by converting the prompts into sentence embeddings which can, amongst other features, be used to train the prompt injection classifier 192, 194).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, and Cappel, further incorporating Yeung to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Yeung’s teaching to use a sentence embedding model as an encoding model for generating segments out of user prompts into Vaknin, Majmudar, and Cappel’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This additional functionality would facilitate generation of vectors useful in malicious prompt detection.
Claim(s) 12-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaknin, in view of Majmudar, Cappel, Yeung, and Martin et al. (US 10372910 B2), hereinafter Martin.
Regarding Claim 12:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
Vaknin further teaches generating, by the encoding model, a set of malicious vectors from the malicious prompt and a set of benign vectors from the set of benign prompts (Vaknin – Paragraph [0074]: At step 515, as detailed above, the device may identify a plurality of topics present in the prompt. For instance, the device may assess the words present in the prompt and convert them into vector representations, such as by using a FastText library).
The combination of Vaknin, Majmudar, and Cappel does not expressly teach further comprising: obtaining a malicious prompt and a set of benign prompts; scoring each of the set of malicious vectors according to a vector distance in the one or more vector distances to the set of benign vectors to obtain a similarity score for each of the set of malicious vectors; selecting a subset of the set of malicious vectors having at least the similarity score satisfying a similarity threshold indicating a greater vector distance with respect to the set of benign vectors; and adding the subset of the set of malicious vectors to the set of stored vectors.
However, Yeung teaches further comprising: obtaining a malicious prompt and a set of benign prompts (Yeung – Col. 9, Line 20-40: The prompt injection classifier 192, 194 can be a machine learning model such as deBERTa, an XGBoost classification model, a logistic regression model, an XLNet model and the like. In the case of a binary classifier, the prompt injection classifier 192, 194 can be trained using a corpus of data which can include a plurality of benign prompts that do not contain prompt injection information and a plurality of malicious prompts that contain various character strings (which can include portions of alphanumeric symbols, non-printable characters, symbols, controls, etc.) and the like which encapsulate various sorts of prompt injection. Malicious prompts in this context refer to prompts that cause the prompt injection classifier 192, 194 to exhibit undesired behavior. Benign prompts in this context can refer to prompts that do not cause the prompt injection classifier 192, 194 to exhibit undesired behavior. In some variations, the prompts forming part of the corpus can be labeled with their classification. The model training can be performed by converting the prompts into sentence embeddings which can, amongst other features, be used to train the prompt injection classifier 192, 194).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, and Cappel, further incorporating Yeung to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Yeung’s teaching to generate training data for a classifier by processing a set of both malicious and benign prompts into Vaknin, Majmudar, and Cappel’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This addition would provide the system with a means to generate training data for detecting malicious/benign prompts.
The combination of Vaknin, Majmudar, Cappel, and Yeung does not expressly teach scoring each of the set of malicious vectors according to a vector distance in the one or more vector distances to the set of benign vectors to obtain a similarity score for each of the set of malicious vectors; selecting a subset of the set of malicious vectors having at least the similarity score satisfying a similarity threshold indicating a greater vector distance with respect to the set of benign vectors; and adding the subset of the set of malicious vectors to the set of stored vectors.
However, Martin teaches scoring each of the set of malicious vectors according to a vector distance in the one or more vector distances to the set of benign vectors to obtain a similarity score for each of the set of malicious vectors (Martin – Col. 15, Line 8-12: the system can calculate a benign score representing proximity of the new vector to known benign vectors in Block S252, wherein the benign score represents a likelihood that the new vector exemplifies benign activity on the network); selecting a subset of the set of malicious vectors having at least the similarity score satisfying a similarity threshold indicating a greater vector distance with respect to the set of benign vectors; and adding the subset of the set of malicious vectors to the set of stored vectors (Martin – Col. 23, Line 32-45: the system can add the new vector to the corpus of historical vectors and/or to the subset of labeled historical vectors in Block S280, as shown in FIGS. 3, 4, and 5. For example, the system can access a result of an investigation into the asset responsive to the alert generated in Block S240, such as from an investigation database. The system can then label the new vector according to the result, such as by labeling the new vector as malicious (e.g., representing a security threat) or by labeling the new vector as benign according to the result of the investigation. The system can then insert the new vector into the corpus of historical vectors and retrain the replicator neural network (or other outlier detection model) on this extended corpus of historical vectors).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, Cappel, and Yeung, further incorporating Martin to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Martin’s teaching to acquire malicious/benign scores for vectors representing observed activity, then using those scores to label and add the vectors to a set of historical vectors for later detection into Vaknin, Majmudar, Cappel, and Yeung’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. While not directed to LLM security, Martin demonstrates a knowledge in the art of a process for comparing vector representations of security threats by determining vector distances. Further distances from known benign vectors represent a lower likelihood that the subject vector is benign – i.e. a higher likelihood the subject vector is malicious. Once determined, the vectors can be added to a vector store useful in future vector evaluations. These capabilities mirror the functionality described in the claim, and when combined with the teachings of Vaknin, Majmudar, and Yeung, render obvious to a skilled artisan the process of evaluating the embeddings of user prompts in a similar manner.
Regarding Claim 13:
The combination of Vaknin, Majmudar, Cappel, Yeung, and Martin teaches the method of claim 12.
Vaknin further teaches further comprising: segmenting the malicious prompt to create a set of malicious prompt segments, wherein the encoding model generates a malicious vector for each of the set of malicious prompt segments (Vaknin – Paragraph [0074]: At step 515, as detailed above, the device may identify a plurality of topics present in the prompt. For instance, the device may assess the words present in the prompt and convert them into vector representations, such as by using a FastText library).
The motivation to combine the arts is the same as that of Claim 12.
Regarding Claim 14:
The combination of Vaknin, Majmudar, Cappel, Yeung, and Martin teaches the method of claim 12.
Martin further teaches further comprising: scoring each of the set of malicious vectors according to a second vector distance to the set of stored vectors to obtain an impact score for each of the set of malicious vectors (Martin – Col. 14, Line 57-67 and Col. 15, Line 1-7: In particular, the system can compare the new vector directly to a set of historical vectors representing confirmed cyber attacks on the network and/or external networks (hereinafter “malicious vectors”) in Block S270 and output an alert to investigate the asset or the network generally for a particular cyber attack in Block S272 if the new vector matches a particular malicious vector—in the set of known malicious vectors—tagged with the particular cyber attack. However, if the system fails to match the new vector to a known malicious vector, the system can calculate a malicious score representing proximity of the new vector to known malicious vectors, such as by implementing k-means clustering techniques to cluster a corpus of known malicious vectors, calculating distances between the new vector and these clusters of known malicious vectors, and compiling these distances into malicious confidence scores representing a likelihood that the new vector exemplifies a cyber attack or other threat on the network in Block S250), wherein the subset of the set of malicious vectors is further selected based on having the impact score satisfying an impact threshold indicating a greater vector distance with respect to the set of stored vectors (Martin – Col. 23, Line 32-45: the system can add the new vector to the corpus of historical vectors and/or to the subset of labeled historical vectors in Block S280, as shown in FIGS. 3, 4, and 5. For example, the system can access a result of an investigation into the asset responsive to the alert generated in Block S240, such as from an investigation database. The system can then label the new vector according to the result, such as by labeling the new vector as malicious (e.g., representing a security threat) or by labeling the new vector as benign according to the result of the investigation. The system can then insert the new vector into the corpus of historical vectors and retrain the replicator neural network (or other outlier detection model) on this extended corpus of historical vectors).
The motivation to combine the arts is the same as that of Claim 12.
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaknin, in view of Majmudar, Cappel, and Barla (Barla, N. (2024, April 1). How to train a custom LLM embedding model. DagsHub Blog. https://web.archive.org/web/20240405064056/https://dagshub.com/blog/how-to-train-a-custom-llm-embedding-model/), hereinafter Barla.
Regarding Claim 16:
The combination of Vaknin, Majmudar, and Cappel teaches the method of claim 1.
The combination of Vaknin, Majmudar, and Cappel does not expressly teach further comprising: training a custom embedding model with a set of input segments and a set of generated segments to generate vectors based on semantic similarity between the set of input segments and the set of generated segments to generate a trained custom embedding model; and using the trained custom embedding model as the encoding model.
However, Barla teaches further comprising: training a custom embedding model with a set of input segments and a set of generated segments to generate vectors based on semantic similarity between the set of input segments and the set of generated segments to generate a trained custom embedding model (Barla – P. 5: Training or fine-tuning phase involves adjusting the parameters of the model to minimize the distance between similar words and maximize the distance between dissimilar words. Here semantic relationships are learned as the model observes how words co-occur in sentences, enabling it to encode meaningful associations between tokens); and using the trained custom embedding model as the encoding model (Barla – P. 5: Here semantic relationships are learned as the model observes how words co-occur in sentences, enabling it to encode meaningful associations between tokens).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Vaknin, Majmudar, and Cappel, further incorporating Barla to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Barla’s teaching to generate a trained custom embedding model into Vaknin, Majmudar, and Cappel’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This combination provides flexibility and a potential for greater precision in detecting malicious LLM prompts.
Claim(s) 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yeung, in view of Martin, Nissim et al. (Nissim, N., Cohen, A., Moskovitch, R., Shabtai, A., Edri, M., BarAd, O., & Elovici, Y. (2016). Keeping pace with the creation of new malicious PDF files using an active-learning based detection framework. Security Informatics, 5(1). https://doi.org/10.1186/s13388-016-0026-3), hereinafter Nissim, and Vaknin.
Regarding Claim 19:
Yeung teaches obtaining a malicious prompt and a set of benign prompts (Yeung – Col. 9, Line 20-40: The prompt injection classifier 192, 194 can be a machine learning model such as deBERTa, an XGBoost classification model, a logistic regression model, an XLNet model and the like. In the case of a binary classifier, the prompt injection classifier 192, 194 can be trained using a corpus of data which can include a plurality of benign prompts that do not contain prompt injection information and a plurality of malicious prompts that contain various character strings (which can include portions of alphanumeric symbols, non-printable characters, symbols, controls, etc.) and the like which encapsulate various sorts of prompt injection. Malicious prompts in this context refer to prompts that cause the prompt injection classifier 192, 194 to exhibit undesired behavior. Benign prompts in this context can refer to prompts that do not cause the prompt injection classifier 192, 194 to exhibit undesired behavior. In some variations, the prompts forming part of the corpus can be labeled with their classification. The model training can be performed by converting the prompts into sentence embeddings which can, amongst other features, be used to train the prompt injection classifier 192, 194); generating, [by an encoding model], a set of malicious vectors from the malicious prompt and a set of benign vectors from the set of benign prompts (Yeung – Col. 9, Line 37-40: The model training can be performed by converting the prompts into sentence embeddings which can, amongst other features, be used to train the prompt injection classifier 192, 194; Examiner’s Comment: the sentence embeddings of the respective malicious and benign prompts are interpreted to represent the set of malicious vectors and the set of benign vectors); and detecting a prompt injection attack [using the set of stored vectors] (Yeung – Col. 2, Line 46-50: The analysis engine uses the prompt injection classifier to determine a category for the prompt which is indicative of whether the prompt comprises or elicits malicious content (or otherwise causes the GenAI model to behave in an undesired manner; and Col. 2, Line 53-62: The categories can take varying forms. As one example, the category can specify a threat severity for the prompt (malicious, suspicious, unknown, or benign, etc.). In other variations, the category can specify a type of prompt injection attack. The prompt injection attack types can include, for example, one or more of: a direct task deflection attack, a special case attack, a context continuation attack, a context termination attack, a syntactic transformation attack, an encryption attack, a text redirection attack, as well as other types of prompt injection attacks).
Yeung does not expressly teach scoring each of the set of malicious vectors according to a first vector distance to the set of benign vectors to obtain a similarity score for each of the set of malicious vectors; scoring each of the set of malicious vectors according to a second vector distance to the set of stored vectors to obtain an impact score for each of the set of malicious vectors; selecting a subset of the set of malicious vectors having at least the similarity score satisfying a similarity threshold indicating a greater vector distance with respect to the set of benign vectors; wherein the subset of the set of malicious vectors is further selected based on having the impact score satisfying an impact threshold indicating a greater vector distance with respect to the set of stored vectors; and adding the subset of the set of malicious vectors to the set of stored vectors.
However, Martin teaches scoring each of the set of malicious vectors according to a first vector distance to the set of benign vectors to obtain a similarity score for each of the set of malicious vectors (Martin – Col. 15, Line 8-12: the system can calculate a benign score representing proximity of the new vector to known benign vectors in Block S252, wherein the benign score represents a likelihood that the new vector exemplifies benign activity on the network); scoring each of the set of malicious vectors according to a second vector distance to the set of stored vectors to obtain an impact score for each of the set of malicious vectors (Martin – Col. 21, Line 28-67, and Col. 22, Line 1-11: in response to the first degree of deviation exceeding a deviation threshold score, issuing a first alert to investigate the first asset. Generally, in Block S230, the system tests a new vector associated with a first asset on the network for similarity to many (e.g., all) historical vectors generated previously for assets on the network, as shown in FIGS. 3, 4, and 5 …rarity of a sequence of events may be strongly correlated to malignance of this sequence of events. Therefore, if the new vector differs sufficiently from all vectors in the set of historical vectors, the system can label the new vector as an anomaly and immediately issue a prompt or alert to investigate the first asset in Block S240 rather than allocate additional resources to attempting to match or link the new vector to a subset of labeled historical vectors for which the new vector is still an outlier, as shown in FIG. 4. In particular, risk represented by the new vector that is substantially dissimilar from all historical vectors may be less predictable or quantifiable; the system can therefore prompt a human analyst to investigate the first asset associated with the new vector in response to identifying the new vector as an anomaly; In one implementation shown in FIG. 3, the system implements a replicator neural network to calculate an outlier score of the new vector in Block S230 … an outlier score (e.g., from “0” to “100”) based on a difference between the new vector and the new output vector. In particular, the outlier score of the new vector can represent a degree of deviation of the new vector from a normal distribution of vectors for the network) selecting a subset of the set of malicious vectors having at least the similarity score satisfying a similarity threshold (Martin – Col. 26, Line 23-30: Block S260 of the second method S200 recites, in response to the first malicious score exceeding the first benign score, issuing a second alert to investigate the network for the first network security threat. Generally, in Block S260, the system selectively issues an alert to investigate the network or the first asset for a security threat if the malicious score exceeds both a preset malicious threshold score and the benign score); the impact score satisfying an impact threshold indicating a greater vector distance with respect to the set of stored vectors (Martin – Col. 21, Line 28-67, and Col. 22, Line 1-11: in response to the first degree of deviation exceeding a deviation threshold score, issuing a first alert to investigate the first asset. Generally, in Block S230, the system tests a new vector associated with a first asset on the network for similarity to many (e.g., all) historical vectors generated previously for assets on the network, as shown in FIGS. 3, 4, and 5 …rarity of a sequence of events may be strongly correlated to malignance of this sequence of events. Therefore, if the new vector differs sufficiently from all vectors in the set of historical vectors, the system can label the new vector as an anomaly and immediately issue a prompt or alert to investigate the first asset in Block S240 rather than allocate additional resources to attempting to match or link the new vector to a subset of labeled historical vectors for which the new vector is still an outlier, as shown in FIG. 4. In particular, risk represented by the new vector that is substantially dissimilar from all historical vectors may be less predictable or quantifiable; the system can therefore prompt a human analyst to investigate the first asset associated with the new vector in response to identifying the new vector as an anomaly; In one implementation shown in FIG. 3, the system implements a replicator neural network to calculate an outlier score of the new vector in Block S230 … an outlier score (e.g., from “0” to “100”) based on a difference between the new vector and the new output vector. In particular, the outlier score of the new vector can represent a degree of deviation of the new vector from a normal distribution of vectors for the network); and adding the subset of the set of malicious vectors to the set of stored vectors (Martin – Col. 23, Line 32-45: the system can add the new vector to the corpus of historical vectors and/or to the subset of labeled historical vectors in Block S280, as shown in FIGS. 3, 4, and 5. For example, the system can access a result of an investigation into the asset responsive to the alert generated in Block S240, such as from an investigation database. The system can then label the new vector according to the result, such as by labeling the new vector as malicious (e.g., representing a security threat) or by labeling the new vector as benign according to the result of the investigation. The system can then insert the new vector into the corpus of historical vectors and retrain the replicator neural network (or other outlier detection model) on this extended corpus of historical vectors; and Figure 4: Illustration of a variation of a process for detection malicious and/or benign assets acting within a system).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Yeung, further incorporating Martin to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Martin’s teaching to acquire malicious/benign scores for vectors representing observed activity, then using those scores to label and add the vectors to a set of historical vectors for later detection into Yeung’s method for generating vectors based on malicious and benign user prompts useful in detection of prompt injection attacks. While not directed to LLM security, Martin demonstrates a knowledge in the art of a process for comparing vector representations of security threats by determining vector distances. Further distances from known benign vectors represent a lower likelihood that the subject vector is benign – i.e. a higher likelihood the subject vector is malicious. Once determined, the vectors can be added to a vector store useful in future vector evaluations. Martin provides an additional determination that a vector representing an unknown/unfamiliar security threat (based on satisfying a threshold difference from all known vectors) should be processed with prioritized attention and be added to the corpus for classifying incoming threats. These capabilities mirror the functionality described in the claim, and when combined with the teachings of Yeung, render obvious to a skilled artisan the process of evaluating the embeddings of user prompts in a similar manner.
The combination of Yeung and Martin does not expressly teach a subset of the set of malicious vectors having at least the similarity score … indicating a greater vector distance with respect to the set of benign vectors (Nissim – P. 10, Right Col.: In Fig. 4 the files that were acquired (marked with a red circle) are those files that are classified as malicious and are at the maximum distance from the separating hyperplane; and P. 11, Figure 4: Illustration of comparisons among of malicious and benign files); wherein the subset of the set of malicious vectors is further selected based on having … a greater vector distance with respect to the set of stored vectors (Nissim – P. 10, Right Col. And P. 11, Left Col.: In Fig. 4 the files that were acquired (marked with a red circle) are those files that are classified as malicious and are at the maximum distance from the separating hyperplane. Acquiring several new malicious files that are very similar and belong to the same virus family is considered a waste of manual analysis resources, since these files will probably be detected by the same signature. Thus, acquiring one representative file for this set of new malicious files will serve the goal of efficiently updating the signature repository. In order to enhance the signature repository as much as possible, we also check the similarity between the selected files using the kernel farthest-first (KFF) method suggested by Baram et al. [34] which enables us to avoid acquiring examples that are quite similar. Consequently, only the representative files that are most likely malicious are selected).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Yeung and Martin, further incorporating Nissim to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Nissim’s teaching to select only examples of the most malicious threats for future threat detection purposes into Yeung and Martin’s combined method for generating vectors based on malicious and benign user prompts useful in detection of prompt injection attacks. Nissim is also not directed specifically LLM security, but provides establishes a relevant consideration for conserving knowledge base overhead by only including the most impactful classified threats. Nissim reduces redundancy in the maliciousness knowledge base in a similar manner to that of the claimed invention. Additionally, Nissim’s identification of standout threats relative to the broader set of observed files for inclusion in the knowledge base recognizes that these circumstances are more likely to represent legitimate and/or higher risk cases. Combined with Martin, and grounded in the technological environment contributed by Yeung, Nissim’s approaches for building a data store of known malicious threats provide the obvious benefits of limited knowledge base overhead and mitigation of serious system threats. Thus, one of ordinary skill in the art would have been motivated to combine these teachings.
The combination of Yeung, Martin, and Nissim does not expressly teach generating, by the encoding model, a set of [malicious] vectors … and a set of [benign] vectors; and detecting a prompt injection attack using the set of stored vectors.
However, Vaknin teaches generating, by the encoding model, a set of malicious vectors …and a set of benign vectors (Vaknin – Paragraph [0074]: At step 515, as detailed above, the device may identify a plurality of topics present in the prompt. For instance, the device may assess the words present in the prompt and convert them into vector representations, such as by using a FastText library); and detecting a prompt injection attack using the set of stored vectors (Vaknin – Paragraph [0047]: More specifically, prompt analysis process 248 may measure the degree of topic derailment within prompt 304 by quantifying words using word vectors, and then measuring the distance between the vectors such that greater distances represent larger derailment, which is assumed herein to be indicative of a prompt injection attack).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify Yeung, Martin, and Nissim, further incorporating Vaknin to arrive at the conclusion of the claimed invention. One would be motivated to incorporate Vaknin’s teaching to implement an encoding model to generate sets of vectors from malicious and benign prompts to be used in detecting prompt injection attacks into Yeung, Martin, and Nissim’s combined method for segmenting and vectorizing segments of user prompts to detect prompt injection attacks. This addition would provide the system with a relative basis for detecting malicious/benign prompts.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Wildsmith et al. (US 12388849 B2) teaches techniques for performing a similar process for determining user query maliciousness in relation to code inputs
Clement et al. (US 20240386103 A1) teaches techniques for preventing prompt injections attacks using a security agent which inserts secrets to LLM prompts
Cefalu et al. (US 20230359903 A1) teaches a system for preventing malicious user prompt ingestion by an AI model by tokenizing user input and determining whether input tokens are compatible/trusted
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICHOLAS JOSEPH DILUZIO whose telephone number is (703)756-1229. The examiner can normally be reached Mon - Fri -- 7:30 AM - 5 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Yin-Chen Shaw can be reached at 571-272-8878. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NICHOLAS JOSEPH DILUZIO/Examiner, Art Unit 2498
/YIN CHEN SHAW/Supervisory Patent Examiner, Art Unit 2498