DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea mental process without significantly more. The claim(s) recite(s) “[a] method for generating red-teaming data for testing trustworthiness of a target generative model, the method to be implemented by a processor of a testing system, the testing system further including a storage that is electrically connected to the processor and that stores a threat-context dataset related to threats of a generative model, the threat-context dataset including plural pieces of threat- context data that correspond respectively to plural predefined threat categories, each of the pieces of threat-context data including a set of test templates for testing potential threats that belong to the corresponding one of the predefined threat categories, and a threat-test trigger condition related to the corresponding one of the predefined threat categories, the method comprising: sending a reconnaissance prompt to the target generative model for the target generative model to generate a scenario-related response based on the reconnaissance prompt; retrieving the scenario-related response from the target generative model; and generating the red-teaming data based on the scenario-related response and the threat-context dataset stored in the storage”.
The limitation of “generating the red-teaming data based on the scenario-related response and the threat-context dataset stored in the storage” is directed to a step under its broadest reasonable interpretation covers performance of the limitations being a mental process. Therefore, nothing in the claimed elements preclude the step from being performed manually by a human using pencil and paper. If a claim under its broadest reasonable interpretation covers performance in the mind, or by a human using pencil and paper, then the claim falls within the mental process grouping of abstract ideas. Accordingly, claim 1 and 11 recite an abstract idea.
This judicial exception is not integrated into a practical application. Claim 1 [and 11] recite additional elements of:
“the method to be implemented by a processor of a testing system, the testing system further including a storage that is electrically connected to the processor”. The additional element of the processor implementing the method is recited at a high level of generality such that it amounts no more than mere instructions to apply the exception using the generic computing components [processor]. This additional element does not integrate the abstract idea into a practical application and is not an element that is sufficient to amount to significantly more than the judicial exception because it does not impose meaningful limits on practicing the abstract idea.
“the testing system further including a storage …. that stores a threat-context dataset related to threats of a generative model, the threat-context dataset including plural pieces of threat- context data that correspond respectively to plural predefined threat categories, each of the pieces of threat-context data including a set of test templates for testing potential threats that belong to the corresponding one of the predefined threat categories, and a threat-test trigger condition related to the corresponding one of the predefined threat categories”. This additional element is merely storage of data. Storage of data is pre-solution activity that is well-understood, routine, and conventional. Such limitations do not integrate the abstract idea into a practical application and are not elements that are sufficient to amount to significantly more than the judicial exception because the storage of data is pre/solution extra-solution activity.
“sending a reconnaissance prompt to the target generative model for the target generative model to generate a scenario-related response based on the reconnaissance prompt” and “retrieving the scenario-related response from the target generative model”. These additional elements are merely transmission of data/prompt. Transmission of data is activity that is well-understood, routine, and conventional. Such limitations do not integrate the abstract idea into a practical application and are not elements that are sufficient to amount to significantly more than the judicial exception because the transmission of data is extra-solution activity.
“the target generative model to generate a scenario-related response based on the reconnaissance prompt”. The additional element is merely applying a generative model without disclosing any technological advances to the underlying generative model technique, which is not patentable. Merely applying generic machine learning technique without providing technical innovation in the machine learning methods is insufficient for patent eligibility, see Recentive Analytics, Inc. v Fox Corp (April 18, 2025).
Therefore, the additional elements alone, or in combination do not amount to significantly more than the abstract idea, because the elements do not impose meaningful limits on the judicial exception. Thus, claims 1 and 11 are not eligible under 35 USC 101.
Claims 2 and 12 further disclose limitations of additional storage of the vulnerability context dataset and narrowing limitations of the vulnerability context dataset. Storage of data is extra-solution activity that is well-understood, routine, and conventional. Such limitations do not integrate the abstract idea into a practical application and are not elements that are sufficient to amount to significantly more than the judicial exception because the storage of data is pre-solution extra-solution activity. Thus claims 2 and 12 are not eligible under 35 USC 101.
Claims 3 and 13 further disclose limitations of additional storage of the threat identification model and an attack model context dataset and narrowing limitations of the vulnerability context dataset. Storage of data is extra-solution activity that is well-understood, routine, and conventional. Such limitations do not integrate the abstract idea into a practical application and are not elements that are sufficient to amount to significantly more than the judicial exception because the storage of data is pre-solution extra-solution activity. The limitations pertaining to “generating the red-teaming data includes: using the threat identification model to select one of the predefined threat categories based on the scenario-related response and the threat-test trigger conditions respectively of the pieces of threat-context data, where the one of the predefined threat categories thus selected corresponds to one of the threat-test trigger conditions which the scenario-related response involves; using the threat identification model to obtain one of the pieces of threat- context data that corresponds to the one of the predefined threat categories thus selected from the threat-context dataset, and to generate a data-generation instruction based on the set of test templates included in the one of the pieces of threat-context data thus obtained; and using the attack model, based on the data-generation instruction and the vulnerability-context dataset, to generate the red-teaming data” are limitations that can be performed in the human mind using pencil and paper. The additional element of the processor implementing the method is recited at a high level of generality such that it amounts no more than mere instructions to apply the exception using the generic computing components [processor]. This additional element does not integrate the abstract idea into a practical application and is not an element that is sufficient to amount to significantly more than the judicial exception because it does not impose meaningful limits on practicing the abstract idea. Thus claims 3 and 13 are not eligible under 35 USC 101.
Claims 4 and 14 disclose narrowing limitations of generating the red-teaming data which was determined to be an abstract idea. The additional element of the processor implementing the method is recited at a high level of generality such that it amounts no more than mere instructions to apply the exception using the generic computing components [processor]. This additional element does not integrate the abstract idea into a practical application and is not an element that is sufficient to amount to significantly more than the judicial exception because it does not impose meaningful limits on practicing the abstract idea. Thus claims 4 and 14 are not eligible under 35 USC 101.
Claims 5 and 15 further narrows the vulnerability context data and additional steps of “wherein using the attack model to select one of the predefined vulnerability categories includes determining one of the types of generative models indicated by the scenario-related response, and for each of the each of the predefined vulnerability categories, determining whether the ASR that is included in the piece of vulnerability-context data corresponding to the predefined vulnerability category and that corresponds to the one of the types of generative models thus determined is greater than a predetermined threshold value, and selecting the predefined vulnerability category in response to determining that the ASR is greater than the predetermined threshold value.” These limitations are steps that can be achieved in the human mind using pencil and paper. The additional element of the processor implementing the method is recited at a high level of generality such that it amounts no more than mere instructions to apply the exception using the generic computing components [processor]. This additional element does not integrate the abstract idea into a practical application and is not an element that is sufficient to amount to significantly more than the judicial exception because it does not impose meaningful limits on practicing the abstract idea. Thus claims 5 and 15 are not eligible under 35 USC 101.
Claims 6 and 16 further disclose limitations of additional storage of evaluation criterion dataset and narrowing limitations of the evaluation criterion dataset. Storage of data is extra-solution activity that is well-understood, routine, and conventional. Such limitations do not integrate the abstract idea into a practical application and are not elements that are sufficient to amount to significantly more than the judicial exception because the storage of data is pre-solution extra-solution activity. The limitations of “sending the red-teaming data to the target generative model for the target generative model to generate a to-be-evaluated response” and “retrieving the to-be-evaluated response from the target generative model” are limitations that pertains to transmission and gathering of data which is extra solution activity that is well-understood routine and conventional. The limitation of “generating an evaluation result based on the to-be-evaluated response and the evaluation-criteria dataset, the evaluation result indicating whether the target generative model has trustworthiness” is a step can be performed in the human mind using pencil and paper. The additional element of the processor implementing the method is recited at a high level of generality such that it amounts no more than mere instructions to apply the exception using the generic computing components [processor]. This additional element does not integrate the abstract idea into a practical application and is not an element that is sufficient to amount to significantly more than the judicial exception because it does not impose meaningful limits on practicing the abstract idea. Thus claims 6 and 16 are not eligible under 35 USC 101.
Claims 7 and 17 further disclose limitations of additional storage of evaluation criterion dataset and narrowing limitations of the evaluation criterion dataset. Storage of data is extra-solution activity that is well-understood, routine, and conventional. Such limitations do not integrate the abstract idea into a practical application and are not elements that are sufficient to amount to significantly more than the judicial exception because the storage of data is pre-solution extra-solution activity. The limitations of “analyze the to-be-evaluated response so as to determine whether the to-be-evaluated response meets the evaluation standard of one of the evaluation criteria that corresponds to one of the predefined threat categories, where the one of the predefined threat categories corresponds to the piece of threat-context data used to generate the data- generation instruction; in response to determining that the to-be-evaluated response meets the evaluation standard, using the at least one evaluation model to generate an evaluation result indicating that the target generative model has trustworthiness; and in response to determining that the to-be-evaluated response does not meet the evaluation standard, using the at least one evaluation model to generate an evaluation result indicating that the target generative model does not have trustworthiness” are steps that can be achieved in the human mind using pencil and paper. The additional element pertaining to using the model is merely applying a model without disclosing any technological advances to the underlying generative model technique, which is not patentable. Merely applying generic machine learning technique without providing technical innovation in the machine learning methods is insufficient for patent eligibility, see Recentive Analytics, Inc. v Fox Corp (April 18, 2025). The additional element of the processor implementing the method is recited at a high level of generality such that it amounts no more than mere instructions to apply the exception using the generic computing components [processor]. This additional element does not integrate the abstract idea into a practical application and is not an element that is sufficient to amount to significantly more than the judicial exception because it does not impose meaningful limits on practicing the abstract idea. Thus claims 7 and 17 are not eligible under 35 USC 101.
Claims 8, 10, and 18, and 20 further narrow the model limitations recited in the previous claims. The additional element pertaining to using the model is merely applying a model without disclosing any technological advances to the underlying generative model technique, which is not patentable. Merely applying generic machine learning technique without providing technical innovation in the machine learning methods is insufficient for patent eligibility, see Recentive Analytics, Inc. v Fox Corp (April 18, 2025). Thus claims 8, 10, 18 and 20 are not eligible under 35 USC 101.
Claims 9 and 19 further disclose limitations of additional storage of plural evaluation models. The limitations of “generating the evaluation result includes: for each of the evaluation models, using the evaluation model to analyze the to-be-evaluated response so as to determine whether the to-be-evaluated response meets the evaluation standard of one of the evaluation criteria that corresponds to one of the predefined threat categories, the piece of threat-context data corresponding to which is used to generate the data- generation instruction, and based on analysis of the to-be-evaluated response, generate a preliminary result indicating whether the target generative model has trustworthiness; and generating the evaluation result based on the preliminary results that are generated respectively by the evaluation models” are steps that can be performed in the human mind using pencil and paper. The additional element pertaining to using the model is merely applying a model without disclosing any technological advances to the underlying generative model technique, which is not patentable. Merely applying generic machine learning technique without providing technical innovation in the machine learning methods is insufficient for patent eligibility, see Recentive Analytics, Inc. v Fox Corp (April 18, 2025). The additional element of the processor implementing the method is recited at a high level of generality such that it amounts no more than mere instructions to apply the exception using the generic computing components [processor]. This additional element does not integrate the abstract idea into a practical application and is not an element that is sufficient to amount to significantly more than the judicial exception because it does not impose meaningful limits on practicing the abstract idea. Thus claims 9 and 19 are not eligible under 35 USC 101.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
Claims 1-20 are rejected under 35 U.S.C. 112(a) as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor at the time the application was filed, had possession of the claimed invention.
Claim 1 recite the target generative model generates a scenario related response based on the reconnaissance prompt. The written description fails to explain how the claimed function[i.e., the generation of the scenario related responses based on the reconnaissance prompt] is achieved. The written description merely provides broad statements of the function without specifying exactly how the functions steps are actually achieved. To demonstrate that the applicant had possession of the invention, the written description must explain how the claimed function is achieve. The specification does not provide sufficient written description as to how the scenario related responses should be generated.
Claims 3 and 13 recite using the threat identification model to select one of the predefined threat categories based on the scenario-related response and the threat-test trigger conditions respectively of the pieces of threat-context data, using the threat identification model to obtain one of the pieces of threat- context data that corresponds to the one of the predefined threat categories thus selected from the threat-context dataset, and using the attack model, based on the data-generation instruction and the vulnerability-context dataset, to generate the red-teaming data. The written description fails to explain how the claimed functions[i.e., using the threat identification model, the selection of the predefined threat categories, the threat test trigger conditions, using the threat identification model to generate the read teaming data] are achieved. The written description merely provides broad statements of the function without specifying exactly how the functions steps are actually achieved. The written description fails to provide description of how the threat identification model is making the selection of the threat categories and how the threat identification model obtains the threat context data. To demonstrate that the applicant had possession of the invention, the written description must explain how the claimed functions are achieve. The specification does not provide sufficient written description as to how the selection of the predefined threat category is achieved and how the red teaming data should be generated based on using the threat identification model, the selection of the predefined threat categories, the threat test trigger conditions, using the threat identification model to generate the read teaming data.
Claims 4 and 14 also recite using the attack model to select one of the predefined vulnerability categories based on the scenario related response and the attack test trigger conditions, using the attack model to obtain one of the pieces of vulnerability context data, and using the attack model to generate the red teaming data. The written description fails to explain how the claimed functions are accomplished. The written description merely provides broad statements of the function without specifying exactly how the functions steps are actually achieved. To demonstrate that the applicant had possession of the invention, the written description must explain how the claimed functions are achieved.
Claims 5 and 15 recite “determining one of the types of generative models…”. The written description fails to explain how the claimed functions of the determining one of the types of generative models are accomplished. The written description merely provides broad statements of the function without specifying exactly how the functions steps are actually achieved. To demonstrate that the applicant had possession of the invention, the written description must explain how the claimed functions are achieved.
Claims 6 and 16 recite sending the red-teaming data to the target generative model for the target generative model to generate a to-be-evaluated response; retrieving the to-be-evaluated response from the target generative model; and generating an evaluation result based on the to-be-evaluated response and the evaluation-criteria dataset, the evaluation result indicating whether the target generative model has trustworthiness. The written description fails to explain how the claimed functions[i.e., the generation to be evaluated response and the generation of the evaluation result] are achieved. The written description merely provides broad statements of the functions without specifying exactly how the functions steps are actually achieved. To demonstrate that the applicant had possession of the invention, the written description must explain how the claimed function is achieve. The specification does not provide sufficient written description as to how the to be evaluated response and the evaluation result are generated.
Claims 7 and 17 recite using the at least one evaluation model to analyze the to-be-evaluated response so as to determine whether the to-be-evaluated response meets the evaluation standard of one of the evaluation criteria that corresponds to one of the predefined threat categories. The written description fails to explain how the claimed functions[i.e., the analyzing of the to be evaluated response and the determination whether the to-be-evaluated response meets the evaluation standard] are achieved. The written description merely provides broad statements of the function without specifying exactly how the functions steps are actually achieved. To demonstrate that the applicant had possession of the invention, the written description must explain how the claimed function is achieve. The specification does not provide sufficient written description as to how the analyzing of the to be evaluated response and the determination whether the to-be-evaluated response meets the evaluation standard are accomplished.
Claims 2, 8-12, and 18-20 are rejected to because said claims depend upon claim 1.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1 and 11 recite “predefined threat categories” but fail to particularly point out and specify what these categories are. Claims 1 and 11 further recite “threat-test trigger condition” but fail to particularly point out and specify/define the condition.
Claim 1 further recite that a target generative model to generate a scenario related response based on the reconnaissance prompt. The claim fail to specify how the target generative mode generates the scenario related response.
Claims 2 and 12 recite “predefined vulnerability categories” but fail to particularly point out and specify what these categories are.
Claims 3 and 13 recite “using the threat identification model to select one of the predefined threat categories… using the threat identification model to obtain one of the pieces of threat- context data”. The claim failed to particularly point out what the predefined threat categories are, how the threat identification model is selecting/choosing the predefined threat categories, and how the threat identification model obtains the threat context data.
Claims 4 and 14 recite using the attack model to select one of the predefined vulnerability categories based on the scenario related response and the attack test trigger conditions, using the attack model to obtain one of the pieces of vulnerability context data, and using the attack model to generate the red teaming data. The claims fail to particularly point out how the models are achieving said functions.
Claims 5 and 15 recite “determining one of the types of generative models…”. The claims fail to explain how the claimed function of the determining one of the types of generative models is accomplished.
Claims 6 and 16 recite sending the red-teaming data to the target generative model for the target generative model to generate a to-be-evaluated response; retrieving the to-be-evaluated response from the target generative model; and generating an evaluation result based on the to-be-evaluated response and the evaluation-criteria dataset, the evaluation result indicating whether the target generative model has trustworthiness. The claims fail to explain how the claimed functions[i.e., the generation to be evaluated response and the generation of the evaluation result] are achieved.
Claims 7 and 17 recite using the at least one evaluation model to analyze the to-be-evaluated response so as to determine whether the to-be-evaluated response meets the evaluation standard of one of the evaluation criteria that corresponds to one of the predefined threat categories. The claims fail to explain how the claimed functions[i.e., the analyzing of the to be evaluated response and the determination whether the to-be-evaluated response meets the evaluation standard] are achieved.
Claims 2, 8-12, and 18-20 are rejected as being dependent on, and failing to cure the deficiencies of, rejected independent claim 1.
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claims 11-20 are rejected under 35 U.S.C. 112(d) as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Claim 11 is a testing system claim that depends on the method claim 1. The system claim 11 recites limitations that is stored in storage, that is also stated in claim 1. The claim 11 does not provide further limitations that narrows the limitations presented in claim 1. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claims 12-20 are rejected to because said claims depend upon claim 11.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dai et al WO 2025188233 (hereinafter Dai), in view of Apple et al US 20190222585 (hereinafter Apple), and in further view of Blair et al US 20250252192 (hereinafter Blair).
As to claim 1, Dai teaches a method for generating red-teaming data for testing trustworthiness of a target generative model (Figure 4C discloses operation steps of the red teaming system. Paragraph 18 discloses the invention pertains to a security assessment system for a target model, comprising an assessment system, a red teaming system, wherein the red teaming system is configured to submit a red team input to the target model for generating a red team output, and to further receive the first red team output for analysis; wherein a first pair comprising the assessment input and the assessment output is configured to be provided to the red teaming system for improving attack coverage of the red teaming system; and, wherein a second pair comprising the red team input and the red team output is configured to be provided to the assessment system for improving assessment coverage of the assessment system)…
the testing system further including …a threat-context dataset related to threats of a generative model, the threat-context dataset including plural pieces of threat- context data that correspond respectively to plural predefined threat categories (paragraph 53 and Figure 3A disclose the GenAI Protection System is equipped with an array of GenAI Input-based Detection & Prevention Techniques (IDPT) 301 and GenAI Output-based Detection & Prevention Techniques. Paragraph 54 discloses examples categories of IDPT include poisoned data D&P, prompt injection D&P, and model theft or extraction D&P. Examples of ODPT include poisoned model D&P, model jailbreak D&P, Pll (Personally Identifiable Information) or data leakage D&P, and toxicity or harmful content D&P), generating the red-teaming data based on the scenario-related response and the threat-context dataset (Figure 4B, reference number 432 “Generate security testing report. Paragraphs 91-99 disclose that the input and output responses are evaluated to determine model vulnerabilities and the resulting analysis is reported to the GenAI red teaming system. The GenAI red teaming system processes security testing findings/(scenario related responses) and the red teaming techniques/threat context dataset) the method comprising:
sending a reconnaissance prompt to the target generative model for the target generative model to generate a scenario-related response based on the reconnaissance prompt (Figure 4B, reference number 421 “Submit crafted input prompt to GenAI model directly” and reference number 426 “Submit crafted input prompt to configured GenAI model wrapper”. See also paragraphs 89 and 94);
retrieving the scenario-related response from the target generative model (Figure 4B, reference number 422 “Receive generated output from GenAI model directly” and reference number 427 “Receive hardened output from configured GenAI model wrapper”. See also paragraphs 90 and 95); and
generating the red-teaming data based on the scenario-related response and the threat-context dataset …(Figure 4B, reference number 432 “Generate security testing report. Paragraphs 91-99 disclose that the input and output responses are evaluated to determine model vulnerabilities and the resulting analysis is reported to the GenAI red teaming system. The GenAI red teaming system complies and processes security testing finding/(scenario related responses) and the red teaming techniques/threat context dataset).
Dai does not teach a method to be implemented by a processor of a testing system, the testing system further including a storage that is electrically connected to the processor and that stores a threat-context dataset related to threats of a generative model, each of the pieces of threat-context data including a set of test templates for testing potential threats that belong to the corresponding one of the predefined threat categories, and a threat-test trigger condition related to the corresponding one of the predefined threat categories, generating the red-teaming data based on the scenario-related response and the threat-context dataset stored in the storage.
Apple teaches a method to be implemented by a processor of a testing system (abstract discloses system and method predicting and remediating malware threats in an electronic computer network. Paragraph 87 discloses processor is operable to execute instructions ), the testing system further including a storage that is electrically connected to the processor (paragraph 88 discloses the processing system include memory system) and that stores a threat-context dataset related to threats of a generative model (abstract and paragraph 8 discloses the concept of storing in an electronic persistent storage library data representing a plurality of threats, these threats may be threats related to plurality of randomizing algorithms that may include a generative artificial intelligence algorithm).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s method of generating red-teaming data with Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data to anticipate the developments of cyber attackers and malign organizations that intrude upon information security and predict and remediate threats in an electronic computer network (paragraphs 7-8 of Apple).
The combination of Dai in view of Apple does not teach, but Blair teaches each of the pieces of threat-context data including a set of test templates for testing potential threats that belong to the corresponding one of the predefined threat categories (paragraphs 65-66 discloses the concept of customized red teaming processing template that defines use case information, legal issue or compliance information, and/or any other data useful for red teaming), and a threat-test trigger condition related to the corresponding one of the predefined threat categories (paragraph 3 discloses the concept of that red teaming exercises are designed to test the models' robustness, resilience, and reliability by exposing them to a variety of inputs and conditions (trigger conditions) that simulate potential manipulation or attack scenarios), generating the red-teaming data based on the scenario-related response and the threat-context dataset (paragraphs 46-48 disclose generating [scenario related] responses based on the one or more prompt junctures. Paragraph 42 discloses risk taxonomy (threat context dataset), which may include a matrix comprising the set of compliance topics on a first axis and the interaction topics on a second axis, such that the compliance topics and interaction topics intersect at one or more junctures. Paragraph 40 discloses the set of interaction topics related to privacy policies, protecting personal data, and fine avoidance) stored in the storage (paragraph 66 discloses the generative intelligence system stores the data input, recall from paragraphs 40-42 that this data input is the risk taxonomy/threat context dataset. Paragraph 46 also disclose the response is stored in database).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s method of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data with Blair’s teachings of trigger conditions to provide innovative approach to red teaming that identify and exploit subtle vulnerabilities in complex AI systems (paragraphs 4-5 of Blair).
As to claim 11, Dai teaches a testing system for generating red-teaming data for implementing a red team assessment on a target generative model (Figure 4C discloses operation steps of the red teaming system. Paragraph 18 discloses the invention pertains to a security assessment system for a target model, comprising an assessment system, a red teaming system, wherein the red teaming system is configured to submit a red team input to the target model for generating a red team output, and to further receive the first red team output for analysis; wherein a first pair comprising the assessment input and the assessment output is configured to be provided to the red teaming system for improving attack coverage of the red teaming system; and, wherein a second pair comprising the red team input and the red team output is configured to be provided to the assessment system for improving assessment coverage of the assessment system), said testing system comprising: a threat-context dataset related to threats of a generative model, the threat-context dataset including plural pieces of threat-context data that correspond respectively to plural predefined threat categories (paragraph 53 and Figure 3A disclose the GenAI Protection System is equipped with an array of GenAI Input-based Detection & Prevention Techniques (IDPT) 301 and GenAI Output-based Detection & Prevention Techniques. Paragraph 54 discloses examples categories of IDPT include poisoned data D&P, prompt injection D&P, and model theft or extraction D&P. Examples of ODPT include poisoned model D&P, model jailbreak D&P, Pll (Personally Identifiable Information) or data leakage D&P, and toxicity or harmful content D&P), ,wherein implements the method of claim 1 (claim 1 above discloses the method which Dai teaches).
Dai does not teach wherein said testing system comprising: a processor; and a storage that is electrically connected to said processor, and that stores a threat-context dataset related to threats of a generative model, the threat-context dataset including plural pieces of threat-context data that correspond respectively to plural predefined threat categories, each of the pieces of threat-context data including a set of test templates for testing potential threats that belong to the corresponding one of the predefined threat categories, and a threat-test trigger condition related to the corresponding one of the predefined threat categories; wherein said processor implements the method.
Ahmed teaches wherein said testing system comprising: a processor (abstract discloses system and method predicting and remediating malware threats in an electronic computer network. Paragraph 87 discloses processor is operable to execute instructions); and a storage that is electrically connected to said processor (paragraph 88 discloses the processing system include memory system), and that stores a threat-context dataset related to threats of a generative model (abstract and paragraph 8 discloses the concept of storing in an electronic persistent storage library data representing a plurality of threats, these threats may be threats related to plurality of randomizing algorithms that may include a generative artificial intelligence algorithm);wherein said processor implements the method (abstract discloses system and method predicting and remediating malware threats in an electronic computer network. Paragraph 87 discloses processor is operable to execute instructions).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s method of generating red-teaming data with Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data to anticipate the developments of cyber attackers and malign organizations that intrude upon information security and predict and remediate threats in an electronic computer network (paragraphs 7-8 of Apple).
The combination of Dai in view of Apple does not teach, but Blair teaches each of the pieces of threat-context data including a set of test templates for testing potential threats that belong to the corresponding one of the predefined threat categories (paragraphs 65-66 disclose the concept of customized red teaming processing template that defines use case information, legal issue or compliance information, and/or any other data useful for red teaming), and a threat-test trigger condition related to the corresponding one of the predefined threat categories (paragraph 3 discloses the concept of that red teaming exercises are designed to test the models’ robustness, resilience, and reliability by exposing them to a variety of inputs and conditions (trigger conditions) that simulate potential manipulation or attack scenarios).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s method of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data with Blair’s teachings of trigger conditions to provide innovative approach to red teaming that identify and exploit subtle vulnerabilities in complex AI systems (paragraphs 4-5 of Blair).
Claim(s) 2 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dai et al WO 2025188233 (hereinafter Dai), in view of Apple et al US 20190222585 (hereinafter Apple), in further view of Blair et al US 20250252192 (hereinafter Blair), and in further view of Krishnamurthy et al US 20260093551 (hereinafter Krishnamurthy).
As to claim 2, the combination of Dai in view of Apple and Blair teaches all the limitations recited in claim 1 above and further teaches wherein generating the red team data is implemented further based on the vulnerability context dataset (Dai: Figure 4B, reference number 432 “Generate security testing report. Paragraphs 91-99 disclose that the input and output responses are evaluated to determine model vulnerabilities and the resulting analysis is reported to the GenAI red teaming system. The GenAI red teaming system complies and processes security testing finding/(scenario related responses) and the red teaming techniques/threat context dataset).
The combination of Dai in view of Apple and Blair does not teach, but Krishnamurthy teaches the storage further storing a vulnerability-context dataset related to a generative model (Figure 2B, reference number 226 discloses vulnerability detection subsystem of the generative AI model within in the score generating subsystem. Figure 2A reveals the score generating subsystem is stored within memory unit), the vulnerability-context dataset including plural pieces of vulnerability-context data that correspond respectively to plural predefined vulnerability categories, each of the pieces of vulnerability- context data including at least one adversarial attack technique that targets vulnerability belonging to the corresponding one of the predefined vulnerability categories, at least one set of adversarial attack templates that respectively uses the at least one adversarial attack technique for attacking the vulnerabilities belonging to the corresponding one of the predefined vulnerability categories, and an attack-test trigger condition that is related to the corresponding one of the predefined vulnerability categories (paragraphs 100-101 disclose the vulnerability detection subsystem is configured to evaluate how robust the one or more generative AI models are when exposed to jailbreaking or prompt injection attacks [ the jailbreaking and prompt injection are vulnerability categories], where malicious prompts attempt to override the intended behavior of the one or more generative AI models. To accomplish this, the system utilizes the one or more databases of jailbreaking test cases (templates), which have been collected during red teaming efforts conducted on similar large language models (LLMs). The jailbreaking test cases include a plurality of predefined attack prompts and their corresponding expected responses, based on how the LLMs behaved during prior testing under malicious scenarios. The vulnerability detection subsystem compares the first output data (i.e., the actual response generated by the one or more generative AI models in the real-time evaluation) to these predefined attack cases. The vulnerability detection subsystem applies the similarity function to assess how closely the first output data matches the expected responses in the jailbreaking test cases. Specifically, it uses the predefined vulnerability threshold score of 0.8, meaning that if the similarity between the first output data and the second output data is greater than or equal to 0.8, it indicates a strong resemblance to what would be expected if the model had succumbed to a successful prompt injection attack. This similarity metric helps the system determine whether the AI model has successfully resisted or is vulnerable to the attack).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s system of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data and Blair’s teachings of trigger conditions with Krishnamurthy’s vulnerability-context dataset to provide a comprehensive, multi-dimensional evaluation framework that is able to accurately assess various aspects of the AI model's performance and evaluate the AI models across different architectures and deployment scenarios, while providing consistent and comparable results (paragraph 10 of Krishnamurthy).
As to claim 12, the combination of Dai in view of Apple and Blair teaches all the limitations recited in claim 11 above and further teaches wherein said processor (Apple: abstract discloses system and method predicting and remediating malware threats in an electronic computer network. Paragraph 87 discloses processor is operable to execute instructions ) generates the red team data based on the vulnerability context dataset (Dai: Figure 4B, reference number 432 “Generate security testing report. Paragraphs 91-99 disclose that the input and output responses are evaluated to determine model vulnerabilities and the resulting analysis is reported to the GenAI red teaming system. The GenAI red teaming system complies and processes security testing finding/(scenario related responses) and the red teaming techniques/threat context dataset).
The combination of Dai in view of Apple and Blair does not teach, but Krishnamurthy teaches the storage further stores a vulnerability-context dataset related to a generative model (Figure 2B, reference number 226 discloses vulnerability detection subsystem of the generative AI model within in the score generating subsystem. Figure 2A reveals the score generating subsystem is stored within memory unit), the vulnerability-context dataset including plural pieces of vulnerability-context data that correspond respectively to plural predefined vulnerability categories, each of the pieces of vulnerability- context data including at least one adversarial attack technique that targets vulnerability belonging to the corresponding one of the predefined vulnerability categories, at least one set of adversarial attack templates that respectively uses the at least one adversarial attack technique for attacking the vulnerabilities belonging to the corresponding one of the predefined vulnerability categories, and an attack-test trigger condition that is related to the corresponding one of the predefined vulnerability categories (paragraphs 100-101 disclose the vulnerability detection subsystem is configured to evaluate how robust the one or more generative AI models are when exposed to jailbreaking or prompt injection attacks, where malicious prompts attempt to override the intended behavior of the one or more generative AI models. To accomplish this, the system utilizes the one or more databases of jailbreaking test cases (templates), which have been collected during red teaming efforts conducted on similar large language models (LLMs). The jailbreaking test cases include a plurality of predefined attack prompts and their corresponding expected responses, based on how the LLMs behaved during prior testing under malicious scenarios. The vulnerability detection subsystem compares the first output data (i.e., the actual response generated by the one or more generative AI models in the real-time evaluation) to these predefined attack cases. The vulnerability detection subsystem applies the similarity function to assess how closely the first output data matches the expected responses in the jailbreaking test cases. Specifically, it uses the predefined vulnerability threshold score of 0.8, meaning that if the similarity between the first output data and the second output data is greater than or equal to 0.8, it indicates a strong resemblance to what would be expected if the model had succumbed to a successful prompt injection attack. This similarity metric helps the system determine whether the AI model has successfully resisted or is vulnerable to the attack).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s system of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data and Blair’s teachings of trigger conditions with Krishnamurthy’s vulnerability-context dataset to provide a comprehensive, multi-dimensional evaluation framework that is able to accurately assess various aspects of the AI model's performance and evaluate the AI models across different architectures and deployment scenarios, while providing consistent and comparable results (paragraph 10 of Krishnamurthy).
Claim(s) 3-10 and 13-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dai et al WO 2025188233 (hereinafter Dai), in view of Apple et al US 20190222585 (hereinafter Apple), in further view of Blair et al US 20250252192 (hereinafter Blair), in further view of Krishnamurthy et al US 20260093551 (hereinafter Krishnamurthy), and in further view of Ahmed et al US 20250055867 (hereinafter Ahmed).
As to claim 3, the combination of Dai in view of Apple, Blair, and Krishnamurthy teaches all the limitations recited in claim 2 above and further teaches the storage further a threat identification model and an attack model (Krishnamurthy: Figure 2B, reference number 226 discloses vulnerability detection subsystem. Paragraphs 100-101 disclose the vulnerability detection subsystem is configured to evaluate how robust the one or more generative AI models are when exposed to jailbreaking or prompt injection attacks (from an attack model). To accomplish this, the system utilizes the one or more databases of jailbreaking test cases (templates), which have been collected during red teaming efforts conducted on similar large language models (LLMs). The jailbreaking test cases include a plurality of predefined attack prompts and their corresponding expected responses, based on how the LLMs behaved during prior testing under malicious scenarios. The vulnerability detection subsystem compares the first output data (i.e., the actual response generated by the one or more generative AI models in the real-time evaluation) to these predefined attack cases. The vulnerability detection subsystem applies the similarity function to assess how closely the first output data matches the expected responses in the jailbreaking test cases. Specifically, it uses the predefined vulnerability threshold score of 0.8, meaning that if the similarity between the first output data and the second output data is greater than or equal to 0.8, it indicates a strong resemblance to what would be expected if the model had succumbed to a successful prompt injection attack. This similarity metric helps the system determine whether the AI model has successfully resisted or is vulnerable to the attack. Thus identify the threat based on the modeled data) and generating the red-teaming data (Dai: Figure 4B, reference number 432 “Generate security testing report. Paragraphs 91-99 disclose that the input and output responses are evaluated to determine model vulnerabilities and the resulting analysis is reported to the GenAI red teaming system. The GenAI red teaming system complies and processes security testing finding/(scenario related responses).
The combination of Dai in view of Apple, Blair, and Krishnamurthy does not teach, but Ahmed teaches wherein generating the red-teaming data includes: using the threat identification model (paragraph 95 discloses a threat detection model) to select one of the predefined threat categories based on the scenario-related response and the threat-test trigger conditions respectively of the pieces of threat-context data (Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats. The assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. The assessment engine computes a second threat probability score for each of the one or more sub-types of threats by the one or more nano classifiers. In an embodiment, the one or more nano classifiers may be selected based on the aforementioned comparison and a plurality of predefined policies associated with the user. The assessment engine detects the one or more sub-types of the threats in the data based on a comparison of the second threat probability score with a predefined threshold value of the second threat probability score), where the one of the predefined threat categories thus selected corresponds to one of the threat-test trigger conditions which the scenario-related response involves (Figure 5 and paragraphs 94-95 disclose the detected sub type of threats are determined based an intent and category of the prompt and content based on the attributes of the data (trigger conditions); using the threat identification model to obtain one of the pieces of threat-context data that corresponds to the one of the predefined threat categories thus selected from the threat-context dataset (Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats), and to generate a data-generation instruction based on the set of test templates included in the one of the pieces of threat-context data thus obtained (paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. The assessment engine computes a second threat probability score for each of the one or more sub-types of threats by the one or more nano classifiers. In an embodiment, the one or more nano classifiers may be selected based on the aforementioned comparison and a plurality of predefined policies (test templates). The assessment engine detects the one or more sub-types of the threats in the data based on a comparison of the second threat probability score with a predefined threshold value of the second threat probability score. Paragraphs 96 and 99 further reveals on detection of subtypes of threats, the data is selectively moderated based on predefined rules corresponding to each of the one or more sub-type of threats to obtain a moderated data, and the moderated data is used to obtain an improved/modified prompt instructions); and using the attack model, based on the data-generation instruction and the vulnerability-context dataset, to generate the red-teaming data (paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. Paragraph 99 discloses the approved/moderated prompt and content is transmitted to the generative AI model, wherein the generative AI model generate output data based on the moderated/approved prompt and attribute and context data [the context data is the vulnerability-context dataset]). Paragraph 47 discloses examples of attribute and context data wherein examples of such attributes may include but are not limited to, nature of the input data, nature of the output data, usage history and context associated with the data, and similarity with past violations and threats).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s method of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data, Blair’s teachings of trigger conditions, and Krishnamurthy’s vulnerability-context dataset with Ahmed’s threat identification model to dynamically mitigate threats of a generative AI model (paragraph 2 of Ahmed).
As to claim 4, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein using the attack model to generate the red-teaming data includes: using the attack model to select one of the predefined vulnerability categories based on the scenario-related response and the attack-test trigger conditions respectively of the pieces of vulnerability-context data, where the one of the predefined vulnerability categories thus selected corresponds to one of the attack-test conditions which the scenario-related response involves ( Ahmed: Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats. Paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. Paragraph 99 discloses the approved/moderated prompt and content is transmitted to the generative AI model, wherein the generative AI model generate output data based on the moderated/approved prompt and attribute and context data [the context data is the vulnerability-context dataset]). Paragraph 47 discloses examples of attribute and context data wherein examples of such attributes may include but are not limited to, nature of the input data, nature of the output data, usage history and context associated with the data, and similarity with past violations and threats); using the attack model to obtain one of the pieces of vulnerability-context data that corresponds to the one of the predefined attack-test conditions from the vulnerability-context dataset (Ahmed: Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats. Paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data); and using the attack model to generate the red-teaming data based on the data- generation instruction and the at least one set of adversarial attack templates included in the one of the pieces of vulnerability-context data thus obtained ( Ahmed: Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats. Paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. Paragraph 99 discloses the approved/moderated prompt and content is transmitted to the generative AI model, wherein the generative AI model generate output data based on the moderated/approved prompt and attribute and context data [the context data is the vulnerability-context dataset]). Paragraph 47 discloses examples of attribute and context data wherein examples of such attributes may include but are not limited to, nature of the input data, nature of the output data, usage history and context associated with the data, and similarity with past violations and threats). The motivation is similar to the motivation presented in claim 3.
As to claim 5, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein, for each of the predefined vulnerability categories, the attack-test trigger condition of the piece of vulnerability-context data that corresponds to the predefined vulnerability category includes plural attack success rates (ASRs) of adversarial attacks respectively against different types of generative models by targeting the vulnerabilities that belong to the predefined vulnerability category, wherein using the attack model to select one of the predefined vulnerability categories includes determining one of the types of generative models indicated by the scenario-related response, and for each of the each of the predefined vulnerability categories (Ahmed: Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats. Paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. Paragraph 99 discloses the approved/moderated prompt and content is transmitted to the generative AI model, wherein the generative AI model generate output data based on the moderated/approved prompt and attribute and context data [the context data is the vulnerability-context dataset]). Paragraph 47 discloses examples of attribute and context data wherein examples of such attributes may include but are not limited to, nature of the input data, nature of the output data, usage history and context associated with the data, and similarity with past violations and threats), determining whether the ASR that is included in the piece of vulnerability-context data corresponding to the predefined vulnerability category and that corresponds to the one of the types of generative models thus determined is greater than a predetermined threshold value (Ahmed: paragraph 59 discloses when detecting the value of the computed second threat probability score greater than or equal to the predefined threshold score, the system may flag the data as containing the specific subtype of threat. For example, as in previous example, the toxicity second threat probability score of 0.85 exceeds the predefined threshold score of 0.7, leading the system to flag the data as containing a toxicity threat. Paragraph 94 further discloses the assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect one or more sub-type of threats associated with the one or more type of threats), and selecting the predefined vulnerability category in response to determining that the ASR is greater than the predetermined threshold value (Ahmed: paragraph 59 discloses when detecting the value of the computed second threat probability score greater than or equal to the predefined threshold score, the system may flag the data as containing the specific subtype of threat. For example, as in previous example, the toxicity second threat probability score of 0.85 exceeds the predefined threshold score of 0.7, leading the system to flag the data as containing a toxicity threat. Paragraph 94 further discloses the assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect one or more sub-type of threats associated with the one or more type of threats). The motivation is similar to the motivation presented in claim 3.
As to claim 6, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches the storage further storing an evaluation-criterion dataset (Ahmed: Figure 2 reveals the memory 204 holds instructions for assessment engine 208), the evaluation-criterion dataset including plural evaluation criteria that are related respectively to plural predefined assessment items corresponding respectively to the predefined threat categories, each of the evaluation criteria including an evaluation standard for assessing the corresponding one of the predefined assessment items (Ahmed: paragraph 94 discloses the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data to detect threats associated with the prompt and content. The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect one or more sub-type of threats associated with the one or more type of threats), the method further comprising: sending the red-teaming data to the target generative model for the target generative model to generate a to-be-evaluated response (Dai: Figure 4A reveals input data from GenAI Red Teaming system is provided to the GenAI model, and the output is provided to the GenAI assessment system. Ahmed: Figure 3 further discloses that the output response from the generative AI model is provide to the assessment engine to generate an evaluated response such as identify type of threat, see paragraph 99); retrieving the to-be-evaluated response from the target generative model (Dai: Figure 4A reveals input data from GenAI Red Teaming system is provided to the GenAI model, and the output is provided to the GenAI assessment system. Ahmed: Figure 3 further discloses that the output response from the generative AI model is provide to the assessment engine to generate an evaluated response such as identify type of threat, see paragraph 99); and generating an evaluation result based on the to-be-evaluated response and the evaluation-criteria dataset, the evaluation result indicating whether the target generative model has trustworthiness (Ahmed: paragraph 99 discloses the assessment engine associated with the output sub-layer may perform various operations for mitigating threat types of the output response, if present. The assessment engine may analyze the output response to determine an intent and a category of the output response based on the attributes of the output response. The assessment engine may identify one or more sub types of threats associated with the output response based on the analysis of the output response. The one or more sub types of threats may include the profanity and toxicity check, a third-party Intellectual Property (IP) violation check, the organization policy and role-based check, and a hallucination check. Paragraph 100 discloses The assessment engine may further determine, in real-time, a threat score corresponding to each of the sub types of threats, for the output response received via the generative AI model, based on the plurality of attributes. Once the threat score is determined for each of the sub types of threats, the threat probability score may be compared with a predefined threshold for the threshold value of the second threat probability score, by the assessment engine. Further, the assessment engine may select one or more sub types of threats, based on the comparison and a plurality of predefined policies associated with the user. Figure 3 and paragraph 102 disclose the validation engine outputs a validation/evaluation result based on the output response and the output sub layer check/evaluation criteria. The validation result indicates whether the response should be modified, approved or rejected). The motivation is similar to the motivation presented in claim 3.
As to claim 7, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches the storage further storing at least one evaluation model (Ahmed: Figure 2 reveals the memory 204 holds instructions for assessment engine 208), wherein generating the evaluation result includes: using the at least one evaluation model to analyze the to-be-evaluated response so as to determine whether the to-be-evaluated response meets the evaluation standard of one of the evaluation criteria that corresponds to one of the predefined threat categories, where the one of the predefined threat categories corresponds to the piece of threat-context data used to generate the data- generation instruction (Ahmed: paragraph 99 discloses the assessment engine may analyze the output response to determine an intent and a category of the output response based on the attributes of the output response. The assessment engine may identify one or more sub types of threats associated with the output response based on the analysis of the output response. The one or more sub types of threats may include the profanity and toxicity check, a third-party Intellectual Property (IP) violation check, the organization policy and role-based check, and a hallucination check. The sub types of threats corresponds to attributes data and intent which can be threat context data used to generate the data generation instruction/modified response/approved prompt); in response to determining that the to-be-evaluated response meets the evaluation standard, using the at least one evaluation model to generate an evaluation result indicating that the target generative model has trustworthiness (Ahmed: : Figure 3 reveals the output from the validation engine such as rejection of the output response, approved response (corresponds to trustworthiness), and modified response based on the check of the type of threat present in the output response. Paragraph 103 discloses if the validation of the output response is successful without moderation, an approved response may be transmitted to the user through the generative AI model. The approved response may be the output response without any moderation. Otherwise, if the validation of the output response is successful after moderation, a modified response may be transmitted to the user through the generative AI model); and in response to determining that the to-be-evaluated response does not meet the evaluation standard, using the at least one evaluation model to generate an evaluation result indicating that the target generative model does not have trustworthiness (Ahmed: Figure 3 reveals the output from the validation engine such as rejection of the output response (corresponds to the target generative model does not have trustworthiness), approved response, and modified response based on the check of the type of threat present in the output response. Paragraph 102 discloses when the validation of the output response is unsuccessful even after performing some predefined iterations of moderation, the output response may be restricted from further processing, and details of restricting the output response (such as a user rejection message) may be transmitted and rendered to the user. When the validation of the prompt and content is unsuccessful even after performing some predefined iterations of moderation, the prompt and content may be restricted from further processing, and details of restricting the prompt and content (such as a user rejection message) may be transmitted and rendered to the user). The motivation is similar to the motivation presented in claim 3.
As to claim 8, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein said at least one evaluation model is a generative pre-trained transformer (Ahmed: paragraphs 17 and 33 reveal the system includes models that are Generative Pretrained Transformers).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s method of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data, Blair’s teachings of trigger conditions, and Krishnamurthy’s vulnerability-context dataset with Ahmed’s generative pre-trained transformer such that the system can generate new and original content that may be used in various red team applications (paragraph 33 of Ahmed).
As to claim 9, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches the storage further storing plural evaluation models (Ahmed: Figure 2 reveals the memory 204 holds instructions for assessment engine 208), wherein generating the evaluation result includes: for each of the evaluation models, using the evaluation model to analyze the to-be-evaluated response so as to determine whether the to-be-evaluated response meets the evaluation standard of one of the evaluation criteria that corresponds to one of the predefined threat categories, the piece of threat-context data corresponding to which is used to generate the data- generation instruction (Ahmed: paragraph 99 discloses the assessment engine may analyze the output response to determine an intent and a category of the output response based on the attributes of the output response. The assessment engine may identify one or more sub types of threats associated with the output response based on the analysis of the output response. The one or more sub types of threats may include the profanity and toxicity check, a third-party Intellectual Property (IP) violation check, the organization policy and role-based check, and a hallucination check. The sub types of threats corresponds to attributes data and intent which can be threat context data used to generate the data generation instruction/modified response/approved prompt), and based on analysis of the to-be-evaluated response, generate a preliminary result indicating whether the target generative model has trustworthiness (Ahmed: : Figure 3 reveals the output from the validation engine such as rejection of the output response, approved response (corresponds to trustworthiness), and modified response based on the check of the type of threat present in the output response. Paragraph 103 discloses if the validation of the output response is successful without moderation, an approved response may be transmitted to the user through the generative AI model. The approved response may be the output response without any moderation. Otherwise, if the validation of the output response is successful after moderation, a modified response may be transmitted to the user through the generative AI model); and generating the evaluation result based on the preliminary results that are generated respectively by the evaluation models (Ahmed: paragraph 99 discloses the assessment engine may analyze the output response to determine an intent and a category of the output response based on the attributes of the output response. The assessment engine may identify one or more sub types of threats associated with the output response based on the analysis of the output response. The one or more sub types of threats may include the profanity and toxicity check, a third-party Intellectual Property (IP) violation check, the organization policy and role-based check, and a hallucination check. The sub types of threats corresponds to attributes data and intent which can be threat context data used to generate the data generation instruction/modified response/approved prompt). The motivation is similar to the motivation presented in claim 3.
As to claim 10, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein each of the threat identification model and the attack model is a generative pre-trained transformer (Ahmed: paragraphs 17 and 33 reveal the system includes models that are Generative Pretrained Transformers).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s method of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data, Blair’s teachings of trigger conditions, and Krishnamurthy’s vulnerability-context dataset with Ahmed’s generative pre-trained transformer such that the system can generate new and original content that may be used in various red team applications (paragraph 33 of Ahmed).
As to claim 13, the combination of Dai in view of Apple, Blair, and Krishnamurthy teaches all the limitations recited in claim 12 above and further teaches wherein said storage further stores a threat identification model and an attack model (Krishnamurthy: Figure 2B, reference number 226 discloses vulnerability detection subsystem. Paragraphs 100-101 disclose the vulnerability detection subsystem is configured to evaluate how robust the one or more generative AI models are when exposed to jailbreaking or prompt injection attacks (from an attack model). To accomplish this, the system utilizes the one or more databases of jailbreaking test cases (templates), which have been collected during red teaming efforts conducted on similar large language models (LLMs). The jailbreaking test cases include a plurality of predefined attack prompts and their corresponding expected responses, based on how the LLMs behaved during prior testing under malicious scenarios. The vulnerability detection subsystem compares the first output data (i.e., the actual response generated by the one or more generative AI models in the real-time evaluation) to these predefined attack cases. The vulnerability detection subsystem applies the similarity function to assess how closely the first output data matches the expected responses in the jailbreaking test cases. Specifically, it uses the predefined vulnerability threshold score of 0.8, meaning that if the similarity between the first output data and the second output data is greater than or equal to 0.8, it indicates a strong resemblance to what would be expected if the model had succumbed to a successful prompt injection attack. This similarity metric helps the system determine whether the AI model has successfully resisted or is vulnerable to the attack. Thus identify the threat based on the modeled data) and said processor (Apple: abstract discloses system and method predicting and remediating malware threats in an electronic computer network. Paragraph 87 discloses processor is operable to execute instructions) generates the red-teaming data (Dai: Figure 4B, reference number 432 “Generate security testing report. Paragraphs 91-99 disclose that the input and output responses are evaluated to determine model vulnerabilities and the resulting analysis is reported to the GenAI red teaming system. The GenAI red teaming system complies and processes security testing finding/(scenario related responses).
The combination of Dai in view of Apple, Blair, and Krishnamurthy does not teach, but Ahmed teaches generates the red-teaming data by: using the threat identification model (paragraph 95 discloses a threat detection model) to select one of the predefined threat categories based on the scenario-related response and the threat-test trigger conditions respectively of the pieces of threat-context data (Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats. The assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. The assessment engine computes a second threat probability score for each of the one or more sub-types of threats by the one or more nano classifiers. In an embodiment, the one or more nano classifiers may be selected based on the aforementioned comparison and a plurality of predefined policies associated with the user. The assessment engine detects the one or more sub-types of the threats in the data based on a comparison of the second threat probability score with a predefined threshold value of the second threat probability score), where the one of the predefined threat categories thus selected corresponds to one of the threat-test trigger conditions which the scenario-related response involves (Figure 5 and paragraphs 94-95 disclose the detected sub type of threats are determined based an intent and category of the prompt and content based on the attributes of the data (trigger conditions)); using the threat identification model to obtain one of the pieces of threat-context data that corresponds to the one of the predefined threat categories thus selected from the threat-context dataset (Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats), and to generate a data-generation instruction based on the set of test templates included in the one of the pieces of threat-context data thus obtained (paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. The assessment engine computes a second threat probability score for each of the one or more sub-types of threats by the one or more nano classifiers. In an embodiment, the one or more nano classifiers may be selected based on the aforementioned comparison and a plurality of predefined policies (test templates). The assessment engine detects the one or more sub-types of the threats in the data based on a comparison of the second threat probability score with a predefined threshold value of the second threat probability score. Paragraphs 96 and 99 further reveals on detection of subtypes of threats, the data is selectively moderated based on predefined rules corresponding to each of the one or more sub-type of threats to obtain a moderated data, and the moderated data is used to obtain an improved/modified prompt instructions); and using the attack model, based on the data-generation instruction and the vulnerability-context dataset, to generate the red-teaming data (paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. Paragraph 99 discloses the approved/moderated prompt and content is transmitted to the generative AI model, wherein the generative AI model generate output data based on the moderated/approved prompt and attribute and context data [the context data is the vulnerability-context dataset]). Paragraph 47 discloses examples of attribute and context data wherein examples of such attributes may include but are not limited to, nature of the input data, nature of the output data, usage history and context associated with the data, and similarity with past violations and threats).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s system of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data, Blair’s teachings of trigger conditions, and Krishnamurthy’s vulnerability-context dataset with Ahmed’s threat identification model to dynamically mitigating threats of a generative AI model (paragraph 2 of Ahmed).
As to claim 14, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein said processor (Apple: abstract discloses system and method predicting and remediating malware threats in an electronic computer network. Paragraph 87 discloses processor is operable to execute instructions) uses the attack model to generate the red teaming data by: using the attack model to generate the red-teaming data includes: using the attack model to select one of the predefined vulnerability categories based on the scenario-related response and the attack-test trigger conditions respectively of the pieces of vulnerability-context data, where the one of the predefined vulnerability categories thus selected corresponds to one of the attack-test conditions which the scenario-related response involves ( Ahmed: Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats. Paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. Paragraph 99 discloses the approved/moderated prompt and content is transmitted to the generative AI model, wherein the generative AI model generate output data based on the moderated/approved prompt and attribute and context data [the context data is the vulnerability-context dataset]). Paragraph 47 discloses examples of attribute and context data wherein examples of such attributes may include but are not limited to, nature of the input data, nature of the output data, usage history and context associated with the data, and similarity with past violations and threats); using the attack model to obtain one of the pieces of vulnerability-context data that corresponds to the one of the predefined attack-test conditions from the vulnerability-context dataset ( Ahmed: Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats. Paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data); and using the attack model to generate the red-teaming data based on the data- generation instruction and the at least one set of adversarial attack templates included in the one of the pieces of vulnerability-context data thus obtained ( Ahmed: Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect(thus identify/select) one or more sub-type of threats associated with the one or more type of threats. Paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. Paragraph 99 discloses the approved/moderated prompt and content is transmitted to the generative AI model, wherein the generative AI model generate output data based on the moderated/approved prompt and attribute and context data [the context data is the vulnerability-context dataset]). Paragraph 47 discloses examples of attribute and context data wherein examples of such attributes may include but are not limited to, nature of the input data, nature of the output data, usage history and context associated with the data, and similarity with past violations and threats). The motivation is similar to the motivation presented in claim 13.
As to claim 15, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein, for each of the predefined vulnerability categories, the attack-test trigger condition of the piece of vulnerability-context data that corresponds to the predefined vulnerability category includes plural attack success rates (ASRs) of adversarial attacks respectively against different types of generative models by targeting the vulnerabilities that belong to the predefined vulnerability category, wherein said processor (Ahmed: paragraph 44-45 disclose the system includes one or more processors) uses the attack model to select one of the predefined vulnerability categories includes determining one of the types of generative models indicated by the scenario-related response, and for each of the each of the predefined vulnerability categories (Apple: abstract discloses system and method predicting and remediating malware threats in an electronic computer network. Paragraph 87 discloses processor is operable to execute instructions. Ahmed: Figure 5, reference number 302 and paragraphs 94-95 disclose the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data (trigger conditions) to detect threats associated with the prompt and content (content can be the scenario related response). The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect (thus identify/select) one or more sub-type of threats associated with the one or more type of threats. Paragraph 95 discloses the assessment engine dynamically configures the threat detection model by selecting one or more nano classifiers from amongst a plurality of nano classifiers selectively trained to detect the one or more sub-type of threats in the data. Paragraph 99 discloses the approved/moderated prompt and content is transmitted to the generative AI model, wherein the generative AI model generate output data based on the moderated/approved prompt and attribute and context data [the context data is the vulnerability-context dataset]). Paragraph 47 discloses examples of attribute and context data wherein examples of such attributes may include but are not limited to, nature of the input data, nature of the output data, usage history and context associated with the data, and similarity with past violations and threats), determining whether the ASR that is included in the piece of vulnerability-context data corresponding to the predefined vulnerability category and that corresponds to the one of the types of generative models thus determined is greater than a predetermined threshold value (Ahmed: paragraph 59 discloses when detecting the value of the computed second threat probability score greater than or equal to the predefined threshold score, the system may flag the data as containing the specific subtype of threat. For example, as in previous example, the toxicity second threat probability score of 0.85 exceeds the predefined threshold score of 0.7, leading the system to flag the data as containing a toxicity threat. Paragraph 94 further discloses the assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect one or more sub-type of threats associated with the one or more type of threats), and selecting the predefined vulnerability category in response to determining that the ASR is greater than the predetermined threshold value (Ahmed: paragraph 59 discloses when detecting the value of the computed second threat probability score greater than or equal to the predefined threshold score, the system may flag the data as containing the specific subtype of threat. For example, as in previous example, the toxicity second threat probability score of 0.85 exceeds the predefined threshold score of 0.7, leading the system to flag the data as containing a toxicity threat. Paragraph 94 further discloses the assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect one or more sub-type of threats associated with the one or more type of threats). The motivation is similar to the motivation presented in claim 13.
As to claim 16, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein said storage further stores an evaluation-criterion dataset (Ahmed: Figure 2 reveals the memory 204 holds instructions for assessment engine 208), the evaluation-criterion dataset including plural evaluation criteria that are related respectively to plural predefined assessment items corresponding respectively to the predefined threat categories, each of the evaluation criteria including an evaluation standard for assessing the corresponding one of the predefined assessment items (Ahmed: paragraph 94 discloses the assessment engine may determine an intent and category of the prompt and content based on the attributes of the data to detect threats associated with the prompt and content. The threats may include a prompt injection check, a jailbreak check, a profanity and toxicity check, a Personal Identifiable Information (PII) check, an Intellectual Property (IP) violation check, and an organization policy and a role-based check. The assessment engine may determine a first threat probability score corresponding to each of the types of threats for the prompt and content, based on the attributes of the data. The assessment engine may compare the first threat probability score with a predefined threshold score for the each of the types of threats. In case, the first threat probability score of the one or more types of threats is determined to be greater than the predefined threshold score, the assessment engine may further dynamically configure a threat detection model to detect one or more sub-type of threats associated with the one or more type of threats), said processor (Ahmed: paragraphs 44-45 disclose the system includes one or more processors) sends the red-teaming data to the target generative model for the target generative model to generate a to-be-evaluated response (Dai: Figure 4A reveals input data from GenAI Red Teaming system is provided to the GenAI model, and the output is provided to the GenAI assessment system. Ahmed: Figure 3 further discloses that the output response from the generative AI model is provide to the assessment engine to generate an evaluated response such as identify type of threat, see paragraph 99); retrieving the to-be-evaluated response from the target generative model (Dai: Figure 4A reveals input data from GenAI Red Teaming system is provided to the GenAI model, and the output is provided to the GenAI assessment system. Ahmed: Figure 3 further discloses that the output response from the generative AI model is provide to the assessment engine to generate an evaluated response such as identify type of threat, see paragraph 99); and generating an evaluation result based on the to-be-evaluated response and the evaluation-criteria dataset, the evaluation result indicating whether the target generative model has trustworthiness (Ahmed: paragraph 99 discloses the assessment engine associated with the output sub-layer may perform various operations for mitigating threat types of the output response, if present. The assessment engine may analyze the output response to determine an intent and a category of the output response based on the attributes of the output response. The assessment engine may identify one or more sub types of threats associated with the output response based on the analysis of the output response. The one or more sub types of threats may include the profanity and toxicity check, a third-party Intellectual Property (IP) violation check, the organization policy and role-based check, and a hallucination check. Paragraph 100 discloses The assessment engine may further determine, in real-time, a threat score corresponding to each of the sub types of threats, for the output response received via the generative AI model, based on the plurality of attributes. Once the threat score is determined for each of the sub types of threats, the threat probability score may be compared with a predefined threshold for the threshold value of the second threat probability score, by the assessment engine. Further, the assessment engine may select one or more sub types of threats, based on the comparison and a plurality of predefined policies associated with the user. Figure 3 and paragraph 102 disclose the validation engine outputs a validation/evaluation result based on the output response and the output sub layer check/evaluation criteria. The validation result indicates whether the response should be modified, approved or rejected). The motivation is similar to the motivation presented in claim 13.
As to claim 17, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein said storage further stores at least one evaluation model (Ahmed: Figure 2 reveals the memory 204 holds instructions for assessment engine 208), and said processor (Ahmed: paragraphs 44-45 disclose the system includes one or more processors) generates the evaluation result by: using the at least one evaluation model to analyze the to-be-evaluated response so as to determine whether the to-be-evaluated response meets the evaluation standard of one of the evaluation criteria that corresponds to one of the predefined threat categories, where the one of the predefined threat categories corresponds to the piece of threat-context data used to generate the data- generation instruction (Ahmed: paragraph 99 discloses the assessment engine may analyze the output response to determine an intent and a category of the output response based on the attributes of the output response. The assessment engine may identify one or more sub types of threats associated with the output response based on the analysis of the output response. The one or more sub types of threats may include the profanity and toxicity check, a third-party Intellectual Property (IP) violation check, the organization policy and role-based check, and a hallucination check. The sub types of threats corresponds to attributes data and intent which can be threat context data used to generate the data generation instruction/modified response/approved prompt); in response to determining that the to-be-evaluated response meets the evaluation standard, using the at least one evaluation model to generate an evaluation result indicating that the target generative model has trustworthiness (Ahmed: Figure 3 reveals the output from the validation engine such as rejection of the output response, approved response (corresponds to trustworthiness), and modified response based on the check of the type of threat present in the output response. Paragraph 103 discloses if the validation of the output response is successful without moderation, an approved response may be transmitted to the user through the generative AI model. The approved response may be the output response without any moderation. Otherwise, if the validation of the output response is successful after moderation, a modified response may be transmitted to the user through the generative AI model); and in response to determining that the to-be-evaluated response does not meet the evaluation standard, using the at least one evaluation model to generate an evaluation result indicating that the target generative model does not have trustworthiness (Ahmed: Figure 3 reveals the output from the validation engine such as rejection of the output response (corresponds to the target generative model does not have trustworthiness), approved response, and modified response based on the check of the type of threat present in the output response. Paragraph 102 discloses when the validation of the output response is unsuccessful even after performing some predefined iterations of moderation, the output response may be restricted from further processing, and details of restricting the output response (such as a user rejection message) may be transmitted and rendered to the user. When the validation of the prompt and content is unsuccessful even after performing some predefined iterations of moderation, the prompt and content may be restricted from further processing, and details of restricting the prompt and content (such as a user rejection message) may be transmitted and rendered to the user). The motivation is similar to the motivation presented in claim 13.
As to claim 18, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein said at least one evaluation model is a generative pre-trained transformer (Ahmed: paragraphs 17 and 33 reveal the system includes models that are Generative Pretrained Transformers).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s method of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data, Blair’s teachings of trigger conditions, and Krishnamurthy’s vulnerability-context dataset with Ahmed’s generative pre-trained transformer such that the system can generate new and original content that may be used in various red team applications (paragraph 33 of Ahmed).
As to claim 19, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein said storage further stores plural evaluation models (Ahmed: Figure 2 reveals the memory 204 holds instructions for assessment engine 208), and said processor (Ahmed: paragraphs 44-45 disclose the system includes one or more processors) generates the evaluation result by: for each of the evaluation models, using the evaluation model to analyze the to-be-evaluated response so as to determine whether the to-be-evaluated response meets the evaluation standard of one of the evaluation criteria that corresponds to one of the predefined threat categories, the piece of threat-context data corresponding to which is used to generate the data- generation instruction (Ahmed: paragraph 99 discloses the assessment engine may analyze the output response to determine an intent and a category of the output response based on the attributes of the output response. The assessment engine may identify one or more sub types of threats associated with the output response based on the analysis of the output response. The one or more sub types of threats may include the profanity and toxicity check, a third-party Intellectual Property (IP) violation check, the organization policy and role-based check, and a hallucination check. The sub types of threats corresponds to attributes data and intent which can be threat context data used to generate the data generation instruction/modified response/approved prompt), and based on analysis of the to-be-evaluated response, generate a preliminary result indicating whether the target generative model has trustworthiness (Ahmed: Figure 3 reveals the output from the validation engine such as rejection of the output response, approved response (corresponds to trustworthiness), and modified response based on the check of the type of threat present in the output response. Paragraph 103 discloses if the validation of the output response is successful without moderation, an approved response may be transmitted to the user through the generative AI model. The approved response may be the output response without any moderation. Otherwise, if the validation of the output response is successful after moderation, a modified response may be transmitted to the user through the generative AI model); and generating the evaluation result based on the preliminary results that are generated respectively by the evaluation models (Ahmed: paragraph 99 discloses the assessment engine may analyze the output response to determine an intent and a category of the output response based on the attributes of the output response. The assessment engine may identify one or more sub types of threats associated with the output response based on the analysis of the output response. The one or more sub types of threats may include the profanity and toxicity check, a third-party Intellectual Property (IP) violation check, the organization policy and role-based check, and a hallucination check. The sub types of threats corresponds to attributes data and intent which can be threat context data used to generate the data generation instruction/modified response/approved prompt). The motivation is similar to the motivation presented in claim 13.
As to claim 20, the combination of Dai in view of Apple, Blair, Krishnamurthy, and Ahmed teaches wherein each of the threat identification model and the attack model is a generative pre-trained transformer (Ahmed: paragraphs 17 and 33 reveal the system includes models that are Generative Pretrained Transformers).
It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to modify Dai’s method of generating red-teaming data in view of Apple’s teachings of the concept of a threat anticipation method implemented by a processor and storing threat data, Blair’s teachings of trigger conditions, and Krishnamurthy’s vulnerability-context dataset with Ahmed’s generative pre-trained transformer such that the system can generate new and original content that may be used in various red team applications (paragraph 33 of Ahmed).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Miron et al US 20250260708 (hereinafter Miron).
Miron teaches sending a reconnaissance prompt to the target generative model for the target generative model to generate a scenario-related response based on the reconnaissance prompt (paragraphs 98-99); retrieving the scenario-related response from the target generative model (paragraphs 99-100); and generating the red-teaming data (paragraphs 99 and 101) as recited in claim 1.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FELICIA FARROW whose telephone number is (571)272-1856. The examiner can normally be reached M - F 7:30am-4:00pm (EST).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexander Lagor can be reached at (571)270-5143. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/F.F/ Examiner, Art Unit 2437
/BENJAMIN E LANIER/ Primary Examiner, Art Unit 2437