Detailed Action
This Office Action is in response to the remarks entered on 10/08/2025. Claims 2 and 16 have been canceled. New claims 22-32 are added. Claims 1, 3-15, and 17-32 are currently pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 10/08/2025 and 01/28/2026 were filed after the mailing date of the Final Office Action on 08/08/2025. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claim Objections have been withdrawn.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 8-12 and 28-29 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 8 recites “wherein the intermediate result is captured by a proxy of the GenAI model …” It is unclear what constitutes ‘a proxy of the GenAI model’. It is unclear what tasks the proxy performs or what its output data is. Does it classify the input data to detect a prompt seek to cause the GenAI model to behave in an undesired manner?
For purpose of the examination, the examiner interprets the claim to mean: The proxy of the GenAI model is equivalent to the classifiers.
Claims 9-12 depend from Claim 8 and inherit the same deficiency.
Claim 28 recites “wherein the activations are captured by a proxy of the GenAI model that executes in a monitoring environment …” it is unclear what constitutes ‘a proxy of the GenAI model’. It is unclear what tasks the proxy performs or what its output data is. Does it classify the input data to detect a prompt seek to cause the GenAI model to behave in an undesired manner?
For purpose of the examination, the examiner interprets the claim to mean: The proxy of the GenAI model is equivalent to the classifiers.
Claim 29 recites “wherein the proxy is a quantized version of the GenAI model having a same topology and producing intermediate residual streams compatible with the GenAI model, and wherein quantization preserves access to residual stream activations via the runtime hooks.”
First, the claim recites ‘the proxy’. There is insufficient antecedent basis for this limitation in the claim.
Second, it is unclear what constitutes ‘the proxy’. It is unclear what tasks the proxy performs or what its output data is. Does it classify the input data to detect a prompt seek to cause the GenAI model to behave in an undesired manner?
Lastly, it is unclear whether the ‘a quantized version of the GenAI model’ and ‘and wherein quantization preserves’ are results of the same quantization process, as ‘wherein quantization…’ lacks an article ‘a’ or ‘the’. The examiner suggests the applicant to amend the claim to recite ‘wherein the quantization …’
For purpose of the examination, the examiner interprets the claim to mean: The proxy of the GenAI model is equivalent to the classifiers.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 32 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
The claim does not fall within at least one of the four categories of patent eligible subject matter because Claim 32 is a system claim with no hardware in the claim. Claim 32 recites “a computing system comprising: a model computing environment and a computing environment” which encompasses software elements under the broadest reasonable interpretation. However, nowhere in the specification is the “a computing system comprising: a model computing environment and a computing environment” defined as excluding purely software instantiations of the computing environment and system. Therefore, claim 32 is non-statutory. Examiner recommends amending the claim to recite the hardware on which the model runs explicitly.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 14, 21-25 and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Goswami (US 20190238568 A1, hereinafter ‘Goswami’) in view of Shi et al. (“PL-Transformer: a POS-aware and layer ensemble transformer for text classification”, Published online 2022, hereinafter ‘Shi’).
Regarding claim 1, Goswami teaches:
A computer-implemented method comprising: receiving, by a computing system, data [Goswami, 0108, Fig. 9] discloses operating the trained DNN 820 on new data)
capturing, during inference of the comprising activations in layers of the [Goswami, 0108, Fig. 9] The new network activation data 920 which is the means for the intermediate layers of the DNN 820 is input to the SVL classifier 870. And then, the SVL classifier 870 which is trained to detect distortion in the data is used to classify the data. The DNN is interpreted as the GenAI model, and it is more clearly disclosed by the secondary reference Shi)
determining, based on the intermediate result and using a machine learning-based classifier trained using a policy that maps intermediate results to prohibited action by the GenAI model, whether the prompt seeks to cause the GenAI model to behave in an undesired manner; ([Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is distorted or undistorted. [0110-0111] collectively discloses the training process of the SVM classifier which performs selective dropout logic to selectively dropout filters (i.e., policy that maps result to action). [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system) and
providing, by the computing system over a network, data characterizing the determination to a consuming application or process, the consuming application or process (i) allowing an output of the GenAI model to be provided to a requestor when it is determined that the prompt does not seek to cause the GenAI model to behave in an undesired manner, and (ii) initiates at least one remediation action preventing the GenAI model from behaving in an undesired manner when it is determined that the prompt seeks to cause the GenAI model to behave in an undesired manner. ([Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is distorted or undistorted. [0122] discloses if the input image data is determined to not be adversarial, the image processing is performed (i.e., allowing the output to be provided). Otherwise, in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system, or dropout the adversarial image data (i.e., remediation action))
Goswami does not specifically disclose:
receiving, by a computing system, data characterizing a prompt for ingestion by a generative artificial intelligence (GenAI) model, the GenAI model having an input layer, an output layer, and a plurality of intermediate transformer layers positioned immediately after the input layer and immediately before the output layer;
inputting, by the computing system, the prompt into the GenAI model;
capturing, during inference of the GenAI model and using runtime hooks applied to one or more transformer layers of the GenAI model, an intermediate result comprising activations in residual streams derived from one or more of the intermediate transformer layers of the GenAI model;
Shi teaches:
receiving, by a computing system, data characterizing a prompt for ingestion by a generative artificial intelligence (GenAI) model, the GenAI model having an input layer, an output layer, and a plurality of intermediate transformer layers positioned immediately after the input layer and immediately before the output layer; ([Shi, page 1974, Fig. 1] and [Shi, page 1975, left col, 3.2 Embedding layer, lines 1-4] show that the embedding layer obtains the input sequence of sentence representation, which are stream input. Fig.1 further shows that the model having an embedding layer (i.e., input layer), an output layer (i.e., softmax layer), and a plurality of intermediate transformer layers (i.e., PL transformer layer with a plurality of C-Encoder layers). The transformer layer is the GenAI model)
inputting, by the computing system, the prompt into the GenAI model; ([Shi, page 1974, Fig. 1] and [Shi, page 1975, left col, 3.2 Embedding layer, lines 1-4] shows that the embedding layer obtains the input sequence of sentence representation, which are stream input)
capturing, during inference of the GenAI model and using runtime hooks applied to one or more transformer layers of the GenAI model, an intermediate result comprising activations in residual streams derived from one or more of the intermediate transformer layers of the GenAI model; ([Shi, page 1975, left col, 3.2 Embedding layer, lines 1-4] shows that the embedding layer obtains the input sequence of sentence representation, which are stream input. [Shi, page 1974, Fig. 1] and [page 1976, left col, 3.3.2 Layer ensemble mechanisms, line 1 – right col, line 13] collectively discloses capturing PL-Transformer intermediate layers (C-Encoder Layers, which are interpreted as transformer layers) using hooks, and then ensembles the captured layer outputs h1(1)h2(1)…hn(1) and h1(2)h2(2)…hn(2) … h1(l)…hn(l). The transformer layer outputs are combined and sent to Linear Layer and Softmax to generate the final text classification output)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami and Shi to use a method of capturing an intermediate result of the transformer model of Shi to implement the anomaly detection method of Goswami. The suggestion and/or motivation for doing so is to improve the performance of the model by providing more minority class data from the trained latent space.
Regarding claim 3, Goswami in view of Shi teaches:
wherein the GenAI model comprises a plurality of transformer layers and the intermediate result comprises per-layer activations captured from residual streams generated by one or more of the transformer layers via the runtime hooks during a forward pass. ([Shi, page 1974, Fig. 1] and [page 1976, left col, 3.3.2 Layer ensemble mechanisms, line 1 – right col, line 13] collectively discloses capturing PL-Transformer intermediate layers (C-Encoder Layers, which are interpreted as transformer layers) using hooks, and then ensembles the captured layer outputs h1(1)h2(1)…hn(1) and h1(2)h2(2)…hn(2) … h1(l)…hn(l). The transformer layer outputs are combined and sent to Linear Layer and Softmax to generate the final text classification output)
Regarding claim 14, Goswami teaches:
wherein a policy maps intermediate residual stream activations to one or more prohibited actions ([Goswami, 0108] The SVL classifiers 870 is trained to detect adversarial input image data. [Goswami, 0113 and 0115] collectively disclose that the method is not limited to image data and the input data may be textual data (residual stream)).
Regarding claim 21, Goswami teaches:
A computer-implemented method comprising: receiving, by a monitoring computing environment from a model computing environment executing a , ; ([Goswami, 0108, Fig. 9] The new network activation data 920 which is the means for the intermediate layers of the DNN 820 is generated in response to the new input data and input to the SVL classifier 870. The DNN is interpreted as the GenAI, and the GenAI is further disclosed in Shi. And then, the SVL classifier 870 which is trained to detect distortion in the data is used to classify the data)
determining, by the monitoring computing environment and based on the activations and using a machine learning-based classifier trained using a policy mapping residual stream activations to prohibited actions, whether the prompt seeks to cause the [Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is distorted or undistorted. [0109] discloses the pre-training process of the SVM classifier 870 using the undistorted data)
returning, automatically by the monitoring computing environment in response to a determination the the prompt does not elicit undesired actions by the [Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is an distorted or undistorted. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system)
initiating, automatically by the monitoring computing environment in response to a determination that the prompt elicits undesired actions by the GenAI model, at least one remediation action preventing the output as generated by the GenAI model from being returned to the requestor. ([Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is an distorted or undistorted. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system, or dropout the adversarial image data)
Goswami does not specifically disclose:
receiving, from a model computing environment executing a generative artificial intelligence (GenAI) model, activations in residual streams between transformer layers of the GenAI model generated during inference of the GenAI model in response to a prompt, the activations being captured via runtime hooks during a forward pass;
Shi teaches:
receiving, from a model computing environment executing a generative artificial intelligence (GenAI) model, activations in residual streams between transformer layers of the GenAI model generated during inference of the GenAI model in response to a prompt, the activations being captured via runtime hooks during a forward pass; ([Shi, page 1975, left col, 3.2 Embedding layer, lines 1-4] shows that the embedding layer obtains the input sequence of sentence representation, which are stream input. [Shi, page 1974, Fig. 1] and [page 1976, left col, 3.3.2 Layer ensemble mechanisms, line 1 – right col, line 13] collectively discloses capturing PL-Transformer intermediate layers (C-Encoder Layers, which are interpreted as transformer layers) using hooks, and then ensembles the captured layer outputs h1(1)h2(1)…hn(1) and h1(2)h2(2)…hn(2) … h1(l)…hn(l). The transformer layer outputs are combined and sent to Linear Layer and Softmax to generate the final text classification output. The Transformer Layer is the GenAI model)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami and Shi to use a method of capturing an intermediate result of the transformer model of Shi to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the performance of the model by providing more minority class data from the trained latent space.
Regarding claim 22, Goswami in view of Shi teaches:
The method of claim 21, wherein capturing the activations comprises applying runtime hooks to every transformer layer and, after completion of a forward pass responsive to the prompt, querying each hook to retrieve activations from residual streams and storing the retrieved activations for the prompt in a single array of floating-point values. ([Shi, page 1974, Fig. 1] and [page 1976, left col, 3.3.2 Layer ensemble mechanisms, line 1 – right col, line 13] collectively discloses capturing PL-Transformer intermediate layers (C-Encoder Layers, which are interpreted as transformer layers) using hooks, and then ensembles the captured layer outputs h1(1)h2(1)…hn(1) and h1(2)h2(2)…hn(2) … h1(l)…hn(l). The transformer layer outputs are combined and sent to Linear Layer and Softmax to generate the final text classification output)
Regarding claim 23, Goswami teaches:
The method of claim 21, wherein determining whether the [Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is an distorted or undistorted. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system, or dropout the adversarial image data. [Claims 1 and 6] indicates that the output value from the intermediate layers are vectors (array of numerical data))
Goswami does not specifically disclose:
wherein determining whether the prompt
Shi teaches:
wherein determining (classifying) whether the prompt seeks… ([Shi, page 1974, Fig. 1] and [page 1976, left col, 3.3.2 Layer ensemble mechanisms, line 1 – right col, line 13] collectively discloses capturing PL-Transformer intermediate layers (C-Encoder Layers, which are interpreted as transformer layers) using hooks, and then ensembles the captured layer outputs h1(1)h2(1)…hn(1) and h1(2)h2(2)…hn(2) … h1(l)…hn(l). The transformer layer outputs are combined and sent to Linear Layer and Softmax to generate the final text classification output)
Regarding claim 25, Goswami teaches:
The method of claim 21, wherein a policy maps intermediate results to prohibited actions and specifies a remedial action for at least one prohibited action, and wherein the classifier output is used to select the remedial action. ([Goswami, 0090 and 0092] discloses the selective dropout logic based on the classification by the SVM classifier. The selective dropout selects which portion of the input data to dropout, which corresponds to the selection of remedial action)
Regarding claim 28, Goswami teaches:
The method of claim 21, wherein the activations are captured by a proxy of the GenAI model that executes in a monitoring environment and, based on an output or an intermediate result of the proxy, controls whether the GenAI model ingests the prompt. ([Goswami, 0107; Fig. 8 and Fig. 9] The detection module 840 comprises layer-wise comparison logic 844 which captures every network activation data 830 and 860 to generate distance metrics for each of the intermediate layers of the DNN 820. [0109] discloses that the classifier 870 discards the input data if the input is detected to be adversarial)
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of Pardeshi (US 20210407051 A1).
Regarding claim 4, Goswami in view of Shi teaches the method of claim 1.
Goswami in view of Shi does not specifically disclose:
wherein the GenAI model comprises a mixture of experts (MoE) model and the intermediate result comprises outputs from at least a subset of experts in the MoE model.
Pardeshi teaches:
wherein the GenAI model comprises a mixture of experts (MoE) model and the intermediate result comprises outputs from at least a subset of experts in the MoE model ([Pardeshi, 0050] The GAN contains a different trained variational autoencoder VAE for each of a set of object classes, and these VAEs, which corresponds to the subject of experts, are considered experts for different objects classes in a mixture of experts-based approach. [Pardeshi, 0051] An appropriate VAE is selected based on outputs of the VAEs (intermediate result).).
Before the effective filing date of the invention to a person of ordinary skill in the art, it would
have been obvious, having the teachings of Goswami, Shi and Pardeshi to use the mixture of experts (MoE) model of Pardeshi to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the performance of the machine learning system by utilizing machine learning models specialized for different tasks (i.e., different types of subjects, responses, etc.).
Claims 5-7 and 26 are rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of Kuan (US 20230027149 A1, hereinafter ‘Kuan’).
Regarding claim 5, Goswami in view of Shi teaches the method of claim 1.
Goswami in view of Shi does not specifically disclose wherein the consuming application or process flags the prompt based on the determination.
Kuan teaches:
wherein the consuming application or process flags the prompt based on the determination. ([Kuan, 0076] Malicious data 704 is identified by the anomaly prediction system 114. The identification process corresponds to the ‘flags’, as the flags provided to the prompt merely identifies which data contains anomaly)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would
have been obvious, having the teachings of Goswami, Shi and Kuan to use the method of flagging the prompt based on the determination of Kuan to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the performance of the machine learning system by allowing the machine learning model to quickly determine which input data is determined to contain malicious data.
Regarding claim 6, Goswami in view of Shi teaches the method of claim 1.
Goswami in view of Shi does not specifically disclose:
wherein the consuming application or process automatically blocks, at a network gateway, an internet protocol (IP) address of a requester of the prompt based on the determination.
Kuan teaches:
wherein the consuming application or process automatically blocks, at a network gateway, an internet protocol (IP) address of a requester of the prompt based on the determination. ([Kuan, 0080] IP addresses associated with the malicious data 704 identified by the anomaly prediction system 114 is blocked by the control action 708 performed by the system)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would
have been obvious, having the teachings of Goswami, Shi and Kuan to use the method of blocking the IP address of a requester based on the determination of Kuan to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the security of the machine learning system by blocking the prompt from the potentially malicious IP address.
Regarding claim 7, Goswami in view of Shi teaches the method of claim 1.
Goswami in view of Shi does not specifically disclose:
wherein the consuming application or process causes subsequent prompts from an internet protocol (IP) address of a requester of the prompt to be programmatically modified based on the determination and causes the modified prompt to be ingested by the GenAI model.
Kuan teaches:
wherein the consuming application or process causes subsequent prompts from an internet protocol (IP) address of a requester of the prompt to be programmatically modified based on the determination and causes the modified prompt to be ingested by the GenAI model. ([Kuan, 0076] Malicious data 704 is identified by the anomaly prediction system 114. [Kuan, 0088] The system removes the corresponding malicious data 704 from the queue, and not issue any control action against the malicious data again. Removal of the malicious data is interpreted as the process of modifying prompt based on the determination.)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would
have been obvious, having the teachings of Goswami, Shi and Kuan to use the method of blocking the IP address of a requester and modifies the prompt based on the determination of Kuan to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the security of the machine learning system by blocking the prompt from the potentially malicious IP address.
Regarding claim 26, Goswami in view of Shi and further in view of Kuan teaches:
The method of claim 21, Kuan teaches:
wherein the consuming application or process, based on the determination, causes external remediation resources to block, at a network gateway, subsequent requests from an internet protocol (IP) address of a requester of the prompt. ([Kuan, 0080] IP addresses associated with the malicious data 704 identified by the anomaly prediction system 114 is blocked by the control action 708 performed by the system)
Claims 8, 12-13 and 29 are rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of Elthakeb et al. (“Divide and Conquer: Leveraging Intermediate Feature Representations for Quantized Training of Neural Networks”, 2020, hereinafter ‘Elthakeb’).
Regarding claim 8, Goswami in view of Shi teaches:
wherein the intermediate result is captured by a proxy of the GenAI model … transformer layers in the GenAI model ([Shi, page 1975, left col, 3.2 Embedding layer, lines 1-4] shows that the embedding layer obtains the input sequence of sentence representation, which are stream input. [Shi, page 1974, Fig. 1] and [page 1976, left col, 3.3.2 Layer ensemble mechanisms, line 1 – right col, line 13] collectively discloses capturing PL-Transformer intermediate layers (C-Encoder Layers, which are interpreted as transformer layers) using hooks, and then ensembles the captured layer outputs h1(1)h2(1)…hn(1) and h1(2)h2(2)…hn(2) … h1(l)…hn(l). The transformer layer outputs are combined and sent to Linear Layer and Softmax to generate the final text classification output)
However, Goswami in view of Shi does not specifically disclose:
wherein the intermediate result is captured by a proxy of the GenAI model that executes on distinct computing hardware and mirrors a structure of the transformer layers in the GenAI model.
Elthakeb teaches:
wherein the that executes on distinct computing hardware and mirrors a structure of the ([Elthakeb, page 3, 2.1. Matching Activations for Intermediate Layers, lines 1-23], [Figure 3] collectively discloses that the quantized models have the same topology and the produce intermediate streams compatible with the full precision network (the GenAI model). [page 3, 2.2. Splitting, Training and Merging, lines 1-4] discloses that the models are trained in isolation and in parallel)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and Elthakeb to use the method of quantizing the generative model of Elthakeb to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to increase the performance of the system, as quantizing a model makes running inference faster and cheaper by reducing the file size of model weight files.
Regarding claim 12, Goswami in view of Shi in view of Elthakeb teaches:
The method of claim 8, wherein the proxy is a quantized version of the GenAI model having a same topology and producing intermediate residual streams compatible with the GenAI model. ([Elthakeb, page 3, 2.1. Matching Activations for Intermediate Layers, lines 1-23], [Figure 3] collectively discloses that the quantized models have the same topology and the produce intermediate streams compatible with the full precision network (the GenAI model))
Regarding claim 13, Goswami in view of Shi and further in view of Elthakeb teaches:
The method of claim 1 further comprising: quantizing the GenAI model prior to capturing the intermediate result. ([Elthakeb, page 3, 2.1. Matching Activations for Intermediate Layers, lines 1-23], [Figure 3] collectively discloses that the quantized models have the same topology and the produce intermediate streams compatible with the full precision network (the GenAI model))
Regarding claim 29, Goswami in view of Shi teaches:
The method of claim 21.
However, Goswami in view of Shi does not specifically disclose:
wherein the proxy is a quantized version of the GenAI model having a same topology and producing intermediate residual streams compatible with the GenAI model, and wherein quantization preserves access to residual stream activations via the runtime hooks.
Elthakeb teaches:
wherein the proxy is a quantized version of the GenAI model having a same topology and producing intermediate residual streams compatible with the GenAI model, and wherein quantization preserves access to residual stream activations via the runtime hooks. ([Elthakeb, page 3, 2.1. Matching Activations for Intermediate Layers, lines 1-23], [Figure 3] collectively discloses that the quantized models have the same topology and the produce intermediate streams compatible with the full precision network (the GenAI model))
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and Elthakeb to use a method of utilizing a quantized version of the GenAI model having a same topology with the GenAI model of Elthakeb to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the efficiency of the model by offering additional speedup through enabling parallelization [Elthakeb, ABSTRACT].
Claims 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi in view of Elthakeb and further in view of Spalt (US 20200409323 A1).
Regarding claim 9, Goswami in view of Shi and further in view of Elthakeb teaches the method of claim 8.
Goswami in view of Shi and further in view of Elthakeb does not specifically disclose:
wherein the consuming application or process prevents the prompt from being input into the GenAI model based on the determination.
Spalt teaches:
wherein the consuming application or process prevents the prompt from being input into the GenAI model based on the determination ([Spalt, 0100] The matrix generator determines that whether the input matrix can be input to the generator machine learning model. Based on the result of the determination, incomplete input matrix (anomalous prompt) is blocked.).
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi, Elthakeb, and Spalt to use the method of blocking the prompt based on the determination of Spalt to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the security of the machine learning system by blocking the prompt from the potentially malicious IP address.
Regarding claim 10, Goswami in view of Shi and further in view of Elthakeb and further in view of Spalt teaches:
wherein the consuming application or process allows the prompt to be input into the GenAI model based on the determination ([Spalt, 0100] The matrix generator determines that whether the input matrix can be input to the generator machine learning model. Based on the result of the determination, incomplete input matrix (anomalous prompt) is blocked. However, sometimes the input matrix generator 212 can pass the incomplete input matrix to the machine learning model generator 214 (allows the prompt to be input into the generative AI model).).
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi in view of Elthakeb and further in view of Kuan.
Regarding claim 11, Goswami in view of Shi and further in view of Elthakeb does not specifically disclose:
wherein the consuming application or process modifies the prompt based on the determination and causes the modified prompt to be ingested by the GenAI model
Kuan teaches:
wherein the consuming application or process modifies the prompt based on the determination and causes the modified prompt to be ingested by the GenAI model ([Kuan, 0076] Malicious data 704 is identified by the anomaly prediction system 114. [Kuan, 0088] The system removes the corresponding malicious data 704 from the queue, and not issue any control action against the malicious data again. Removal of the malicious data is interpreted as the process of modifying prompt based on the determination.).
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi, Elthakeb and Kuan to use the method of blocking the IP address of a requester and modifies the prompt based on the determination of Kuan to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the security of the machine learning system by blocking the prompt from the potentially malicious IP address.
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of MOZO VELASCO (US 20230049479 A1).
Regarding claim 15, Goswami in view of Shi teaches:
The method of claim 1.
However, Goswami in view of Shi does not specifically disclose:
wherein the one or more prohibited actions include seeking one or more of: a response in a non-approved spoken or written language, computer code, sensitive information, a response to an encrypted prompt, a prompt containing two or modalities, or a prompt in a first modality obfuscating information in a second modality.
MOZO VELASCO teaches:
wherein the one or more prohibited actions include seeking one or more of: a response in a non-approved spoken or written language, computer code, sensitive information, a response to an encrypted prompt, a prompt containing two or modalities, or a prompt in a first modality obfuscating information in a second modality ([MOZO VELASCO, 0080] The generator is trained to generate data without any risk of including sensitive data.).
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and MOZO VELASCO to use the method of blocking the prohibited actions includes a response in a sensitive information of MOZO VELASCO to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the security of the machine learning system by blocking the prompt from the potentially malicious IP address.
Claims 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of Brodie (US 20240005690 A1).
Regarding claim 17, Goswami in view of Shi teaches the method of claim 1.
Goswami in view of Shi does not specifically disclose wherein the GenAI model is a state-based text analysis model.
Brodie teaches:
wherein the GenAI model is a state-based text analysis model ([Brodie, 0035] The neural network can include a deep neural network, a convolutional neural network, a recurrent neural network (e.g., an LSTM), a graph neural network, a transformer neural network, or a generative adversarial neural network. Upon training, such a neural network may become a large language model. The instant specification [0055] explains that the state-based text analysis models include LSTMs)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and Brodie to use a large language model which includes LSTMs of Brodie to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the performance of the model in long-term dependencies in sequential prompt data.
Regarding claim 18, Goswami in view of Shi and further in view of Brodie teaches:
wherein the state-based text analysis model comprises: a large language model. ([Brodie, 0035] The neural network can include a deep neural network, a convolutional neural network, a recurrent neural network (e.g., an LSTM), a graph neural network, a transformer neural network, or a generative adversarial neural network. Upon training, such a neural network may become a large language model. The instant specification [0055] explains that the state-based text analysis models include LSTMs)
Regarding claim 19, Goswami in view of Shi and further in view of Brodie teaches:
wherein the state-based text analysis model comprises: a long short-term memory model (LSTM) ([Brodie, 0035] The neural network can include a deep neural network, a convolutional neural network, a recurrent neural network (e.g., an LSTM), a graph neural network, a transformer neural network, or a generative adversarial neural network. Upon training, such a neural network may become a large language model. The instant specification [0055] explains that the state-based text analysis models include LSTMs).
Claims 20 is rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi in view of O’CONNOR et al. (US 20220092404 A1, hereinafter ‘O’CONNOR’) and further in view of Elthakeb.
Regarding claim 20, Goswami teaches:
A computer-implemented method comprising:
receiving, by a computing system, data c[Goswami, 0108, Fig. 9] The new network activation data 920 which is the means for the intermediate layers of the DNN 820 is generated in response to the new input data and input to the SVL classifier 870. And then, the SVL classifier 870 which is trained to detect distortion in the data is used to classify the data)
capturing, during inference of the comprising activations in residual streams of one or more of the intermediate layers of the [Goswami, 0108, Fig. 9] The new network activation data 920 which is the means for the intermediate layers of the DNN 820 is input to the SVL classifier 870. And then, the SVL classifier 870 which is trained to detect distortion in the data is used to classify the data)
determining, using a machine learning-based classifier trained using a policy that maps intermediate residual stream activations to acceptable and undesired actions and based on the intermediate result, whether the prompt elicits undesired actions by the GenAI model, the classifier being trained using a data set generated according to a policy which maps acceptable and undesired actions to intermediate results; and ([Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is distorted or undistorted. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system)
returning, automatically by the computing system in response to a determination the the prompt does not elicit undesired actions by the GenAI model, an output of the GenAI model to a requestor; and ([Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is distorted or undistorted. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system, or dropout the adversarial image data)
initiating, automatically by the computing system in response to a determination that the prompt elicits undesired actions by the GenAI model, at least one remediation action preventing the output as generated by the GenAI model from being returned to the requestor. ([Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is distorted or undistorted. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system, or dropout the adversarial image data)
Goswami does not specifically disclose:
receiving, by a computing system, data characterizing a prompt
redirecting, by the computing system, the prompt from the GenAI model to a quantized version of the GenAI model having an input layer, an output layer, and a plurality of intermediate transformer layers positioned immediately after the input layer and immediately before the output layer, the quantized version of the GenAI model executing on separate computing resources and being different than the GenAI model;
capturing, during inference of the quantized version of the GenAI model, an intermediate result from residual streams derived from one or more of the intermediate transformer layers of the quantized version of the GenAI model;
Shi teaches:
receiving, by a computing system, data characterizing a prompt ([Shi, page 1974, Fig. 1] and [Shi, page 1975, left col, 3.2 Embedding layer, lines 1-4] show that the embedding layer obtains the input sequence of sentence representation, which are stream input. Fig.1 further shows that the model having an embedding layer (i.e., input layer), an output layer (i.e., softmax layer), and a plurality of intermediate transformer layers (i.e., PL transformer layer with a plurality of C-Encoder layers))
capturing, during inference of the transformer layers of the [Shi, page 1975, left col, 3.2 Embedding layer, lines 1-4] shows that the embedding layer obtains the input sequence of sentence representation, which are stream input. [Shi, page 1974, Fig. 1] and [page 1976, left col, 3.3.2 Layer ensemble mechanisms, line 1 – right col, line 13] collectively discloses capturing PL-Transformer intermediate layers (C-Encoder Layers, which are interpreted as transformer layers) using hooks, and then ensembles the captured layer outputs h1(1)h2(1)…hn(1) and h1(2)h2(2)…hn(2) … h1(l)…hn(l). The transformer layer outputs are combined and sent to Linear Layer and Softmax to generate the final text classification output)
Goswami in view of Shi does not specifically disclose:
redirecting, by the computing system, the prompt from the GenAI model to a quantized version of the GenAI model having an input layer, an output layer, and a plurality of intermediate transformer layers positioned immediately after the input layer and immediately before the output layer, the quantized version of the GenAI model executing on separate computing resources and being different than the GenAI model;
O’CONNOR teaches:
redirecting, by the computing system, the prompt from the GenAI model to a quantized version of the GenAI model having an input layer, an output layer ([O’CONNOR, 0011], [0041] and [0043-0046] The student model 106 are interpreted as the ‘quantized version’ of the teacher model as it is trained based on the teacher model 105. [O’CONNOR, Fig. 4 and Fig. 7] The teacher network 105 receives input data 103 (Step S220), processes the input data 103 (Step S230), and then transmits the input data to the student neural network 106 (Step S200). [0051] shows that the input data are prompts (words spoken by speaker))
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and O’CONNOR to use the method of redirecting the output of the GenAI model (teacher model) to a quantized model (student model) of O’CONNOR to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to increase the efficiency and performance of the quantized model by updating the simplified quantized (student) model based on the original GenAI model (teacher model) with higher complexity.
However, Goswami in view of Shi and further in view of Li does not specifically disclose:
the quantized version of the GenAI model executing on separate computing resources
Elthakeb teaches:
the quantized version of the GenAI model executing on separate computing resources ([Elthakeb, page 3, 2.1. Matching Activations for Intermediate Layers, lines 1-23], [Figure 3] collectively discloses that the quantized models have the same topology and the produce intermediate streams compatible with the full precision network (the GenAI model). [page 3, 2.2. Splitting, Training and Merging, lines 1-4] discloses that the models are trained in isolation and in parallel)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi, O’CONNOR and Elthakeb to use the method of quantizing the generative model of Elthakeb to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to increase the performance of the system, as quantizing a model makes running inference faster and cheaper by reducing the file size of model weight files.
Claim 24 is rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of Clement et al. (US 20240386103 A1, hereinafter ‘Clement’).
Regarding claim 24, Goswami in view of Shi teaches:
The method of claim 21.
However, Goswami in view of Shi does not specifically disclose:
wherein the classifier is a multi-class classifier executed by an analysis engine in a monitoring environment and configured to output a label identifying exactly one prohibited action selected from: a response in a non-approved spoken or written language or computer code.
Clement teaches:
wherein the classifier is a multi-class classifier executed by an analysis engine in a monitoring environment and configured to output a label identifying exactly one prohibited action selected from: a response in a non-approved spoken or written language or computer code. ([Clement, 0048 and Fig. 4 block 422-424] The security agent determines (i.e., classifies) whether the secret (i.e., data characterizing the determination) is included in the response. If the response has secret (a prompt that seeks to cause the GenAI model to behave in a desired manner), the secret is extracted, if not (a prompt seeks to cause the GenAI model to behave in an undesired manner), then an error message is returned back to the user application)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and Clement to use a multi-class classifier outputting a label identifying prohibited action selected from: a response in a non-approved spoken or written language or computer code of Clement to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the security of the model by preventing the malicious prompt to be ingested by the system or model.
Claim 27 is rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of Rand et al. (US 20070204341 A1, hereinafter ‘Rand’)
Regarding claim 27, Goswami teaches:
The method of claim 21, wherein the consuming application or process, based on the determination, causes a local remediation engine in a model environment to [Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is distorted or undistorted. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system, or dropout the adversarial image data)
However, Goswami in view of Shi does not specifically disclose:
delete portions of subsequent prompts from an internet protocol (IP) address of a requester of the prompt and to cause a modified prompt to be ingested by the GenAI model
Rand teaches:
delete portions of subsequent prompts from an internet protocol (IP) address of a requester of the prompt and to cause a modified prompt to be ingested by the [Rand, 0034] discloses modifying malicious or prohibited activities from specific IP addresses to be ingested by the system)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and Rand to use a method of modifying prompts from an IP address to be ingested by the system of Rand to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the security of the model by preventing the malicious prompt to be ingested by the system or model.
Claim 30 is rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of Schmidtler et al. (US 20160335435 A1, hereinafter ‘Schmidtler’)
Regarding claim 30, Goswami in view of Shi teaches:
The method of claim 21, [Goswami, 0107; Fig. 8 and Fig. 9] The detection module 840 comprises layer-wise comparison logic 844 which captures every network activation data 830 and 860 to generate distance metrics for each of the intermediate layers of the DNN 820. [0109] discloses that the classifier 870 discards the input data if the input is detected to be adversarial. [0110] discloses that the SVM classifier is trained using adversarial data and based on the responses from intermediate layers)
However, Goswami in view of Shi does not specifically disclose:
wherein the policy used to train the classifier defines that encrypted instructions or instructions to encrypt an output constitute a prohibited action
Schmidtler teaches:
wherein the policy used to train the classifier defines that encrypted instructions or instructions to encrypt an output constitute a prohibited action ([Schmidtler, 0009] discloses that the classifier is trained from a collection of data comprising known malicious files, and the classifier is trained such that it can handle encrypted and/or compressed files without decrypting it)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and Schmidtler to use a method of train the classifier defines that encrypted instructions or instructions to encrypt of Schmidtler to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the security of the model by preventing attackers from accessing the output data and/or input data.
Claim 31 is rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of Cai et al. (US 20060026675 A1, hereinafter ‘Cai’).
Regarding claim 31, Goswami teaches:
The method of claim 21, wherein the policy used to train the classifier [Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is distorted or undistorted. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system, or dropout the adversarial image data)
Goswami in view of Shi does not specifically disclose:
the classifier defines detection of obfuscation patterns including one or more of base64, hexadecimal, and rot13 encodings
Cai teaches:
the classifier defines detection of obfuscation patterns including one or more of base64, hexadecimal, and rot13 encodings ([Cai, 0028] discloses classifying malicious hexadecimal files and extracting single byte patterns using a SVM classifier)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and Schmidtler to use a method of detecting obfuscation patterns using the classifier of Cai to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to improve the security of the model by enabling the classifier to capture malicious patterns in various types of input data.
Claim 32 is rejected under 35 U.S.C. 103 as being unpatentable over Goswami in view of Shi and further in view of Elthakeb.
Regarding claim 32, Goswami teaches:
A computing system comprising: a model computing environment executing a [Goswami, 0108, Fig. 9] The new network activation data 920 which is the means for the intermediate layers of the DNN 820 is generated in response to the new input data and input to the SVL classifier 870. The computing environment in which the DNN and the SVL classifier run is the model computing environment)
[Goswami, 0108, Fig. 9] The new network activation data 920 which is the means for the intermediate layers of the DNN 820 is generated in response to the new input data and input to the SVL classifier 870. And then, the SVL classifier 870 which is trained to detect distortion in the data is used to classify the data)
an analysis engine configured to assemble the activations into an array of numerical data having a consistent shape and to apply a machine learning-based classifier trained using a policy mapping intermediate results to prohibited actions comprising at least a response in a non-approved spoken or written language and computer code; and ([Goswami, 0108, Fig. 9] The new network activation data 920 which is the means for the intermediate layers of the DNN 820 is generated in response to the new input data and input to the SVL classifier 870. And then, the SVL classifier 870 which is trained to detect distortion in the data is used to classify the data. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system)
a remediation engine configured, in response to a classifier output indicating a prohibited action, to block, at a network gateway, transmission of an output of the GenAI model to a requestor. ([Goswami, Fig. 9 and 0108] discloses generating an output 930 indicating whether or not the SVM classifier 870 has determined that the new input data 910 is distorted or undistorted. [0122] discloses in response to the output of the SVM classifier, a responsive operation may be performed such as alerting contingent systems or operating or shutting down the downstream computing system)
Goswami does not specifically disclose:
a model computing environment executing a generative artificial intelligence (GenAI) model having an input layer, an output layer, and a plurality of intermediate transformer layers;
a monitoring computing environment distinct from the model computing environment; wherein the monitoring computing environment comprises:
a proxy of the GenAI model executing in the monitoring computing environment, the proxy comprising a quantized version of the GenAI model having a same topology and producing intermediate residual streams compatible with the GenAI model;
runtime hooks applied to one or more transformer layers of the proxy and configured to capture, during a forward pass responsive to a prompt, activations in residual streams;
Shi teaches:
a model computing environment executing a generative artificial intelligence (GenAI) model having an input layer, an output layer, and a plurality of intermediate transformer layers; ([Shi, page 1974, Fig. 1] and [Shi, page 1975, left col, 3.2 Embedding layer, lines 1-4] show that the embedding layer obtains the input sequence of sentence representation, which are stream input. Fig.1 further shows that the model having an embedding layer (i.e., input layer), an output layer (i.e., softmax layer), and a plurality of intermediate transformer layers (i.e., PL transformer layer with a plurality of C-Encoder layers). The transformer layer is the GenAI model)
runtime hooks applied to one or more transformer layers of the proxy and configured to capture, during a forward pass responsive to a prompt, activations in residual streams; ([Shi, page 1975, left col, 3.2 Embedding layer, lines 1-4] shows that the embedding layer obtains the input sequence of sentence representation, which are stream input. [Shi, page 1974, Fig. 1] and [page 1976, left col, 3.3.2 Layer ensemble mechanisms, line 1 – right col, line 13] collectively discloses capturing PL-Transformer intermediate layers (C-Encoder Layers, which are interpreted as transformer layers) using hooks, and then ensembles the captured layer outputs h1(1)h2(1)…hn(1) and h1(2)h2(2)…hn(2) … h1(l)…hn(l). The transformer layer outputs are combined and sent to Linear Layer and Softmax to generate the final text classification output)
Goswami in view of Shi does not specifically disclose:
a monitoring computing environment distinct from the model computing environment; wherein the monitoring computing environment comprises:
a proxy of the GenAI model executing in the monitoring computing environment, the proxy comprising a quantized version of the GenAI model having a same topology and producing intermediate residual streams compatible with the GenAI model;
Elthakeb teaches:
a monitoring computing environment distinct from the model computing environment; wherein the monitoring computing environment comprises: ([Elthakeb, page 3, 2.1. Matching Activations for Intermediate Layers, lines 1-23], [Figure 3] collectively discloses that the quantized models have the same topology and the produce intermediate streams compatible with the full precision network (the GenAI model). [page 3, 2.2. Splitting, Training and Merging, lines 1-4] discloses that the models are trained in isolation and in parallel)
a proxy of the GenAI model executing in the monitoring computing environment, the proxy comprising a quantized version of the GenAI model having a same topology and producing intermediate residual streams compatible with the GenAI model; ([Elthakeb, page 3, 2.1. Matching Activations for Intermediate Layers, lines 1-23], [Figure 3] collectively discloses that the quantized models have the same topology and the produce intermediate streams compatible with the full precision network (the GenAI model))
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Goswami, Shi and Elthakeb to use the method of quantizing the generative model of Elthakeb to implement the machine learning method of Goswami. The suggestion and/or motivation for doing so is to increase the performance of the system, as quantizing a model makes running inference faster and cheaper by reducing the file size of model weight files.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claim 20 is rejected on the ground of nonstatutory double patenting as being obvious under claim 16 of US 12572777 in view of Elthakeb et al. (“Divide and Conquer: Leveraging Intermediate Feature Representations for Quantized Training of Neural Networks”, 2020)
Present application 18/669,366
Reference patent US 12572777
20
A computer-implemented method comprising: receiving, by a computing system, data characterizing a prompt for ingestion by a generative artificial intelligence (GenAI);
16
A computer-implemented method for policy-based gating of multimodal prompts in a generative artificial intelligence (GenAI) system, the method comprising: receiving, by a proxy computing system, data characterizing a multimodal prompt for ingestion by a production GenAI model, the GenAI model comprising a transformer architecture with a plurality of layers including at least one intermediate layer and residual streams;
redirecting, by the computing system, the prompt from the GenAI model to a quantized version of the GenAI model having an input layer, an output layer, and a plurality of intermediate transformer layers positioned immediately after the input layer and immediately before the output layer,
redirecting the multimodal prompt to a quantized proxy model that shares a layer topology with the production GenAI model, the quantized proxy model configured to process the prompt without generating a final output;
capturing, during inference of the quantized version of the GenAI model and using runtime hooks, an intermediate result comprising activations in residual streams of one or more of the intermediate transformer layers of the quantized version of the GenAI model;
capturing, by the proxy computing system and via forward-pass hooks applied to at least one transformer layer of the quantized proxy model, per-token activation vectors from residual streams at a predetermined intermediate layer, wherein the capturing occurs after encoding of the prompt and before any output is generated by the production GenAI model;
determining, using a machine learning-based classifier trained using a policy that maps intermediate residual stream activations to acceptable and undesired actions and based on the intermediate result, whether the prompt elicits undesired actions by the GenAI model, the classifier being trained using a data set generated according to a policy which maps acceptable and undesired actions to intermediate results;
applying, by the proxy computing system, a classifier trained using a policy-labeled dataset that maps internal activation patterns to a plurality of enumerated prohibited actions, the prohibited actions comprising at least: (i) generation of computer code, (ii) response in a non-approved language, (iii) response to an encrypted prompt, or (iv) cross-modal obfuscation;
returning, automatically by the computing system in response to a determination that the prompt does not elicit undesired actions by the GenAI model, an output of the GenAI model to a requestor; and
determining, by the proxy computing system and based on the per-token activation vectors, whether the prompt is likely to elicit a prohibited action by the production GenAI model; and
initiating, automatically by the computing system in response to a determination that the prompt elicits undesired actions by the GenAI model, at least one remediation action preventing the output as generated by the GenAI model from being returned to the requestor.
in response to the determination, gating the prompt by preventing the prompt from being ingested by the production GenAI model when the determination indicates a prohibited action, and allowing ingestion otherwise, such that the production GenAI model does not process prompts likely to elicit prohibited actions.
However, the reference patent US 12572777 does not suggest or disclose:
the quantized version of the GenAI model executing on separate computing resources and being different than the GenAI model;
Elthakeb teaches:
the quantized version of the GenAI model executing on separate computing resources and being different than the GenAI model; ([Elthakeb, page 3, 2.1. Matching Activations for Intermediate Layers, lines 1-23], [Figure 3] collectively discloses that the quantized models have the same topology and the produce intermediate streams compatible with the full precision network (the GenAI model). [page 3, 2.2. Splitting, Training and Merging, lines 1-4] discloses that the models are trained in isolation and in parallel)
It would have been obvious before the effective filing date of the claimed invention to a person
having ordinary skill in the art to apply the method of running the proxy of the GenAI model on distinct computing hardware by Elthakeb to improve the performance of the machine learning method of the present invention. The suggestion and/or motivation for doing so is to improve the efficiency of the method by offering additional speedup through enabling parallelization [Elthakeb, ABSTRACT].
Response to Arguments
Claim Objections
Claim Objections have been withdrawn.
Response to Arguments under 35 U.S.C. 101
35 U.S.C. 101 rejections based on MPEP 2106.04(a) Abstract Ideas have been withdrawn.
Response to Arguments under 35 U.S.C. 103
Examiner respectfully disagrees. Applicant’s arguments with respect to claim(s) 1, 3-15 and 17-32 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
New Claims
New Claims are rejected under 35 U.S.C. 103 as being unpatentable over Goswami, Shi, Clement, Rand, Schmidtler, and Cai. See 35 U.S.C. 103 rejection above.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUN KWON whose telephone number is (571)272-2072. The examiner can normally be reached Monday – Friday 7:30AM – 4:30PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JUN KWON/Examiner, Art Unit 2127
/ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127