Prosecution Insights
Last updated: October 02, 2026
Application No. 18/545,804

GUIDANCE SIGNALS FOR ACCELERATING INFERENCING IN GENERATIVE ARTIFICIAL INTELLIGENCE MODELS

Non-Final OA §103
Filed
Dec 19, 2023
Priority
Jul 13, 2023 — provisional 63/513,488
Examiner
DORVIL, RICHEMOND
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Qualcomm Incorporated
OA Round
3 (Non-Final)
36%
Grant Probability
At Risk
3-4
OA Rounds
8m
Est. Remaining
68%
With Interview

Examiner Intelligence

Grants only 36% of cases
36%
Career Allowance Rate
22 granted / 61 resolved
-25.9% vs TC avg
Strong +32% interview lift
Without
With
+32.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
15 currently pending
Career history
81
Total Applications
across all art units

Statute-Specific Performance

§101
12.2%
-27.8% vs TC avg
§103
55.0%
+15.0% vs TC avg
§102
11.7%
-28.3% vs TC avg
§112
16.5%
-23.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 61 resolved cases

Office Action

§103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 6, 10, 11, 19, 24, 28, 29, 37, and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Horesh et al. (U.S. Patent No. 11,972,333) in view of Kim et al. (U.S. Patent Publication 2024/0311405). Concerning independent claims 1 and 19, Horesh et al. discloses a system and method for generative artificial intelligence, comprising: “at least one memory having executable instructions stored thereon; and one or more processors configured to execute the executable instructions to cause the processing system to:” – system 100 includes a processor 130 and a memory 135 coupled to processor 130 (column 5, lines 12 to 24: Figure 1); processor 130 may be capable of executing scripts or instructions of one or more software programs stored in memory 135 (column 6, lines 30 to 33: Figure 1); “generate, based on an input query and a first generative artificial intelligence model, a sequence of tokens corresponding to a candidate response to the input query” – first generative AI model 170 generates a first output based on a user input provided via a prompt (column 8, lines 45 to 50: Figure 1); model input 202 is input provided to a generative AI model to prompt the generative AI model to generate an output answering the input; system 100 may provide an input prompt to a user to provide a natural language input to be received by first generative AI model 210 (column 12, lines 47 to 52: Figure 2); BERT model 240 is trained to tokenize first output 212 to generate a first output vector 242 of tokens and to tokenize second output 222 to generate a second output vector 244 of tokens (“a sequence of tokens corresponding to a candidate response”) (column 13, lines 48 to 57: Figure 2); through management of another generative AI model, a number of irrelevant answers to a query may be reduced (column 24, lines 16 to 27: Figure 1); model input 202 of natural language input from a user, then, can be "an input query"; “output, to a second generative artificial intelligence model trained to generate guidance signals, the sequence of tokens and the input query for verification” – second generative AI model is trained using fresher training data (Abstract); managing a generative AI model includes using a second generative AI model to generate outputs from similar inputs and comparing the output of the generative AI models to determine their similarity (Abstract); second generative AI model is trained on a second training set and at least a portion of the second training set is subsequent in time to the first training set; a method also includes generating a second output by the second generative AI model based on the input (“trained to generate guidance signals”) (column 1, line 67 to column 2, line 5); computing system is configured to receive a first output of the first generative AI model based on the input; provide the input to a second generative AI model with the second generative AI model trained on a second training set and at least a portion of the second training set being subsequent in time to the first training set (“trained to generate guidance signals”); generate a second output by the second generative AI model based on the input (column 4, lines 21 to 30); a system is implemented to manage a generative AI model to optimize the performance of the model including improving the relevancy of outputs generated by the model; to ensure that the outputs generated and provided by a generative AI model are relevant, a second generative AI model is implemented to generate outputs based on the same prompts to the generative AI model being managed, and the outputs from the managed generative AI model and the second generative AI model are compared to determine their similarity; a second generative AI model is trained using fresher data than the managed generative AI model so that outputs from the second generative AI model may be more relevant (“trained to generate guidance signals”); to ensure that outputs generated and provided by a generative AI model are relevant (“for verification”), a second generative model is implemented to generate outputs based on the same prompts to the generative AI model being managed (column 4, lines 53 to 66); second generative AI model 140 is implemented for managing the first generative AI model 170; second generative AI model 140 is trained to generate an output to compare to an output from the first generative AI model 170 to determine if the first model’s output is relevant (“trained to generate guidance signals”) (column 7, line 65 to column 8, line 6: Figure 1); “receive, from the second generative artificial intelligence model, one or more guidance signals for the generated sequence of tokens” – an output of the second generative AI model is compared to an output of the managed generative AI model by a classification model to determine a relevance of the output from the managed generative Al model, and to perform suitable policies to optimize performance of the managed generative AI model, e.g., providing alternative outputs or preventing providing the output (Abstract); NN classifier 250 generates similarity indication 252 indicating the similarity between the first output and the second output; a similarity indication 252 may be used to execute various policies for managing first generative AI model 210 (column 14, lines 47 to 51: Figure 2); policy engine 270 receives identification signal 262 or similarity signal 252 regarding the similarity of first output 212 and second output 222 (column 15, lines 21 to 27: Figure 2); second generative AI model 220 then, provides “one or more first guidance signals" for managing output by first generative AI model 210; that is, second generative AI model 220 includes policy engine 270 to manage (‘guide’) first generative AI model 210 by ‘signaling’ first generative AI model 210 whether to provide alternative outputs or prevent providing the outputs; equivalently, it is a nature of supervision by a second generative AI model in managing a first generative AI model that a first generative AI model is ‘guided’ by a second generative AI model; “revise the candidate response to the input query based on the generated sequence of tokens and the one or more first guidance signals” – first generative model 170 is the model to be managed from system 100; managing the generative AI model may include any action to manage the outputs of the generative AI model including adjusting ('revising') the output to be provided or providing an alternative output (column 7, lines 1 to 10: Figure 1); an output of a classification model is used to perform various suitable policies to optimize the performance of the managed generative AI model, e.g. providing alternate outputs or prevent providing the outputs (column 1, lines 54 to 58); policy engine 160 may instruct system 100 to provide the input again to first generative AI model 170, and first generative AI model 170 generates a third output that is different than the first output generated previously based on the input (column 10, lines 54 to 59: Figure 1); here, providing an alternate output for a response is equivalent to “revise the candidate response”; “output the revised candidate response as a response to the input query” –interface 110 may provide outputs from one or more generative AI models including second comparison as a second similarity threshold being greater than a defined threshold, policy engine 160 may instruct system 100 to output the third output for use to a user (column 11, lines 6 to 11: Figure 1); policy engine 270 may instruct system 100 to output first output 212 for use in response to identifying that first output 212 is similar to second output 222, or policy engine 270 may instruct system 100 to output second output 222 from second generative Al model 220 (column 15, lines 39 to 53: Figure 2). Concerning independent claims 1 and 19, Horesh et al. discloses a supervisory system for generative AI model that includes a first generative AI model 170 that is a managed model and a second model 140 that is a supervisory model, and that supervisory model 140 and managed model 170 are trained. The second supervisory model 140 then generates an output to determine if output from managed model 170 should be revised or output as it is. It is maintained that supervisory model 140 is equivalently sending “a guidance signal” to managed model 170, and that supervisory model 140 is “trained to generate guidance signals” because it is equivalently a nature of supervision to guide an entity that is being managed or supervised. Horesh et al., then, is maintained to disclose all of the limitations of these independent claims with the exception of “wherein the first generative artificial intelligence model is a smaller version of the second artificial intelligence model”. If second generative AI model 140 is trained on a different set of training data to supervise and manage first generative AI model 170, then second generative AI model 140 is “trained to generate guidance signals”. That is, supervision or management by second generative AI model 140 ‘guides’ output by first generative AI model 170. However, Horesh et al. does not disclose anything about the relevant sizes of the first and second generative AI models 140 and 170 in the limitation of “wherein the first generative artificial intelligence model is a smaller version of the second generative artificial intelligence model. Instead, Horesh et al. describes an embodiment of a plurality of generative AI models being trained on various time ranges of years from ten years of training data. (Column 20, Line 54 to Column 21, Line 53: Figure 5) Still, if a second generative AI model is trained on more recent years of training data that a first generative AI model, then second generative AI model is able to supervise and manage output of first generative AI model with data that is more accurate from the standpoint if it being more up-to-date. Concerning independent claims 1 and 19, Kim et al. teaches dynamic selection from among multiple candidate generative models with different computational efficiencies. (Abstract) Many generative models can be of a very large size including billions of parameters and smaller size counterparts can be separately trained with less parameters or pruned and/or quantized by applying one or more pruning techniques and/or one or more quantization techniques. Due to the large size of a generative model, there can be significant resource utilization and latency, but smaller size models can be less robust and/or less accurate than their larger size counterparts. (¶[0002] - ¶[0003]) An implementation can dynamically select between at least a smaller LLM and a larger LLM on a request-by-request basis to achieve reduced latency and/or improved computational efficiency while mitigating occurrences of any inaccurate and/or under-specified responses. (¶[0007]) Selection between smaller and larger LLMs can be made based on first and second measures that characterize a probability of generating a correct response to a request, where a first measure characterizes generating a correct response using the smaller LLM and a second measure characterizes generating a correct response using the larger LLM. (¶[0008]) A trained machine learning (ML) model can be used in selecting from among multiple candidate generative models, e.g., from between at least a smaller LLM and a larger LLM. (¶[0012]) Candidate generative models 150 can include LLM 150A with less than 100 billion parameters, LLM 150B with between 100 billion and 250 billion parameters, and LLM 150C with over 250 billion parameters. (¶[0027]: Figure 1) Selection engine 126 utilizes request features to select which of multiple candidate generative models 150 should be utilized in responding to a request. (¶[0042]: Figure 1) Kim et al., then, teaches “wherein the first generative artificial intelligence model is a smaller version of the second generative artificial intelligence model” because LLM 150A is smaller than LLM 150B or LLM 150C. An objective is to reduce latency and/or conserve computational resources to mitigate occurrences of a generated response being inaccurate or under-specified. (¶[0004) It would have been obvious to one having ordinary skill in the art to provide a first generative artificial intelligence model that is a smaller version of a second artificial intelligence model as taught by Kim et al. as first generative AI model and second generative AI model of Horesh et al. for a purpose of conserving computational resources and mitigating occurrences of a generated response being inaccurate. Concerning claims 6 and 24, Horesh et al. discloses that policy engine 160 instructs system 100 to generate a third output that is different from first output by applying first input again to first generative AI model 170. (Column 10, Lines 41 to 59: Figure 2) Policy engine 160, then, provides “a guidance signal” to first generative AI model 170 based on output from second generative AI model 140. Concerning claims 10 and 28, Horesh et al. discloses that an input 202 is provided to a second generative artificial intelligence model 220 in Figure 2, and this input may represent a query. Consequently, an output of classification model may manage a first artificial intelligence model to provide alternative outputs, prevent providing the output, adjusting the output, and retraining the model (Abstract; Column 1, Lines 54 to 58; Column 7, Lines 1 to 10: Figure 1) Broadly, providing alternative outputs, preventing providing the output, adjusting the output, and retraining the model are “a list of actions to be performed by the first generative artificial intelligence model to generate the candidate response”. Concerning claims 11 and 29, Horesh et al. discloses first generative AI model 170 and second generative AI model 140, and that a system may host first generative AI model if the model is not included in system 100; an externally developed generative AI model may be hosted by a system associated with the model’s developer and a user may interface with the system 100 in order to use the externally developed generative model (column 5, lines 37 to 58: Figure 1). Generally, Horesh et al. discloses that at least one of a first and second generative AI model may be hosted externally (“on a system remote from the processing system”), and it would be an obvious expedient to host one of the first and second generative AI models locally and one remotely under principles of distributed processing in a client/server architecture. Concerning claims 37 and 40, Kim et al. teaches generative models can be of a very large size including billions of parameters, and smaller size counterparts can be separately trained with less parameters or pruned and/or quantized by applying one or more pruning techniques and/or one or more quantization techniques. (¶[0003]) A smaller LLM can be a quantized and/or pruned version of the larger LLM (“wherein the first generative artificial intelligence model is a pruned version of the second generative artificial intelligence model”). (¶[0005]) Claims 2 to 5 and 20 to 23 are rejected under 35 U.S.C. 103 as being unpatentable over Horesh et al. (U.S. Patent No. 11,972,333) in view of Kim et al. (U.S. Patent Publication 2024/0311405) as applied to claims 1 and 19 above, and further in view of Madaan et al. (U.S. Patent Publication 2022/0261535). Concerning claims 2, 4, 20, and 22, Horesh et al. discloses generating tokens for output and providing an alternative output or adjusting an output of a first generative artificial intelligence model, but does not expressly “identify a token location within the candidate response at which at least one incorrect token is to be replaced” “or replacing “one incorrect token”. That is, Horesh et al. generally discloses adjusting a response, but is not specifically directed to replacing individual tokens at a location in an adjusted response. However, Madaan et al. teaches automatically modifying responses from generative models by replacing at least a portion of proposed words by at least a portion of alternate words. (Abstract; ¶[0002]) Step 308 includes modifying at least a portion of the words proposed by the at least one automated conversation exchange software program in connection with the at least one conversation by replacing at least a portion of the one or more identified words with at least a portion of one or more alternate words. (¶[0024]: Figure 3: Step 308) Implicitly, words that are being replaced correspond to “tokens”, and a word that is being replaced in a portion implicitly has some specific “location” as a portion in a response of Madaan et al. An objective is to improve text generation in automated conversation exchange software programs, or chatbots, to reduce human biases. (¶[0001]) It would have been obvious to one having ordinary skill in the art to generate an alternate response in Horesh et al. by identifying a location of a portion of a token to be replaced as taught by Madaan et al. for a purpose of improving text generation in chatbots to reduce issues of human bias. Concerning claims 3 and 21, Horesh et al. discloses an iterative procedure of policy engine 160 instructing first generative AI model 170 to generate a third output that is different than the first output previously based on the input (“generate a second candidate response”), that second generative AI model 140 then generates a new output to compare to the third output (“output, to the second generative artificial intelligence model, tokens corresponding to the second candidate response”), and that if the outputs are similar during the second comparison to determine if the second output is relevant (“receive one or more second guidance signals for the tokens corresponding to the second candidate response”). Then policy engine 160 may instruct system 100 to output the third output (“generate a third candidate response based on the tokens corresponding to the second candidate response and the one or more second guidance signals”). (Column 10, Line 41 to Column 11, Line 11: Figure 1) Policy engine 270 may instruct system 100 to perform any number of iterations of checking the outputs of first generative AI model 210, and may generate a new output to model input 202 a defined number of iterations. (Column 15, Line 54 to Column 16, Line 15: Figure 2) Concerning claims 5 and 23, Madaan et al. teaches modifying responses from a generative model using artificial intelligence to reduce human bias in conversations with a chatbot, e.g., as directed to nationality. (¶[0001]) Bias-neural BERT embeddings are used to predict if a word belongs to a predetermined category, e.g., an age-related category, a gender-related category, a race-related category, and a nationality-related category. (¶[0021]) Madaan et al. replaces words having predetermined categories of bias with words that are more “semantically acceptable” because the replacement words reduce the bias of the words. That is, replacement words are more “semantically acceptable” because they have more acceptable meanings with lesser bias. Claims 7 to 8 and 25 to 26 are rejected under 35 U.S.C. 103 as being unpatentable over Horesh et al. (U.S. Patent No. 11,972,333) in view of Kim et al. (U.S. Patent Publication 2024/0311405) as applied to claims 1 and 19 above, and further in view of Zorn et al. (U.S. Patent Publication 2023/0418815). Horesh et al. discloses “instructing the first artificial intelligence model to generate the revised candidate response according to one or more rules.” A policy engine instructs a first generative artificial intelligence model to provide an alternate response based on a similarity and a defined number of iterations, and preventing all of the generated outputs if a new output is not sufficiently similar. (Column 15, Lines 54 to Column 16, Line 15: Figure 2) That is, “one or more rules” can be construed as determining to revise output or prevent output based on a similarity and a predetermined number of iterations. However, Horesh et al. does not disclose characteristics of a guidance prompt as “one or more structured grammar commands” or “a natural language command”. Still, Zorn et al. teaches that it is known that prompts to artificial intelligence models may be generated in natural language or in a structured grammar. Specifically, Zorn et al. teaches that conventional large language models can receive natural language text and generate an appropriate response using a prompt in the form of a natural language description of what imperative code could be able to do. (¶[0002]) A task prompt may be expressed using natural language, but this need not be the case. The task prompt could be expressed using a particular language for expressing the task prompt. The task prompt may be expressed in natural language or some query language including Structured Query Language (SQL). (¶[0038]) Here, structured query language may be construed as “one or more structured grammar commands”. Zorn et al., then, teaches that a prompt to a large language model can be provided as “one or more structured grammar commands”, e.g., in SQL, or as “a natural language command”. An objective is to generate declarative code to expand a utility of language models to aid in the generation of additional declarative code. (¶[0019]) It would have been obvious to one having ordinary skill in the art to generate an alternate response in Horesh et al. from a natural language prompt or a structured grammar prompt as taught by Zorn et al. for a purpose of expanding a utility of language models to aid in the generation of declarative code. Claims 9 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Horesh et al. (U.S. Patent No. 11,972,333) in view of Kim et al. (U.S. Patent Publication 2024/0311405) as applied to claims 1 and 19 above, and further in view of Callegari et al. (U.S. Patent Publication 2024/0362422). Horesh et al. discloses that policy engine 270 may instruct system 100 to perform any number of iterations of checking the outputs of first generative AI model 210, and that first generative AI model 210 may be configured to generate a new output to model input 202 for a defined number of iterations. If the defined number of iterations is reached with all of the new outputs identified as not being similar to the second output, policy engine 270 may instruct system 100 to prevent all of the generated outputs from first generative AI model 210 from being output for use. (Column 15, Line 54 to Column 16, Line 15: Figure 2) Horesh et al., then, discloses “to revise the candidate response to the input query, the one or more processors are configured to cause the processing system to determine that a threshold number of revisions have been performed with respect to a response generated by the first generative artificial intelligence model to the input query”. However, Horesh et al. only prevents output if a defined number of iterations is reached, and does not clearly “output the candidate response as the response to the input query based on determining that the threshold number of revisions have been performed.” Still, Callegari et al. teaches revising large language model prompts to generate a revised prompt in response to second input for the large language model to revise the prompt in view of an assessment report, generate a final response to the revised prompt, and output the final response. (Abstract) If a user is not satisfied with the next response generated based on the revised prompt, if a predetermined number of iterations have not been performed, or if the assessment report 34 for a current response has not met a predefined assessment threshold, then the assessment and revision is repeated for at least another iteration. However, if the current response is acceptable by either the user or predefined criteria, then revised response 36 is output to the user. (¶[0018]: Figure 3) Assessment and revision may be iterated once or a number of times, and the revised response of the final iteration is referred to as the final response 56 of the final iteration. (¶[0030]: Figure 5) Callegari et al., then, teaches outputting a final response after a predetermined number of iterations are performed instead of preventing output of the final response in Horesh et al. An objective is to address a problem that a usefulness of a response is greatly influenced by a quality of a prompt and the technical challenge of crafting a right prompt in order for a large language model to respond with a level of detail and precision that a user desires. (¶[0003]) It would have been obvious to one having ordinary skill in the art to output a revised candidate response based on a threshold number of revisions as taught by Callegari et al. in a generative artificial intelligence model of Horesh et al. for a purpose of enabling a large language model to respond with a level of detail and precision that a user desires. Claims 38 and 41 are rejected under 35 U.S.C. 103 as being unpatentable over Horesh et al. (U.S. Patent No. 11,972,333) in view of Kim et al. (U.S. Patent Publication 2024/0311405) as applied to claims 1 and 19 above, and further in view of Wang et al. (U.S. Patent Publication 2020/0134506). Kim et al. teaches first and second generative LLMs with a smaller LLM that can in one embodiment be derived from a larger LLM by pruning or quantization, but does not provide “wherein the first generative artificial intelligence model and the second generative artificial intelligence model have similar probability distributions.” However, Wang et al. teaches model training of a student model corresponding to a teacher model. (Abstract) Once training of a complex network model is completed, a simplified model may be extracted from the complex model through knowledge distillation that includes forcing the smaller neural network to output the same result. The small neural network is referred to as a ‘student’ model and the large neural network is referred to as a ‘teacher’ model. A difference between output of the teacher model and output of the student model may be indicated by a loss function. Logit loss indicates a difference between probability distributions generated by the teacher model and the student model. (¶[0046] - ¶[0048]) The student model is trained by iteratively decreasing the total loss. (¶[0069]) Wang et al., then, teaches “wherein the first generative artificial intelligence model and the second generative artificial intelligence model have similar probability distributions” because a student model is forced to have a ‘similar’ probability distribution to a teacher model by minimizing a loss function of a logit loss in knowledge distillation. An objective is to increase a robustness of a trained student model without retraining the teacher model by knowledge distillation. (¶[0011]) It would have been obvious to one having ordinary skill in the art to train a smaller student model so that it has a similar probability distribution to a larger teacher model as taught by Wang et al. for the smaller LLM and larger LLM of Kim et al. in order to increase robustness of a student model in generating a same result as a teacher model. Claims 39 and 42 are rejected under 35 U.S.C. 103 as being unpatentable over Horesh et al. (U.S. Patent No. 11,972,333) in view of Kim et al. (U.S. Patent Publication 2024/0311405) as applied to claims 1 and 19 above, and further in view of Sridhar et al. (U.S. Patent Publication 2021/0279595). Kim et al. teaches first and second generative LLMs with a smaller LLM that can in one embodiment be derived from a larger LLM by pruning or quantization, and that a smaller LLM has fewer parameters than a larger LLM, but does not provide “wherein the first generative artificial intelligence model was trained using a smaller dataset than the second generative artificial intelligence model.” However, Sridhar et al. teaches training one or more student-teacher modules as part of teacher neural network training. (Abstract) Knowledge distillation (KD) is a compression technique used to transfer knowledge of a bigger neural network with many learned parameters to a smaller neural network with fewer learned parameters. Output of a larger neural network is a ‘soft target’ used as a supervision signal for training a smaller model called the student sub-network. The student sub-network receives both soft targets and hard targets as supervision signals to enable a student model to be trained using a smaller training dataset (“wherein the first generative artificial intelligence model was trained using a smaller dataset than the second generative artificial intelligence model”), as the soft targets provide higher entropy and less variation, e.g., better generalization, than the hard targets. (¶[0003] - ¶[0004]) Integrated system 300 uses a cascade of supervised signals and provides a soft target for knowledge distillation training of student sub-networks. (¶[0058]: Figure 2B) An objective is to improve feature representations of student sub-networks in distillation frameworks and enable sharing computation to eliminate redundant computation and achieve faster inference. (¶[0027]) It would have been obvious to one having ordinary skill in the art to train a smaller student model with a smaller dataset than a larger teacher model as taught by Sridhar et al. for the smaller LLM and larger LLM of Kim et al. in order to improve feature representations in knowledge distillation and achieve faster inference. Response to Arguments Applicants’ arguments filed on 09 June 2026 are being considered but are not persuasive. Applicants amend independent claims 1 and 19 to include a new limitation of a second generative artificial intelligence model “trained to generate guidance signals”, and presents an argument directed against the prior rejection of these independent claims as being obvious under 35 U.S.C. §103 over Horesh et al. (U.S. Patent No. 11,972,333) in view of Kim et al. (U.S. Patent Publication 2024/0311405). Generally, Applicants argue that this new limitation of a second generative artificial intelligence model “trained to generate guidance signals” is not disclosed by Horesh et al. Specifically, Applicants argue that none of a second generative model, a classification model, or a policy engine is “trained to generate guidance signals”. Applicants appear to admit that a second generative model of Horesh et al. is trained to generate outputs based on a second training set. However, Applicants contend that Horesh et al. is silent with respect to a policy engine being trained at all. This argument is not persuasive and an obviousness rejection is being maintained for independent claims 1 and 19 under 35 U.S.C. §103 over Horesh et al. (U.S. Patent No. 11,972,333) in view of Kim et al. (U.S. Patent Publication 2024/0311405). Generally, it is maintained that Horesh et al. discloses a second generative AI model that is trained and that this second generative AI model equivalently generates “guidance signals” to supervise a first generative AI model. Horesh et al. discloses that a first generative AI model 170 is a managed model and a second generative AI model 140 is a supervisory model. Overall, Horesh et al. is directed to “a supervisory system” to manage a generative AI model, and second generative AI model 140 acts as a supervisor that manages output of first generative AI model 170. Moreover, Horesh et al. includes various descriptions of training second generative AI model 140. Admittedly, Horesh et al. does not expressly use the term “guidance” in describing how second generative AI model 140 evaluates and manages an output from first generative AI model 170. However, supervision and management are simply another way of saying that second generative AI model 140 is ‘guiding’ first generative AI model 170. Conventionally, a supervisor manages or ‘guides’ an action of a worker who is being managed or supervised in accordance with ordinary usage of language. Similarly, it may be true that Horesh et al. provides this supervision, management, or ‘guidance’ in a preferred embodiment by training second generative AI model 140 with data from more recent time ranges to ensure that output is more relevant. Still, managing a first generative AI model 170 by a second generative AI model 140 that is trained on data from more recent time ranges is one way of managing or ‘guiding’ first generative AI model 170 to ensure that output reflects more up-to-date information. Horesh et al.’s Title is “Supervisory Systems For Generative Artificial Intelligence”, and its Abstract states: Systems and methods are disclosed for managing a generative artificial intelligence (AI) model to improve performance. Managing the generative AI model includes using a second generative AI model to generate outputs from similar inputs and comparing the outputs of the generative AI models to determine their similarity. The second generative AI model is trained using fresher training data . . . . (emphasis added) Horesh et al., then, clearly discloses that a second generative AI model is trained to supervise and manage (‘guide’) first generative AI model. Similarly, Horesh et al., at Column 7, Line 65 to Column 8, Line 6, states: The second generative AI model 140 is a generative AI model trained to generate desired content prompted from an input provided to the generative AI model. The second generative AI model 140 is implemented for managing the first generative AI model 170 (which may be included in or external to the system 100). In particular, the second generative AI model 140 is trained to generate an output to compare to an output from the first generative AI model 170 to determine if the first model's output is relevant. (emphasis added) Here, Horesh et al. discloses that a second generative AI model 140 is trained to generate an output for managing a first generative AI model 170. This output for managing first generative AI model 170 is “a guidance signal” from second generative AI model 140, and second generative AI model 140 is trained to generate this output to manage first generative AI model 170. Consequently, Horesh et al. discloses a second generative artificial intelligence model “trained to generate guidance signals”. Applicants’ argument that “none of the second generative model, the classification model, or the policy engine are ‘trained to generate guidance signals’”, then, is not persuasive. Horesh et al. is maintained to equivalently disclose that at least a second generative model, i.e., second generative AI model 140, is trained to generate a supervisory output that manages first generative AI 170. This supervisory managing output is equivalent to “a guidance signal”. That is, supervision or management is equivalent to ‘guidance’ provided by a second model to a first model. Admittedly, Horesh et al. provides a somewhat more complex configuration in Figure 2 that includes a BERT model 240, a classifier 250, and a policy engine 270. Still, Horesh et al. provides an overall functionality that discloses at least all of the limitations of the comprising language of the independent claims. Even if BERT model 240, classifier 250, and policy engine 270 were not described as being trained to generate ‘guidance signals’ in Horesh et al., it is maintained that second generative AI model 140 is disclosed as being trained to generate supervisory output, or ‘guidance signals’. Consequently, Applicants’ arguments are not persuasive that Horesh et al. fails to disclose their new limitation of a second generative artificial intelligence model “trained to generate guidance signals”. Horesh et al. fails to disclose only a limitation of “wherein the first generative artificial intelligence model is a smaller version of the second generative artificial intelligence model”, but Kim et al. teaches that generative models may be of larger and smaller sizes. However, a rationale for a modification can be based on KSR International Co. v. Teleflex Inc. (KSR), 550 U.S. 398, 82 USPQ2d 1385 (2007): (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results or (E) “Obvious to try”, i.e., choosing from a finite number of identified, predictable solutions, with a reasonable expectation of success. Intuitively, a supervisory model should have more knowledge that a managed model, so that a managed model should be smaller than a supervisory model. Kim et al. teaches that generative models may be of larger and smaller sizes, so that it would be obvious to select a larger size for a supervisory model from a finite set of alternatives of a supervisory model being larger or smaller. Similarly, it would be predictable to apply a known technique of providing a larger supervisory model as compared to a size of a managed model so as to improve computation efficiency by only calling upon a larger supervisory model when it is necessary to correct output of a smaller managed model. There would be a predictable advantage to utilizing a larger supervisory model only when necessary to correct output of a smaller managed model to mitigate occurrence of inaccurate responses. Applicants’ arguments are not persuasive. This Office Action is NON-FINAL. Conclusion The prior art made of record and not relied upon is considered pertinent to Applicants’ disclosure. Jin discloses related prior art directed to guidance signals. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARTIN LERNER whose telephone number is (571) 272-7608. The examiner can normally be reached Monday-Thursday 8:30 AM-6:00 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MARTIN LERNER/Primary Examiner Art Unit 2658 July 13, 2026
Read full office action

Prosecution Timeline

Dec 19, 2023
Application Filed
Oct 23, 2025
Non-Final Rejection mailed — §103
Feb 23, 2026
Response Filed
Mar 16, 2026
Final Rejection mailed — §103
May 05, 2026
Response after Non-Final Action
Jun 09, 2026
Request for Continued Examination
Jun 12, 2026
Response after Non-Final Action
Jul 15, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731586
EMOTION DETECTION IN BARGE-IN ANALYSIS
4y 2m to grant Granted Sep 08, 2026
Patent 12718012
Message Display Method and Electronic Device
3y 7m to grant Granted Aug 25, 2026
Patent 12688362
TRANSLATION DEVICE
3y 2m to grant Granted Jul 21, 2026
Patent 12658188
SYSTEM AND METHOD FOR ROBOT INITIATED PERSONALISED CONVERSATION WITH A USER
2y 8m to grant Granted Jun 16, 2026
Patent 12651593
INTENT RECOGNITION METHODS, APPARATUSES, AND DEVICES
3y 1m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
36%
Grant Probability
68%
With Interview (+32.0%)
3y 5m (~8m remaining)
Median Time to Grant
High
PTA Risk
Based on 61 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month