Prosecution Insights
Last updated: August 17, 2026
Application No. 18/966,018

MACHINE LEARNING MODEL WITH CONSTRAINED OUTPUT TOKEN VOCABULARY

Non-Final OA §101§102§103
Filed
Dec 02, 2024
Priority
May 20, 2024 — provisional 63/649,908
Examiner
CASTILLO-TORRES, KEISHA Y
Art Unit
Tech Center
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
84 granted / 113 resolved
+14.3% vs TC avg
Strong +31% interview lift
Without
With
+31.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
29 currently pending
Career history
147
Total Applications
across all art units

Statute-Specific Performance

§101
28.3%
-11.7% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
13.1%
-26.9% vs TC avg
§112
5.7%
-34.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 113 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION Claims 1-20 of the instant application are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 09/18/2025 was filed. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Specification The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim(s) 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. More specifically directed to the abstract idea grouping of: mathematical concept and/or mental process. The independent claim(s) recite(s): 1. A computing system comprising: one or more processing devices configured to: receive a prompt; at a machine learning model that has an output token vocabulary including a plurality of candidate output tokens, compute a plurality of output token probabilities over the output token vocabulary based at least in part on the prompt; at a decoder plugin, compute a constrained output token vocabulary as a proper subset of the output token vocabulary; select one or more output tokens based at least in part on the computed output token probabilities, wherein the one or more output tokens are selected from among the candidate output tokens included in the constrained output token vocabulary; and transmit an output including the one or more output tokens to an additional computing process. 12. A method for use with a computing system, the method comprising: [the limitations as in claim 1, above]. This reads on a human (e.g., mentally and/or using pen and paper): Receiving a prompt from another human (e.g., text, utterance); Using a predetermined set of rules (i.e., model) that comprises an output (e.g., a response) including a plurality of tokens (e.g., words) to compute a probability (i.e., mathematical concept) for each of the tokens (e.g., words) based at least in part on the received text/utterance; Using a predetermined set of rules to compute output tokens; Selecting one or more tokens (e.g., words) based on the probabilities, and Transmitting (e.g., by writing down on a piece of paper) the output including the tokens. 20. A computing system comprising: one or more processing devices configured to: receive a prompt; at a machine learning model that has an output token vocabulary including a plurality of candidate output tokens, compute a plurality of output token probabilities over the output token vocabulary based at least in part on the prompt; at a decoder plugin, modify respective output token probabilities associated with the plurality of candidate output tokens to thereby obtain constrained output token probabilities; select one or more output tokens based at least in part on the constrained output token probabilities; and transmit an output including the one or more output tokens to an additional computing process. This reads on a human (e.g., mentally and/or using pen and paper): Receiving a prompt from another human (e.g., text, utterance); Using a predetermined set of rules (i.e., model) that comprises an output (e.g., a response) including a plurality of tokens (e.g., words) to compute a probability (i.e., mathematical concept) for each of the tokens (e.g., words) based at least in part on the received text/utterance; Using a predetermined set of rules to compute output tokens; Selecting one or more tokens (e.g., words) based on the probabilities, and Transmitting (e.g., by writing down on a piece of paper) the output including the tokens. This judicial exception is not integrated into a practical application because for example: claims 1, 12, and 20 similarly recite: a computing system, processing devices, machine learning model, a decoder plugin, and/or an additional computing process. As an example, in ¶ [0095] of the as filed specification, it is disclosed: Computing system 300 may embody the computing system 10 described above and illustrated in FIG. 1. Components of computing system 300 may be included in one or more personal computers, server computers, tablet computers, home-entertainment computers, network computing devices, video game devices, mobile computing devices, mobile communication devices (e.g., smartphone), and/or other computing devices, and wearable computing devices such as smart wristwatches and head mounted augmented reality devices. Therefore, a general-purpose computer or computing device is described and mainly used as an application thereof. Accordingly, these additional elements do not integrate the abstract idea into a practical idea because it does not impose any meaningful limits on practicing the abstract idea. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional elements of using a computer is listed as a general computing device as noted. The claim is not patent eligible. With respect to claims 2 and 13, the claim(s) recite: 2. The computing system of claim 1, wherein the one or more processing devices are further configured to: based at least in part on the prompt, compute a tokenized prompt including a plurality of input tokens; at the machine learning model, compute the output token probabilities in each of a plurality of autoregressive generation iterations; and at the decoder plugin, iteratively update the constrained output token vocabulary at each of a plurality of the autoregressive generation iterations based at least in part on a context including: the tokenized prompt; and a prior output sequence including one or more prior output tokens computed at prior autoregressive generation iterations. This reads on a human (e.g., mentally and/or using pen and paper): Based on the received prompt (e.g., text or utterance), tokenize the prompt (e.g., segment a sentence into words); Using a predetermined set of rules, compute probabilities for each token/word; Using a predetermined set of rules to compute output tokens for each token/word including: The tokenized prompt Prior tokens No additional limitations are present. With respect to claims 3 and 14, the claim(s) recite: 3. The computing system of claim 2, wherein: the decoder plugin includes an oversight machine learning model; and the one or more processing devices are further configured to: at the oversight machine learning model, compute a predicted classification of the output conditioned on the context; and select the constrained output token vocabulary based at least in part on the predicted classification. This reads on a human (e.g., mentally and/or using pen and paper): Using another predetermined set of rules to compute or determine a classification for the output based on context and selecting an output based on the classification. No additional limitations are present. With respect to claim 4, the claim(s) recite: 4. The computing system of claim 3, wherein, at the oversight machine learning model, the one or more processing devices are further configured to: compute a predicted completion of the context; and compute the predicted classification based at least in part on the predicted completion. This reads on a human (e.g., mentally and/or using pen and paper): wherein the predetermined set of rules comprise a computation of a prediction for completion and wherein the prediction classification is based on the prediction for completion. No additional limitations are present. With respect to claims 5 and 15, the claim(s) recite: 5. The computing system of claim 2, wherein: the decoder plugin includes an oversight machine learning model; and the one or more processing devices are further configured to: at the oversight machine learning model, compute a predicted completion of the context; compute a reward value associated with the predicted completion; and select the constrained output token vocabulary based at least in part on the reward value. This reads on a human (e.g., mentally and/or using pen and paper): wherein the predetermined set of rules comprise a computation of a prediction for completion and computing a reward value (e.g., score – a mathematical concept) with the predicted completion and selecting output based on the reward value (e.g., score – a mathematical concept). No additional limitations are present. With respect to claim 6, the claim(s) recite: 6. The computing system of claim 1, wherein, at the decoder plugin, the one or more processing devices are further configured to modify one or more sampling parameters of the machine learning model. This reads on a human (e.g., mentally and/or using pen and paper): modifying parameters within the predetermined set of rules No additional limitations are present. With respect to claims 7 and 16, the claim(s) recite: 7. The computing system of claim 1, wherein, at the decoder plugin, the one or more processing devices are further configured to: execute a search algorithm over a predefined search domain; and select the constrained output token vocabulary based at least in part on a search result returned by the search algorithm. This reads on a human (e.g., mentally and/or using pen and paper): use a predetermined set of rules (i.e., search algorithm – mathematical concept) and selecting output based on the result. No additional limitations are present. With respect to claim 8, the claim(s) recite: 8. The computing system of claim 7, wherein, when executing the search algorithm at the decoder plugin, the one or more processing devices are configured to perform a Monte Carlo tree search (MCTS) over a plurality of branches of the predefined search domain. This reads on a human (e.g., mentally and/or using pen and paper): using mathematical concepts (i.e., Monte Carlo tree search) No additional limitations are present. With respect to claims 9 and 17, the claim(s) recite: 9. The computing system of claim 1, wherein, at the decoder plugin, the one or more processing devices are configured to specify the constrained output token vocabulary with a regular expression or a context-free grammar. This reads on a human (e.g., mentally and/or using pen and paper): using predetermined set of rules (i.e., specifying output with a regular expression or context-free grammar) No additional limitations are present. With respect to claims 10 and 18, the claim(s) recite: 10. The computing system of claim 1, wherein, at the machine learning model, the one or more processing devices are further configured to: rescale the output token probabilities to obtain constrained output token probabilities normalized over the constrained output token vocabulary; and select the one or more output tokens at least in part by sampling the output tokens from the constrained output token probabilities. This reads on a human (e.g., mentally and/or using pen and paper): using a predetermined set of rules (i.e., to rescale output probabilities – mathematical concept) to obtain normalized output token and selecting output tokens by sampling the output tokens (i.e., mathematical concept) No additional limitations are present. With respect to claims 11 and 19, the claim(s) recite: 11. The computing system of claim 1, wherein, during a decoder plugin generation phase performed prior to computing the constrained output token vocabulary, the one or more processing devices are further configured to: receive a decoder plugin generation prompt; and based at least in part on the decoder plugin generation prompt, compute at least a portion of the decoder plugin at the machine learning model. This reads on a human (e.g., mentally and/or using pen and paper): receiving a prompt (e.g., from another human or after using a predetermined set of rules) and based on said prompt, define the predetermined set of rules. No additional limitations are present. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 10-12, and 18-20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Smith et al. (US 12632676 B1). As to independent claim 1, Smith et al. teaches: 1. A computing system (see ¶ col. 2, lines 28-31: “(13) The disclosed implementations describe systems and methods to guide an LLM in determining when to initiate a request to a user for clarification and when to cause execution of a determined directive based on an utterance…”) comprising: one or more processing devices (see ¶ col. 26, lines 41-46: “(138) Each of these devices (110/120/125) may include one or more controllers/processors (804/904), which may each include a central processing unit (CPU) for processing data…”) configured to: receive a prompt (see ¶ col. 3, lines 39-47: “(20) At some point, the device 110a may receive audio (utterance) corresponding to a spoken natural language input originating from the user 5, as in 150. The device 110a may generate audio data corresponding to the audio and may send the audio data to the remote system 120. Alternatively, the device 110b may receive a typed natural language input from the user 5. The device 110b may generate text data corresponding to the typed input and may send the text data to the remote system 120.”); at a machine learning model that has an output token vocabulary including a plurality of candidate output tokens, compute a plurality of output token probabilities over the output token vocabulary based at least in part on the prompt (see ¶ col. 13, lines 15-22: “74) Based on the list of possible devices, possible actions that may be performed by those devices, the generated text, and the user context and user history, a constrained token list may be generated, as in 410. For example, the constrained token list may indicate all domain tokens, such as SMART HOME, MUSIC, CHATBOT, or other skill that is associated with the user or the devices associated with the user.” and ¶ col. 14, lines 2-10: “(76) The generated LLM input may then be provided to an LLM tuned with the tokens, as discussed above with respect to FIG. 3, as in 416. (77) The LLM, for each pass through the LLM, generates and outputs probability scores for each of a plurality of tokens, including the tokens generated by the example process 300 (FIG. 3), indicating a probability that the token corresponds to the element for which the tokens are processed, as in 418…”); at a decoder plugin, compute a constrained output token vocabulary as a proper subset of the output token vocabulary (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in limitations above and further ¶ col. 10, lines 55-62: “(60) In some implementations, the LLM component 260 may be a decoder only LLM. Decoder only LLMs only allow information to flow forward in time, and each forward pass through the LLM results in a single token…” claim 1: “… receiving an utterance from a first device; converting the utterance into text; determining a first constrained token list indicating tokens that can be included as a first element of a structured output; processing, as a first pass through a decoder only large language model (“LLM”), the text with instructions to generate, for each of a plurality of tokens, a first probability score indicative of a first probability that the token is accurately selected by the decoder only LLM as the first element of the structured output; constraining the plurality of tokens to only those tokens indicated on the first constrained token list; determining, a first token from the first constrained token list having a highest first probability score; …”); select one or more output tokens based at least in part on the computed output token probabilities, wherein the one or more output tokens are selected from among the candidate output tokens included in the constrained output token vocabulary (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in limitations above and more specifically claim 1: “… receiving an utterance from a first device; converting the utterance into text; determining a first constrained token list indicating tokens that can be included as a first element of a structured output; processing, as a first pass through a decoder only large language model (“LLM”), the text with instructions to generate, for each of a plurality of tokens, a first probability score indicative of a first probability that the token is accurately selected by the decoder only LLM as the first element of the structured output; constraining the plurality of tokens to only those tokens indicated on the first constrained token list; determining, a first token from the first constrained token list having a highest first probability score; ”); and transmit an output including the one or more output tokens to an additional computing process (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in limitations above and further ¶ col. 18, lines 33-35: “(100) If it is determined at decision block 527 that a request for clarification is not to be generated, the structured output is sent by the example process 500 for execution, as in 528…” and claim 1: “…sending the structured output for execution.”). As to independent claim 12, Smith et al. teaches: 12. A method for use with a computing system (see ¶ col. 2, lines 28-31: “(13) The disclosed implementations describe systems and methods to guide an LLM in determining when to initiate a request to a user for clarification and when to cause execution of a determined directive based on an utterance…”), the method comprising: [the limitations as in claim 1, above]. As to independent claim 20, Smith et al. teaches: 20. A computing system (see ¶ col. 2, lines 28-31: “(13) The disclosed implementations describe systems and methods to guide an LLM in determining when to initiate a request to a user for clarification and when to cause execution of a determined directive based on an utterance…”) comprising: one or more processing devices (see ¶ col. 26, lines 41-46: “(138) Each of these devices (110/120/125) may include one or more controllers/processors (804/904), which may each include a central processing unit (CPU) for processing data…”) configured to: receive a prompt (see ¶ col. 3, lines 39-47: “(20) At some point, the device 110a may receive audio (utterance) corresponding to a spoken natural language input originating from the user 5, as in 150. The device 110a may generate audio data corresponding to the audio and may send the audio data to the remote system 120. Alternatively, the device 110b may receive a typed natural language input from the user 5. The device 110b may generate text data corresponding to the typed input and may send the text data to the remote system 120.”); at a machine learning model that has an output token vocabulary including a plurality of candidate output tokens, compute a plurality of output token probabilities over the output token vocabulary based at least in part on the prompt (see ¶ col. 13, lines 15-22: “74) Based on the list of possible devices, possible actions that may be performed by those devices, the generated text, and the user context and user history, a constrained token list may be generated, as in 410. For example, the constrained token list may indicate all domain tokens, such as SMART HOME, MUSIC, CHATBOT, or other skill that is associated with the user or the devices associated with the user.” and ¶ col. 14, lines 2-10: “(76) The generated LLM input may then be provided to an LLM tuned with the tokens, as discussed above with respect to FIG. 3, as in 416. (77) The LLM, for each pass through the LLM, generates and outputs probability scores for each of a plurality of tokens, including the tokens generated by the example process 300 (FIG. 3), indicating a probability that the token corresponds to the element for which the tokens are processed, as in 418…”); at a decoder plugin, modify respective output token probabilities associated with the plurality of candidate output tokens to thereby obtain constrained output token probabilities (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in limitations above and further ¶ col. 10, lines 55-62: “(60) In some implementations, the LLM component 260 may be a decoder only LLM. Decoder only LLMs only allow information to flow forward in time, and each forward pass through the LLM results in a single token…” claim 1: “… receiving an utterance from a first device; converting the utterance into text; determining a first constrained token list indicating tokens that can be included as a first element of a structured output; processing, as a first pass through a decoder only large language model (“LLM”), the text with instructions to generate, for each of a plurality of tokens, a first probability score indicative of a first probability that the token is accurately selected by the decoder only LLM as the first element of the structured output; constraining the plurality of tokens to only those tokens indicated on the first constrained token list; determining, a first token from the first constrained token list having a highest first probability score; …”); select one or more output tokens based at least in part on the constrained output token probabilities (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in limitations above and more specifically claim 1: “… receiving an utterance from a first device; converting the utterance into text; determining a first constrained token list indicating tokens that can be included as a first element of a structured output; processing, as a first pass through a decoder only large language model (“LLM”), the text with instructions to generate, for each of a plurality of tokens, a first probability score indicative of a first probability that the token is accurately selected by the decoder only LLM as the first element of the structured output; constraining the plurality of tokens to only those tokens indicated on the first constrained token list; determining, a first token from the first constrained token list having a highest first probability score;”); and transmit an output including the one or more output tokens to an additional computing process (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in limitations above and further ¶ col. 18, lines 33-35: “(100) If it is determined at decision block 527 that a request for clarification is not to be generated, the structured output is sent by the example process 500 for execution, as in 528…” and claim 1: “…sending the structured output for execution.”). Regarding claims 10 and 18, Smith et al. teaches the limitations as in claim 1 and 10, above. Smith et al. further teaches: 10. The computing system of claim 1, wherein, at the machine learning model (see ¶ col. 13, lines 15-22 and ¶ col. 14, lines 2-10 citations as in claim 1, above), the one or more processing devices (see ¶ col. 26, lines 41-46 citations as in claim 1, above) are further configured to: rescale the output token probabilities to obtain constrained output token probabilities normalized over the constrained output token vocabulary (see ¶ col. 14, lines 10-34: “…The output from the LLM may then be constrained to just the tokens indicated on the constrained token list and the probability scores for those tokens re-normalized across the tokens indicated on the constrained token list to generate a probability distribution based on just the constrained tokens, as in 419. For example, the LLM may output probability scores for thousands of tokens, including the tokens upon which it was tuned (FIG. 3). If the determined constrained token list includes four tokens, the LLM output will be constrained to those four tokens and the probability scores for those four tokens re-normalized with respect to those four tokens. A determination may then be made as to whether the highest probability score determined for a token of the constrained token list, once re-normalized, exceeds a probability threshold, as in 420. The probability threshold may be any value and may be different for different users, different actions, different domains, different devices, etc. (78) If it is determined that the highest re-normalized probability score, or at least one of the re-normalized probability scores, exceeds the probability threshold, in some implementations, a confidence score may be determined for the token with the highest probability score, the confidence score indicating a confidence that the token with the highest probability score is the correct token, as in 421.”); and select the one or more output tokens at least in part by sampling the output tokens from the constrained output token probabilities (see ¶ col. 14, lines 10-34 citations as in limitation above: “…(78) If it is determined that the highest re-normalized probability score, or at least one of the re-normalized probability scores, exceeds the probability threshold, in some implementations, a confidence score may be determined for the token with the highest probability score, the confidence score indicating a confidence that the token with the highest probability score is the correct token, as in 421.”). 18. The method of claim 12, further comprising, [the limitations as in claim 10, above]. Regarding claims 11 and 19, Smith et al. teaches the limitations as in claims 1 and 12, above. Smith et al. further teaches: 11. The computing system of claim 1, wherein, during a decoder plugin generation phase performed prior to computing the constrained output token vocabulary (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in claim 1, above and further ¶ col. 10, lines 55-62: “(60) In some implementations, the LLM component 260 may be a decoder only LLM. Decoder only LLMs only allow information to flow forward in time, and each forward pass through the LLM results in a single token…” claim 1: “… receiving an utterance from a first device; converting the utterance into text; determining a first constrained token list indicating tokens that can be included as a first element of a structured output; processing, as a first pass through a decoder only large language model (“LLM”), the text with instructions to generate, for each of a plurality of tokens, a first probability score indicative of a first probability that the token is accurately selected by the decoder only LLM as the first element of the structured output; constraining the plurality of tokens to only those tokens indicated on the first constrained token list; determining, a first token from the first constrained token list having a highest first probability score; …”), the one or more processing devices (see ¶ col. 26, lines 41-46 citation(s) as in claim 1, above.) are further configured to: receive a decoder plugin generation prompt (see ¶ col. 3, lines 39-47: “(20) At some point, the device 110a may receive audio (utterance) corresponding to a spoken natural language input originating from the user 5, as in 150…” and see ¶ col. 13, lines 15-22: “74) Based on the list of possible devices, possible actions that may be performed by those devices, the generated text, and the user context and user history, a constrained token list may be generated, as in 410. For example, the constrained token list may indicate all domain tokens, such as SMART HOME, MUSIC, CHATBOT, or other skill that is associated with the user or the devices associated with the user.” and ¶ col. 14, lines 2-10: “(76) The generated LLM input may then be provided to an LLM tuned with the tokens, as discussed above with respect to FIG. 3, as in 416. (77) The LLM, for each pass through the LLM, generates and outputs probability scores for each of a plurality of tokens, including the tokens generated by the example process 300 (FIG. 3), indicating a probability that the token corresponds to the element for which the tokens are processed, as in 418…”); and based at least in part on the decoder plugin generation prompt, compute at least a portion of the decoder plugin at the machine learning model (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in limitations above and further ¶ col. 10, lines 55-62: “(60) In some implementations, the LLM component 260 may be a decoder only LLM. Decoder only LLMs only allow information to flow forward in time, and each forward pass through the LLM results in a single token…” claim 1: “… receiving an utterance from a first device; converting the utterance into text; determining a first constrained token list indicating tokens that can be included as a first element of a structured output; processing, as a first pass through a decoder only large language model (“LLM”), the text with instructions to generate, for each of a plurality of tokens, a first probability score indicative of a first probability that the token is accurately selected by the decoder only LLM as the first element of the structured output; constraining the plurality of tokens to only those tokens indicated on the first constrained token list; determining, a first token from the first constrained token list having a highest first probability score; …”). 19. The method of claim 12, further comprising, [the limitations as in claim 11, above]. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 2 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (US 12632676 B1) as applied to claims 1 and 12 above, and further in view of Lott et al. (US 20260119905 A1). Regarding claims 2 and 13, Smith et al. teaches the limitations as in claims 1 and 12, above. Smith et al. further teaches: 2. The computing system of claim 1, wherein the one or more processing devices (see ¶ col. 26, lines 41-46 citation(s) as in claim 1, above.) are further configured to: based at least in part on the prompt, compute a tokenized prompt including a plurality of input tokens (see ¶ col. 13, lines 15-22: “74) Based on the list of possible devices, possible actions that may be performed by those devices, the generated text, and the user context and user history, a constrained token list may be generated, as in 410. For example, the constrained token list may indicate all domain tokens, such as SMART HOME, MUSIC, CHATBOT, or other skill that is associated with the user or the devices associated with the user.” and); at the machine learning model, compute the output token probabilities in each of a plurality of (see ¶ col. 13, lines 15-22 citation(s) as in limitation(s) above and further ¶ col. 14, lines 2-10: “(76) The generated LLM input may then be provided to an LLM tuned with the tokens, as discussed above with respect to FIG. 3, as in 416. (77) The LLM, for each pass through the LLM, generates and outputs probability scores for each of a plurality of tokens, including the tokens generated by the example process 300 (FIG. 3), indicating a probability that the token corresponds to the element for which the tokens are processed, as in 418…”); and at the decoder plugin, iteratively update the constrained output token vocabulary at each of a plurality of the (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as claim 1, above and further ¶ col. 10, lines 55-62: “(60) In some implementations, the LLM component 260 may be a decoder only LLM. Decoder only LLMs only allow information to flow forward in time, and each forward pass through the LLM results in a single token…” claim 1: “… receiving an utterance from a first device; converting the utterance into text; determining a first constrained token list indicating tokens that can be included as a first element of a structured output; processing, as a first pass through a decoder only large language model (“LLM”), the text with instructions to generate, for each of a plurality of tokens, a first probability score indicative of a first probability that the token is accurately selected by the decoder only LLM as the first element of the structured output; constraining the plurality of tokens to only those tokens indicated on the first constrained token list; determining, a first token from the first constrained token list having a highest first probability score; …”) including: the tokenized prompt (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as claim 1, above and further ¶ col. 2, lines 31-37: “…As discussed further below, at each pass through the LLM it may be determined whether the probability score and/or confidence score of a determined token exceeds respective thresholds. Based on those determinations, the system may decide whether to proceed with the determined tokens or to request clarification.”); and a prior output sequence including one or more prior output tokens computed at prior (see ¶ col. 11, lines 27-37: “… By generating tokens, tuning the LLM with those tokens, and then constraining the LLM output at each pass through the LLM to only a defined set of tokens, generation of the directive is performed with more efficiency and accuracy than traditional systems. For example, by constraining the LLM output at each pass to only tokens that correspond to a selected token from a previous pass, the accuracy is increased because the output will be constrained to only include tokens that are related to other aspects of the directive that have been selected for inclusion in a structured output...”). However, Smith et al. does not explicitly teach, but Lott et al. does teach: autoregressive generation iterations (see [0033]: “Speculative decoding, with an acceptance rate of α, may result in cost reductions relative to using a single autoregressive model to generate tokens iteratively. Inference cost savings, relative to iterative token generation, may be represented by the expression: [0034] Consider the example, for N=1000, C.sup.target=10, C.sup.draft=1, n=4, α=3, wherein N corresponds to a number of tokens, C.sup.AR corresponds to a computational cost using an acceptance rate of α, C.sup.target corresponds to a computational cost of generating a set of tokens using the target model, C.sup.draft corresponds to a computational cost of generating a set of tokens using the draft model, and n corresponds to a number of tokens generated speculatively generated tokens generated through a single pass through an autoregressive model. In such an example, speculative decoding may result in a 35% reduction in computational expense relative to autoregressive iterative token generation alone.”) Smith et al. and Lott et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. to incorporate the teachings of Lott et al. of autoregressive generation iterations which provides the benefit of improving the efficiency and throughout of large language models ([0025] of Lott et al.). 13. The method of claim 12, further comprising: [the limitations as in claim 2 and taught by Smith et al. in combination with Lott et al., above]. Smith et al. and Lott et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. to incorporate the teachings of Lott et al. of autoregressive generation iterations which provides the benefit of improving the efficiency and throughout of large language models ([0025] of Lott et al.). Claims 3 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (US 12632676 B1) in view of Lott et al. (US 20260119905 A1) as applied to claims 2 and 13 above, and further in view of Liu et al. (Liu, Michael Xieyang, et al. ""We need structured output": Towards user-centered constraints on large language model output." Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 2024. https://arxiv.org/pdf/2404.07362). Regarding claims 3 and 14, Smith et al. in combination with Lott et al. teaches the limitations as in claims 2 and 13, above. Smith et al. further teaches: 3. The computing system of claim 2, wherein: the one or more processing devices (see ¶ col. 26, lines 41-46 citation(s) as in claim 1, above.) are further configured to: However, Smith et al. in combination with Lott et al. does not explicitly teach, but Liu et al. does teach: the decoder plugin includes an oversight machine learning model (see ¶ 1 and 4 under RQ1: Real-world use cases that necessitate output constraints: “Table 1 presents a taxonomy of six primary categories of use cases that require output constraints, each with representative real-world examples and quotes submitted by our respondents… Giving an answer without extra conversational prose. When asking an LLM to perform data classification or labeling, such as “[classifying sentiments as] Positive, Negative, Neutral, etc.,” respondents typically expect the model to only output the classification result (e.g. “Positive.”) without a trailing “explanation” (e.g., “Positive, since it referred to the movie as a ‘timeless masterpiece’...”),as the addition of explanation could potentially confuse the downstream parsing logic. This indicates a potential misalignment between a common training objective — where LLMs are often tailored to be conversational and provide rich details [2, 17, 33] — and certain specialized downstream use cases where software developers need LLMs to be succinct. Such use cases necessitate output constraints that are independent of the prompt that would help adapt a general-purpose model to meet specific user requirements.”); and at the oversight machine learning model, compute a predicted classification of the output conditioned on the context (see ¶ 1 and 4 under 3. RQ1: Real-world use cases that necessitate output constraints citation(s) as limitation above and further Figure 2 and ¶ 6.1.4 Automatically inferring constraints based on prompts: “One interesting feature request for ConstraintMaker is the ability to automatically infer constraints from user-written prompts, simi lar to previous intelligent prediction or auto-completion systems and tools [3, 19, 37]. For instance, for a prompt shown in Fig. 2-1, ConstraintMaker could proactively suggest to the users if they’d like to constrain the model output to a JSON object with specific fields. …”); and select the constrained output token vocabulary based at least in part on the predicted classification (see ¶ 1 and 4 under 3. RQ1: Real-world use cases that necessitate output constraints and Figure 2 and ¶ 6.1.4 Automatically inferring constraints based on prompts citation(s) as limitation(s) above. More specifically: “Figure 2: ConstraintMaker’s user interfaces (1-4) & use cases (5-6). After writing the prompt (1), users can easily specify output constraints using a graphical user interface (2 & 3) provided by ConstraintMaker, and the resulting output (4) is guaranteed to follow the constraints…”). Smith et al., Lott et al., and Liu et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Lott et al. to incorporate the teachings of Liu et al. of the decoder plugin includes an oversight machine learning model; and at the oversight machine learning model, compute a predicted classification of the output conditioned on the context; and select the constrained output token vocabulary based at least in part on the predicted classification which provides the benefit of enabling users to prototype and test output constraints iteratively ([conclusion] of Liu et al.). 14. The method of claim 13, wherein: the decoder plugin includes an oversight machine learning model (see ¶ 1 and 4 under RQ1: Real-world use cases that necessitate output constraints citation(s) as in claim 3, above.); and the method further comprises: [the limitations as in claim 3 and taught by Smith et al. in combination with Lott et al. and Liu et al., above]. Smith et al., Lott et al., and Liu et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Lott et al. to incorporate the teachings of Liu et al. of the decoder plugin includes an oversight machine learning model; and at the oversight machine learning model, compute a predicted classification of the output conditioned on the context; and select the constrained output token vocabulary based at least in part on the predicted classification which provides the benefit of enabling users to prototype and test output constraints iteratively ([conclusion] of Liu et al.). Claims 4-5 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (US 12632676 B1) in view of Lott et al. (US 20260119905 A1) and Liu et al. (Liu, Michael Xieyang, et al. ""We need structured output": Towards user-centered constraints on large language model output." Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 2024. https://arxiv.org/pdf/2404.07362) as applied to claims 3 and 13 above, and further in view of Willard et al. (Willard, Brandon T., and Rémi Louf. "Efficient guided generation for large language models." arXiv preprint arXiv:2307.09702 (2023).). Regarding claim 4, Smith et al. in combination with Lott et al. and Liu et al. teaches the limitations as in claim 3, above. Smith et al. further teaches: 4. The computing system of claim 3, wherein, the one or more processing devices (see ¶ col. 26, lines 41-46 citation(s) as in claim 1, above.) are further configured to: Liu et al. further teaches: at the oversight machine learning model (see ¶ 1 and 4 under 3. RQ1: Real-world use cases that necessitate output constraints citation(s) as in claim 3, above), Smith et al., Lott et al., and Liu et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Lott et al. to incorporate the teachings of Liu et al. of the oversight machine learning which provides the benefit of enabling users to prototype and test output constraints iteratively ([conclusion] of Liu et al.). However, Smith et al. in combination with Lott et al. and Liu et al. do not explicitly teach, but Willard et al. does teach: compute a predicted completion of the context (see ¶ 5 of 3. Iterative FSM Processing and Indexing: “This formulation allows us to determine the exact states in Q in which the guiding regular expression’s FSM stops after sampling a single vocabulary token ˜st+1. These FSM states can then be tracked during the LLM token sampling process in Algorithm 2 and used to efficiently continue the state machine without reading from the beginning of the growing sample sequence each time.”); and compute the predicted classification based at least in part on the predicted completion (see ¶ 5 of 3. Iterative FSM Processing and Indexing citation as in limitation above and further Example 1: “Example 1. We illustrate the FSM sampling process in Figure 1 for the reg ular expression ([0-9]*)?\.?[0-9]*, which can be used to generate floating point numbers. For simplicity, let the vocabulary, V, consist of only the strings: "A", ".", "42", ".2", and "1". When the generation begins, the FSM is in state 0, so our algorithm masks the string "A", since it would not be accepted by the FSM. We can only sample ".", "42", ".2", and "1" in this case. If we sample ".2", we advance the FSM to state 3. In this case, only "42" and "1" are valid completions, so we mask the other values before sampling. If we sample "1" instead, we advance the FSM to state 1, in which case ".", ".42", ".2", and "1" are valid completions and the mask remains unchanged.”). Smith et al., Lott et al., Liu et al., and Willard et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Lott et al. and Liu et al. to incorporate the teachings of Willard et al. of compute a predicted completion of the context; and compute the predicted classification based at least in part on the predicted completion which provides the benefit of allowing one to enforce domain-specific knowledge and constraints, and enabling the construction of reliable interfaces by guaranteeing the structure of the generated text ([abstract] of Willard et al.). Regarding claims 5 and 15, Smith et al. in combination with Lott et al. teaches the limitations as in claims 2 and 13, above. Smith et al. further teaches: 5. The computing system of claim 2, wherein: the one or more processing devices (see ¶ col. 26, lines 41-46 citation(s) as in claim 1, above.) are further configured to: However, Smith et al. in combination with Lott et al. does not explicitly teach, but Liu et al. does teach: the decoder plugin includes an oversight machine learning model (see ¶ 1 and 4 under RQ1: Real-world use cases that necessitate output constraints citation as in claim 3, above.); and Smith et al., Lott et al., and Liu et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Lott et al. to incorporate the teachings of Liu et al. of the decoder plugin includes an oversight machine learning model which provides the benefit of enabling users to prototype and test output constraints iteratively ([conclusion] of Liu et al.). However, Smith et al. in combination with Lott et al. and Liu et al. do not explicitly teach, but Willard et al. does teach: at the oversight machine learning model, compute a predicted completion of the context (see ¶ 5 of 3. Iterative FSM Processing and Indexing: “This formulation allows us to determine the exact states in Q in which the guiding regular expression’s FSM stops after sampling a single vocabulary token ˜st+1. These FSM states can then be tracked during the LLM token sampling process in Algorithm 2 and used to efficiently continue the state machine without reading from the beginning of the growing sample sequence each time.”); compute a reward value associated with the predicted completion (see ¶ 5 of 3. Iterative FSM Processing and Indexing citation as in limitation above and further Example 1: “Example 1. We illustrate the FSM sampling process in Figure 1 for the regular expression ([0-9]*)?\.?[0-9]*, which can be used to generate floating point numbers. For simplicity, let the vocabulary, V, consist of only the strings: "A", ".", "42", ".2", and "1". When the generation begins, the FSM is in state 0, so our algorithm masks the string "A", since it would not be accepted by the FSM. We can only sample ".", "42", ".2", and "1" in this case. If we sample ".2", we advance the FSM to state 3. In this case, only "42" and "1" are valid completions, so we mask the other values before sampling. If we sample "1" instead, we advance the FSM to state 1, in which case ".", ".42", ".2", and "1" are valid completions and the mask remains unchanged.”); and select the constrained output token vocabulary based at least in part on the reward value (see ¶ 5 of 3. Iterative FSM Processing and Indexing and Example 1 citations as in limitations above. More specifically: “Example 1. …When the generation begins, the FSM is in state 0, so our algorithm masks the string "A", since it would not be accepted by the FSM. We can only sample ".", "42", ".2", and "1" in this case. If we sample ".2", we advance the FSM to state 3. In this case, only "42" and "1" are valid completions, so we mask the other values before sampling. If we sample "1" instead, we advance the FSM to state 1, in which case ".", ".42", ".2", and "1" are valid completions and the mask remains unchanged.”). Smith et al., Lott et al., Liu et al., and Willard et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Lott et al. and Liu et al. to incorporate the teachings of Willard et al. of at the oversight machine learning model, compute a predicted completion of the context; compute a reward value associated with the predicted completion; and select the constrained output token vocabulary based at least in part on the reward value which provides the benefit of allowing one to enforce domain-specific knowledge and constraints, and enabling the construction of reliable interfaces by guaranteeing the structure of the generated text ([abstract] of Willard et al.). 15. The method of claim 13, wherein: However, Smith et al. in combination with Lott et al. does not explicitly teach, but Liu et al. does teach: the decoder plugin includes an oversight machine learning model (see ¶ 1 and 4 under RQ1: Real-world use cases that necessitate output constraints citation as in claim 3, above.); and Smith et al., Lott et al., and Liu et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Lott et al. to incorporate the teachings of Liu et al. of the decoder plugin includes an oversight machine learning model which provides the benefit of enabling users to prototype and test output constraints iteratively ([conclusion] of Liu et al.). However, Smith et al. in combination with Lott et al. and Liu et al. do not explicitly teach, but Willard et al. does teach: the method further includes: [the last three limitations as in claim 5 and taught by Willard et al., above]. Smith et al., Lott et al., Liu et al., and Willard et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Lott et al. and Liu et al. to incorporate the teachings of Willard et al. of at the oversight machine learning model, compute a predicted completion of the context; compute a reward value associated with the predicted completion; and select the constrained output token vocabulary based at least in part on the reward value which provides the benefit of allowing one to enforce domain-specific knowledge and constraints, and enabling the construction of reliable interfaces by guaranteeing the structure of the generated text ([abstract] of Willard et al.). Claims 6-7, 9, and 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (US 12632676 B1) as applied to claims 1 and 12 above, and further in view of Willard et al. (Willard, Brandon T., and Rémi Louf. "Efficient guided generation for large language models." arXiv preprint arXiv:2307.09702 (2023).). Regarding claim 6, Smith et al. teaches the limitations as in claim 1, above. Smith et al. further teaches: 6. The computing system of claim 1, wherein, at the decoder plugin (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in claim 1, above), the one or more processing devices (see ¶ col. 26, lines 41-46 citations as in claim 1, above.) are further configured However, Smith et al. does not explicitly teach, but Willard et al. does teach: to modify one or more sampling parameters of the machine learning model (see Algorithm 1 and ¶ 1-3 of 2.1 Sampling sequences: “Let F ⊂P(V), where P is the powerset operator, be subsets of multi-token strings that end with a special token EOS ∈ V. The text generation task is to draw samples from F. Several procedures have been considered to generate elements of F. Greedy decoding consists in generating tokens recursively, choosing the to ken with highest probability at each step. Beam search also generates tokens recursively, using a heuristic to find the mode of the distribution. More re cently, SMC sampling has also been used to generate sequences [Lew et al., 2023]. The sampling procedure is described in generality by Algorithm 1. Often called multinomial sampling, the procedure recursively generates new tokens by sampling from the categorical distribution defined above until the EOS token is found.”). Smith et al. and Willard et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Lott et al. and Liu et al. to incorporate the teachings of Willard et al. of modify one or more sampling parameters of the machine learning model which provides the benefit of allowing one to enforce domain-specific knowledge and constraints, and enabling the construction of reliable interfaces by guaranteeing the structure of the generated text ([abstract] of Willard et al.). Regarding claim 7 and 16, Smith et al. teaches the limitations as in claim 1, above. Smith et al. further teaches: 7. The computing system of claim 1, wherein, at the decoder plugin (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in claim 1, above), the one or more processing devices (see ¶ col. 26, lines 41-46 citations as in claim 1, above.) are further configured to: However, Smith et al. does not explicitly teach, but Willard et al. does teach: execute a search algorithm over a predefined search domain (see Algorithm 1 and ¶ 1-3 of 2.1 Sampling sequences citations as in claim 6, above: “beam search”); and select the constrained output token vocabulary based at least in part on a search result returned by the search algorithm (see Algorithm 1 and ¶ 1-3 of 2.1 Sampling sequences citations as in claim 6, above: “beam search” and “…Often called multinomial sampling, the procedure recursively generates new tokens by sampling from the categorical distribution defined above until the EOS token is found). Smith et al. and Willard et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. to incorporate the teachings of Willard et al. of execute a search algorithm over a predefined search domain and select the constrained output token vocabulary based at least in part on a search result returned by the search algorithm which provides the benefit of allowing one to enforce domain-specific knowledge and constraints, and enabling the construction of reliable interfaces by guaranteeing the structure of the generated text ([abstract] of Willard et al.). 16. The method of claim 12, further comprising, at the decoder plugin (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in claim 1, above): [the limitations as in claim 7, above]. Smith et al. and Willard et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. to incorporate the teachings of Willard et al. of execute a search algorithm over a predefined search domain and select the constrained output token vocabulary based at least in part on a search result returned by the search algorithm which provides the benefit of allowing one to enforce domain-specific knowledge and constraints, and enabling the construction of reliable interfaces by guaranteeing the structure of the generated text ([abstract] of Willard et al.). Regarding claims 9 and 17, Smith et al. teaches the limitations as in claims 1 and 12, above. Smith et al. further teaches: 9. The computing system of claim 1, wherein, at the decoder plugin (see ¶ col. 3, lines 39-47, ¶ col. 13, lines 15-22, and ¶ col. 14, lines 2-10 citations as in claim 1, above), the one or more processing devices (see ¶ col. 26, lines 41-46 citations as in claim 1, above.) are configured to However, Smith et al. does not explicitly teach, but Willard et al. does teach: specify the constrained output token vocabulary with a regular expression or a context-free grammar (see Example 1 (We illustrate the FSM sampling process in Figure 1 for the regular expression…), Figure 1 (FSM masking for the regular expression), and ¶ 1-7 of 4. Extensions to Iterative Parsing: “In this section, we move our focus to general parser-guided generation and start with a simple walk-through for a Python-like grammar provided as a CFG. 12 Consider a vocabulary consisting of strings like "d" and "ef" that can be combined to produce Python-like syntax according to an implicit CFG, and assume that these strings are sequentially sampled and concatenated according to a process like Algorithm 1. Furthermore, consider a terminal symbol DEF in the CFG that corre sponds to the string "def" and is given by the trivial regular expression def. Also, consider a NAME symbol given by the regular expression [^\W\d]\w* (e.g. Python identifiers). We want to sequentially parse strings sampled from the aforementioned vocabulary in a way that adheres the Python syn tax. For example, the following could be one such sequence: ["d", "ef", " f", "oo(", "):", " ", "pass"]. All the elements of the sequence are by definition elements of the vocabulary. Concatenating the sequence produces "def foo(): pass", which is a valid sequence of tokens defining a function… In general, the next valid strings that can be sampled from the vocabu lary are ones that either 1. continue expanding/advancing the NAME currently starting with "f" (as the full sequence in our example does), and/or 2. anything that begins with "("–i.e. an LPAR symbol with regular ex pression (–and proceeds to specify a valid argument signature. In the first case, the "f" can be seen as a partially matched NAME symbol in Python, and–recalling that its regular expression is [^\W\d]\w*–we can say that it matches both sub-patterns (i.e. [^\W\d] and \w*) in the regular expression. Our use of FSMs formalize the notion of sub-patterns by way of an FSM’s states. In this case, the regex for NAME can be represented by an FSM, M, with three states: 0 (i.e. the initial state q0), 1 (i.e. [^\W\d]), and 2 (i.e. \w*), where 1,2 ∈ F.”). Smith et al. and Willard et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. to specify the constrained output token vocabulary with a regular expression which provides the benefit of allowing one to enforce domain-specific knowledge and constraints, and enabling the construction of reliable interfaces by guaranteeing the structure of the generated text ([abstract] of Willard et al.). 17. The method of claim 12, further comprising, [the limitations as in claim 9 as taught by Smith et al. in combination with Willard et al., above]. Smith et al. and Willard et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. to specify the constrained output token vocabulary with a regular expression which provides the benefit of allowing one to enforce domain-specific knowledge and constraints, and enabling the construction of reliable interfaces by guaranteeing the structure of the generated text ([abstract] of Willard et al.). Claim 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (US 12632676 B1) in view of Willard et al. (Willard, Brandon T., and Rémi Louf. "Efficient guided generation for large language models." arXiv preprint arXiv:2307.09702 (2023).) as applied to claim 7 above, and further in view of Zhang et al. (Zhang, Maosen, et al. "Language generation via combinatorial constraint satisfaction: A tree search enhanced Monte-Carlo approach." Findings of the Association for Computational Linguistics: EMNLP 2020. 2020. https://arxiv.org/pdf/2011.12334). Regarding claim 8, Smith et al. in combination with Willard et al. teaches the limitations as in claim 7, above. Smith et al. further teaches: 8. The computing system of claim 7, wherein, when executing the search algorithm at the decoder plugin, the one or more processing devices are configured to However, Smith et al. in combination with Willard et al. does not explicitly teach, but Zhang et al. does teach: perform a Monte Carlo tree search (MCTS) over a plurality of branches of the predefined search domain (see Figure 1 ((a) Natural language generation via con straint satisfaction (bottom), comparing to supervised approach (up). (b) Our proposed tree search enhanced MCMC (TSMH, pink line) traverses the probabilistic space of high-quality sentences more effectively than the baseline (blue line)) and ¶ 5-6 of 1. Introduction: “To better handle the combinatorial constraints, a tree search is embedded into the proposal process of the Markov chain Monte Carlo (MCMC) for constrained language generation, which suggests candidate proposals that satisfy more constraints…[…] …2. We propose a Tree Search enhanced Metropolis-Hastings approach (TSMH) for the proposed task, which mixes faster than standard MCMC in the presence of combinatorial constraints…” and ¶ 3. Tree Search Enhanced MCMC: “Markov chain Monte Carlo (MCMC) is a classical approach to sample sentences from probability dis tribution π(x) as defined in Equation 1…”). Smith et al., Willard et al., and Zhang et al. are considered to be analogous to the claimed invention because they are in the same field of endeavor in lexical analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith et al. in combination with Willard et al. to incorporate the teachings of Zhang et al. of perform a Monte Carlo tree search (MCTS) over a plurality of branches of the predefined search domain which provides the benefit of improving language generation tasks ([abstract] of Zhang et al.). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Keisha Y Castillo-Torres whose telephone number is (571)272-3975. The examiner can normally be reached Monday - Friday, 9:00 am - 4:00 pm (EST). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571)272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. Keisha Y. Castillo-Torres Examiner Art Unit 2659 /Keisha Y. Castillo-Torres/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Dec 02, 2024
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694878
METHODS AND SYSTEMS FOR COMBINED VOICE AND GESTURE CONTROL
3y 2m to grant Granted Jul 28, 2026
Patent 12694217
ENTITY RECOGNITION METHOD, MODEL TRAINING METHOD, ELECTRONIC DEVICE, AND MEDIUM
2y 3m to grant Granted Jul 28, 2026
Patent 12682159
INSTRUCTION FOLLOWING IN LARGE LANGUAGE MODELS TO REDUCE COMPUTATIONAL RESOURCE CONSUMPTION
2y 11m to grant Granted Jul 14, 2026
Patent 12682170
CONDENSING A DOCUMENT FOR ENHANCED ANALYSIS AND PROCESSING
2y 11m to grant Granted Jul 14, 2026
Patent 12664971
AUDIO-BASED MEDIA EDIT POINT SELECTION
2y 6m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
99%
With Interview (+31.3%)
2y 10m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 113 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month