DETAILED ACTION
This communication is in response to the Application filed on 03/01/2025. Claims 1-20 are pending and have been examined. Claims 1, 13 and 18 are independent. This Application was published as US20260260065A1.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 03/01/2025 was filed. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 7 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Lott et al. (US Pub 2024/0354346) in view of Bachmann et al. ("Judge decoding: Faster speculative sampling requires going beyond model alignment." International Conference on Learning Representations. Vol. 2025. 2025).
Regarding Claim 1,
Lott discloses a method for executing operations by a first machine-trained model (Lott, Fig.1, par [033], "…an example 100 of recursive speculative decoding in generative artificial intelligence models.."), comprising:
receiving a set of draft tokens that have been produced by a second machine- trained model (Lott, Fig.1, paras [033-035], "…a draft model and a target model can be used in conjunction (or otherwise together) to perform recursive speculative decoding of tokens...The draft model generally selects a plurality of high probability nodes (tokens) from a probability distribution for an output over a set of potential tokens..."; Fig.4, par [059], "…at block 410, with receiving a plurality of sets of tokens... the plurality of sets of tokens may be tokens generated ( e.g., by a draft model) based on an input prompt and a first generative artificial intelligence model...");
verifying correctness of the draft tokens in the set of draft tokens by comparing probability information generated by the first machine-trained model with probability information generated by the second machine-trained model (par [028], "…The target model takes the generated n tokens and processes the n tokens in parallel to generate probability distributions for each of the n tokens..."; par [029], "…The target model can then verify the tokens generated by the draft model by comparing distributions from the draft model and target model to determine whether a token is accepted or rejected…"; par [022], "…the target model can perform rejection sampling on a per-token basis to accept or reject individual tokens generated by the draft model such that the draft model and the target model have similar probability distributions...");
Lott does not disclose a context-matching determination process applied to the rejected draft tokens.
However, Bachmann, in the analogous field of endeavor, discloses determining whether any draft token in the set of draft tokens that is rejected by the verifying is accepted based on a specified matching criterion (Bachman, Abstract, "…Speculative decoding has been proposed as a technique to accelerate autoregressive
generation, leveraging a fast draft model to propose candidate tokens, which are then verified in parallel based on their likelihood under the target model…."; 4.1 Versatile and Accurate Verification with Token Embedding, Token embeddings signal error, "…we find that the model’s reaction to processing the incorrect token itself reveals surprisingly valuable information. Specifically, our experiments show that last hidden layer embeddings of erroneous tokens effectively ”flag” errors..."; Inference, "…we use fjudge as an additional evaluator for a given token..."; i.e., a trained evaluator module as an additional check is applied to the candidate token's embedding, and the candidate token is accepted when the module's confidence output exceeds a chosen threshold value; "…We take the logical OR between the two, zstand (i.e., standard speculative decoding verification) or zjudge (i.e., judge head of a trained evaluation module)..." ); and
providing draft tokens of the set of draft tokens that have been accepted by the verifying or the determining to the second machine-trained model for use by the second machine-trained model in generating another set of draft tokens (Bachmann, 3.1 Background, "…The accepted tokens are then added to the current context, and we repeat the steps until completion..."). Lott also discloses the autoregressive iteration process for sampling additional tokens (par [022], "…the draft model can generate speculatively additional tokens in sequence and probabilities used for sampling these additional tokens based on a current set of accepted tokens...").
Therefore, It would have been obvious to a person of ordinary skill in the art to apply Bachmann's known judge-dead technique of embedding-driven re-evaluation of a rejected draft token to Lott's batch draft and probability verification architecture, with the expectation of yielding a system that recovers contextually valid tokens a probability-only criterion would otherwise discard and carry forward accepted tokens.
Regarding Claim 2,
Lott in view of Bachmann discloses the method of claim 1, wherein the first machine-trained model and the second machine-trained model are respective language models having different respective total numbers of parameters (Lott, par [023], "…the draft model may be a smaller version of the target model (e.g., trained on millions of tokens, instead of hundreds of millions or even billions of tokens)...").
Regarding Claim 7,
Lott in view of Bachmann discloses the method of claim 1, further comprising invoking one or more matching components, from a set of available matching components, to perform the determining (Bachmann, Section 4, Inference, "…two accept/reject masks from the target model: zstand as in standard SD verification and zjudge from the judge head. We take the logical OR between the two, z = zstand V zjudge ...").
Claim 13 is a computing system claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 1.
Claims 3-5, 14 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Lott in view of Bachmann further in view of Li et al. ("Nearest neighbor speculative decoding for LLM generation and attribution." Advances in Neural Information Processing Systems 37 (2024): 80987-81015).
Regarding Claim 3,
Lott in view of Bachmann discloses the method of claim 1.
Bachmann further discloses rejecting a particular draft token in the set of draft tokens that is rejected by both the verifying and the determining (Bachmann, Section 4.1, Inference, "…we use fjudge as an additional evaluator for a given token..."; i.e., a trained evaluator module as an additional check is applied to the candidate token's embedding, and the candidate token is accepted when the module's confidence output exceeds a chosen threshold value; "…We take the logical OR between the two, zstand (i.e., standard speculative decoding verification) or zjudge (i.e., judge head of a trained evaluation module)..."; i.e., it is construed that a token is rejected when both standard speculative decoding verification and the judge head of the evaluation module reject the token) but does not explicitly disclose the limitation of "rejecting any draft tokens in the set of draft tokens that follows the particular draft token."
However, Li, in the analogous field of endeavor, discloses rejecting any draft tokens in the set of draft tokens that follows the particular draft token (Li, Abstract, "…NEST performs token-level retrieval at each inference step to compute a semi-parametric mixture distribution and identify promising span continuations in a corpus..."; Introduction, "…Relaxed speculative decoding. If a span of more than one token is selected, it undergoes evaluation based on the mixture probability...only a prefix deemed highly likely by the mixture probability is accepted."; 3.4 Relaxed Speculative Decoding, "…the relaxation factor...If token...is rejected, we will remove all the tokens from that point to the end of span…"; i.e., Li discloses a relaxed version of speculative decoding that upper-bounds the acceptance probability for each token in a retrieved span using a tunable relaxation factor).
Therefore, It would have been obvious to a person of ordinary skill in the art to apply Li's known span-threshold (i.e., relaxation factor) technique for bounding consecutively accepted retrieved tokens to the machine learning-based probability-based verification and matching-based determination of draft tokens taught by Lott in view of Bachmann, with the expectation of yielding a system that limits the number of matching tokens in a row are allowed.
Regarding Claim 4,
Lott in view of Bachmann discloses the method of claim 1, further comprising rejecting a particular draft token in the set of draft tokens when: the particular draft token is rejected by the verifying; and
Li discloses the particular draft token is preceding by a prescribed number of draft tokens that have been accepted by the determining (Li, 3.4 Relaxed Speculative Decoding, "…the relaxation factor...If token...is rejected, we will remove all the tokens from that point to the end of span…"; i.e., Li discloses a relaxed version of speculative decoding that upper-bounds the acceptance probability for each token in a retrieved span using a tunable relaxation factor), the prescribed number being specified by a token span threshold parameter (Li, Section 3.4, a tunable relaxation factor).
Rationale for combination is similar to that provided for Claim 3.
Regarding Claim 5,
Lott in view of Bachmann further in view of Li discloses the method of claim 4, wherein the token span threshold parameter is determined by a configuration setting and remains fixed through the operations (Li, Appendix A, "…For relaxed speculative decoding, we set (the relaxation factor) for Pile of Law tasks for all model...For Wikipedia-based tasks, we set (a different relaxation factor) for the 7B model...for the 13B model, and for the 70B model...").
Claim 14 is a computing system claim with limitations similar to the limitations of Claim 4 and is rejected under similar rationale.
Claim 18 is a computer-readable storage medium claim with limitations similar to the limitations of Claims 1, 3, and 4 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 3.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Lott in view of Bachmann further in view of Li further in view of Goel et al. (US Pub 2025/0245530).
Regarding Claim 6,
Lott in view of Bachmann further in view of Li discloses the method of claim 4 but does not explicitly disclose dynamically varying token span threshold parameter.
However, Goel, in the analogous field of speculative decoding of LLM, discloses wherein the token span threshold parameter dynamically varies over the operations (Goel, par [018] "…Subsequent rounds of inferencing may involve adjustments to the number of tokens included in a draft set of tokens (also referred to as a token length). These adjustments may be performed to maximize, or at least increase, the rate at which tokens are generated as a response to the input query…"; par [037], "… Adjustments to the length of the draft set of tokens to be generated by the generative artificial intelligence model may be performed based on the (measured) operational parameter(s) of the device… the acceptance rate of the previously generated draft set of tokens, historical acceptance rates of other sets of tokens generated by the generative artificial intelligence model for the input query and historical input queries, the scheduling function…").
Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have applied Goel’s known technique of dynamically redetermining a draft token count parameter each round via a scheduling function to a span threshold parameter fixed by configuration of Lott in view of Bachmann further in view of Li, with the reasonable expectation of yielding a span threshold that varies dynamically over the course of operation.
Allowable Subject Matter
Claims 8-12, 15-17, and 19-20 are objected to as being dependent upon a rejected base claims 1, 13 and 18 but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claims 8-12, 15-17 and 19-20 disclose identified semantic attributes (e.g., topic, theme, category, style, intent, syntactic structure, or organization of parts) as a set of matching components that are extracted from context information and a draft token and then compared for a match. Prior art literature search within speculative decoding failed to locate a token-level matching and recovery mechanism for probability-rejected draft tokens within a two-model speculative decoding model of LLM using one or more semantic attributes.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Agrawal et al. ("AdaEDL: Early draft stopping for speculative decoding of large language models via an entropy-based lower bound on token acceptance probability." arXiv preprint arXiv:2410.18351 (2024)) discloses Adaptive Entropy-based Draft Length (AdaEDL), a simple, training and parameter-free criteria which allows for early stopping of the token drafting process by approximating a lower bound on the expected acceptance probability of the drafted token based on the currently observed entropy of the drafted logits (Agrawal, Abstract).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JANGWOEN LEE whose telephone number is (703)756-5597. The examiner can normally be reached Monday-Friday 8:00 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, BHAVESH MEHTA can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JANGWOEN LEE/ Examiner, Art Unit 2656
/BHAVESH M MEHTA/ Supervisory Patent Examiner, Art Unit 2656