DETAILED ACTION
This is responsive to the application filed 24 January 2025.
Claims 1-20 are pending and considered below.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 1-20 are objected to because of the following informalities: in line 19 of claim 1, it is believed the limitation “the one of the draft tokens are accepted” should be ‘the one of the draft tokens is accepted’. Independent claim 16 suffers from a similar deficiency and is likewise objected to. The dependent claims are objected to for depending upon an objected claim without providing a remedy.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
In lines 8-14, claim 1 recites the limitations “verifying at least one of the draft tokens, wherein the verification result indicates that, when a ratio of a first probability of one of the draft tokens to a second probability of the one of the draft tokens is larger than a threshold which is lower than 1, or the one of the draft tokens is identical to a corresponding one of the answer tokens, the one of the draft tokens passes the verification” (emphasis added). It is unclear if and how “the verification result” and/or “the verification” refers back to the claimed “verifying”.
Independent claim 16 recites similar limitations and is likewise rejected.
The dependent claims are rejected for depending upon a rejected claim without providing a remedy.
The limitations will be interpreted as:
verifying at least one of the draft tokens, wherein a verification result indicates that, when a ratio of a first probability of one of the draft tokens to a second probability of the one of the draft tokens is larger than a threshold which is lower than 1, or the one of the draft tokens is identical to a corresponding one of the answer tokens, the one of the draft tokens passes the verification
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-7, 9, 12 and 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Lott et al. (US 2024/0320433) in view of Shazeer et al. (US 2020/0082226).
Claim 1:
Lott discloses a method for larger language model (LLM) inference, comprising:
receiving an input prompt; generating sequentially a plurality of draft tokens based on the input prompt by a draft model (“generating a response to an input query using a generative artificial intelligence model. The method generally includes generating, based on an input query and a first generative model, a plurality of sets of tokens, each set of tokens in the plurality of sets of tokens corresponding to a candidate response to the input query”, [0005], see also “generating responses to a query input into a large language model on a group basis using speculative decoding techniques. Generally, the draft model can generate one or more sets of tokens as candidate responses to the query. The target model, in turn, can perform sampling rejection on a per-set basis. Tokens within a group can be selected based on conditional probabilities of tokens within the group”, [0027] and “the draft model may speculatively generate n tokens autoregressively”, [0030], note autoregressive models generate tokens sequentially, see “a single token may be generated each time an autoregressive model is executed, which means that N inferences may be performed to generate a sequence of N tokens”, [0029]);
generating a plurality of answer tokens in parallel based on the plurality of draft tokens by a target model (“The plurality of sets of tokens are output to a second generative model for verification. An indication of a selected set of tokens from the plurality of sets of tokens is received from the second generative model based on the input query and the plurality of sets of tokens. The selected set of tokens is output as a response to the input query”, [0005], see also “The target model, in turn, can perform sampling rejection on a per-set basis. Tokens within a group can be selected based on conditional probabilities of tokens within the group”, [0027], see also “The target model takes the generated n tokens and processes the n tokens in parallel to generate probability distributions for each of the n tokens”, [0031]); and
verifying at least one of the draft tokens (“The plurality of sets of tokens are output to a second generative model for verification”, [0005]), wherein a first probability of one of the draft tokens is generated by the draft model, and a second probability of the one of the draft tokens is generated by the target model (“A probability distribution associated with each respective set of tokens in the plurality of sets of tokens is compared to a corresponding probability distribution generated by a second generative model for the respective set of tokens”, [0006], see also “the draft model can generate speculatively additional tokens and probabilities used for sampling these additional tokens based on a current set of accepted tokens. The target model can generate tokens based on the tokens generated by the draft model. To generate a result, the target model can perform rejection sampling on a per-token basis to accept or reject tokens generated by the draft model such that the draft model and the target model have similar probability distributions”, [0025]),
wherein, when the one of the draft tokens passes the verification, the one of the draft tokens are accepted, and a token, obtained from the one of the draft tokens or the corresponding one of the answer tokens, is appended into a response for the input prompt (“The plurality of sets of tokens are output to a second generative model for verification. An indication of a selected set of tokens from the plurality of sets of tokens is received from the second generative model based on the input query and the plurality of sets of tokens. The selected set of tokens is output as a response to the input query.”, [0005]).
Lott does not explicitly disclose wherein a verification result indicates that, when a ratio of the first probability of one of the draft tokens to the second probability of the one of the draft tokens is larger than a threshold which is lower than 1, or the one of the draft tokens is identical to a corresponding one of the answer tokens, the one of the draft tokens passes the verification.
In a similar system generating a plurality draft tokens by a draft model (auxiliary model), generating a plurality of answer tokens based on the plurality of draft tokens by a target model (base model p1) and verifying at least one of draft tokens (“each auxiliary model p2, . . . , pk receives the input x and predicts a single output token ŷi for i=2, . . . , k”, [0045], see also “the base model p1 is used as a scoring model to determine which of the tokens ŷi for i=2, . . . , k, should be accepted”, [0046]), Shazeer discloses wherein a verification result indicates that, when a ratio of a first probability of one of the draft tokens to a second probability of the one of the draft tokens is larger than a threshold which is lower than 1, or the one of the draft tokens (“in” and “the”, which were predicted by the models p1 and p2 during the prediction) is identical to a corresponding one of the answer tokens, the one of the draft tokens passes the verification (“finds a largest n such that (i) a prediction from model p1 of a next token for an input of the current input concatenated with the first through the (n−1)st tokens independently predicted by models p1 through pn-1 matches (ii) the independent prediction of the n-th token by model pn”, [0050], see also “In the verification substep 420, the base model p1 scores each of the independent predictions, conditioning on the previous independent predictions where applicable. In the example of FIG. 4, the highest probability prediction, or scoring prediction, for the third position is “car”. The scoring prediction “car” is predicted by p1 using an input 414, i.e., the words “I saw a dog ride in the”. The input 414 is the words of prefix 412 concatenated with the independently predicted “in” and “the”, which were predicted by the models p1 and p2 during the prediction substep 410. The scoring prediction “car” is different from the prediction by p3 of “bus” for the third position”, [0056], see also “In the acceptance substep 430, the system extends the prefix, ŷ, to include the predictions for the first and second positions, i.e., “in” and “the”, but not the prediction for the third position, i.e., “bus””, [0057]).
It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to combine the references to yield the predictable result of passing Lott’s verification when the one of the draft tokens is identical to a corresponding one of the answer tokens as disclosed by Shazeer because both references address the same problem of accelerating autoregressive/Transformer token generation by generating candidate tokens and verifying them in parallel. Substituting Lott’s token-verification using Shazeer’s exact match criterion would have been a predictable implementation of a known token-verification technique.
Claim 2:
Lott in view of Shazeer discloses the method according to claim 1, wherein the token is the one of the draft tokens (Lott, [0005], note the selected set of tokens is selected from the plurality of sets of tokens).
Claim 3:
Lott in view of Shazeer discloses the method according to claim 1, wherein when the one of the draft tokens passes the verification and all draft tokens before the one of the draft tokens pass the verification, the one of the draft tokens is accepted (Lott, [0005], note a verified/selected token is passed as a response).
Claim 4:
Lott in view of Shazeer discloses the method according to claim 1, wherein the plurality of draft tokens is γ draft tokens, wherein an ith draft token is generated by the draft model based on an (i−1)th draft token input to the draft model, i is greater than 1 and less than or equal to γ (Lott, [0030]).
Claim 5:
Lott in view of Shazeer discloses the method according to claim 1, wherein generating the plurality of answer tokens in parallel based on the plurality of draft tokens comprises generating a plurality of tokens in parallel based on the plurality of draft tokens, wherein the plurality of tokens comprise the plurality of answer tokens and an additional token, wherein the additional token is generated based on the last draft token in the plurality of draft tokens (Lott, [0032] and [0043], see also Shazeer, [0046] and Fig. 2B, last p1 output).
Claim 6:
Lott in view of Shazeer discloses the method according to claim 5, wherein the plurality of tokens comprise γ answer tokens and the additional token, wherein a first answer token is generated based on the input prompt, an ith answer token is generated based on the first to i−1th draft tokens, and the additional token is generated based on the first to Yth draft tokens (Shazeer, [0046] and Fig. 2B, last p1 output).
Claim 7:
Lott in view of Shazeer discloses the method according to claim 5, further comprising: if all of the draft tokens are accepted according to the verification result, the additional token is appended into the response (Shazeer, [0046], [0057] and Fig. 2B, last p1 output).
Claim 9:
Lott in view of Shazeer discloses the method according to claim 1, further comprising: if not all of the draft tokens are accepted according to the verification result, an accepted draft token subset is appended into the response (Shazeer, [0057]), a token probability distribution of the next position of the accepted draft token subset is adjusted, and a next token is obtained based on the adjusted token probability distribution and appended into the response (Shazeer, [0039], [0058], note that iterations adjust parameters, see Lott, [0039], [0055]).
Claim 12:
Lott in view of Shazeer discloses the method according to claim 1, wherein the ratio of the first probability of the one of the draft tokens to the second probability of the one of the draft tokens is referred as … (further narrowing a limitation claimed in the alternative and not mapped by the references).
Claim 14:
Lott in view of Shazeer discloses the method according to claim 5, further comprising: generating sequentially a second group of draft tokens based on the input prompt by the draft model, and generating a second group of answer tokens in parallel based on the second group of draft tokens and the additional token generated by the target model if all of the draft tokens are accepted (Lott, [0066], [0070], [0072]).
Claim 15:
Lott in view of Shazeer discloses the method according to claim 9, further comprising: generating sequentially a second group of draft tokens based on the input prompt by the draft model, and outputting a second group of answer tokens in parallel based on the second group of draft tokens and the obtained next token generated by the target model if not all of the draft tokens are accepted (Lott, [0066], [0070], [0072]).
Claims 8, 10-11, 16-17 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Lott et al. (US 2024/0320433) in view of Shazeer et al. (US 2020/0082226) and Ramanujan et al. (US 2024/0419493).
Claim 16:
Lott discloses a system for LLM inference, comprising:
a processor ([0122]), configured to
receive an input prompt, generate sequentially a plurality of draft tokens based on the input prompt by running a draft model (“generating a response to an input query using a generative artificial intelligence model. The method generally includes generating, based on an input query and a first generative model, a plurality of sets of tokens, each set of tokens in the plurality of sets of tokens corresponding to a candidate response to the input query”, [0005], see also “generating responses to a query input into a large language model on a group basis using speculative decoding techniques. Generally, the draft model can generate one or more sets of tokens as candidate responses to the query. The target model, in turn, can perform sampling rejection on a per-set basis. Tokens within a group can be selected based on conditional probabilities of tokens within the group”, [0027] and “the draft model may speculatively generate n tokens autoregressively”, [0030], note autoregressive models generate tokens sequentially, see “a single token may be generated each time an autoregressive model is executed, which means that N inferences may be performed to generate a sequence of N tokens”, [0029]),
generate a plurality of answer tokens in parallel based on the plurality of draft tokens by running a target model (“The plurality of sets of tokens are output to a second generative model for verification. An indication of a selected set of tokens from the plurality of sets of tokens is received from the second generative model based on the input query and the plurality of sets of tokens. The selected set of tokens is output as a response to the input query”, [0005], see also “The target model, in turn, can perform sampling rejection on a per-set basis. Tokens within a group can be selected based on conditional probabilities of tokens within the group”, [0027], see also “The target model takes the generated n tokens and processes the n tokens in parallel to generate probability distributions for each of the n tokens”, [0031]),
verifying at least one of the draft tokens (“The plurality of sets of tokens are output to a second generative model for verification”, [0005]), wherein a first probability of one of the draft tokens is generated by the draft model, and a second probability of the one of the draft tokens is generated by the target model (“A probability distribution associated with each respective set of tokens in the plurality of sets of tokens is compared to a corresponding probability distribution generated by a second generative model for the respective set of tokens”, [0006], see also “the draft model can generate speculatively additional tokens and probabilities used for sampling these additional tokens based on a current set of accepted tokens. The target model can generate tokens based on the tokens generated by the draft model. To generate a result, the target model can perform rejection sampling on a per-token basis to accept or reject tokens generated by the draft model such that the draft model and the target model have similar probability distributions”, [0025]),
wherein, when the one of the draft tokens passes the verification, the one of the draft tokens are accepted, and a token, obtained from the one of the draft tokens or the corresponding one of the answer tokens, is appended into a response for the input prompt (“The plurality of sets of tokens are output to a second generative model for verification. An indication of a selected set of tokens from the plurality of sets of tokens is received from the second generative model based on the input query and the plurality of sets of tokens. The selected set of tokens is output as a response to the input query.”, [0005]).
Lott does not explicitly disclose wherein a verification result indicates that, when a ratio of the first probability of one of the draft tokens to the second probability of the one of the draft tokens is larger than a threshold which is lower than 1, or the one of the draft tokens is identical to a corresponding one of the answer tokens, the one of the draft tokens passes the verification.
In a similar system generating a plurality draft tokens by a draft model (auxiliary model), generating a plurality of answer tokens based on the plurality of draft tokens by a target model (base model p1) and verifying at least one of draft tokens (“each auxiliary model p2, . . . , pk receives the input x and predicts a single output token ŷi for i=2, . . . , k”, [0045], see also “the base model p1 is used as a scoring model to determine which of the tokens ŷi for i=2, . . . , k, should be accepted”, [0046]), Shazeer discloses wherein a verification result indicates that, when a ratio of a first probability of one of the draft tokens to a second probability of the one of the draft tokens is larger than a threshold which is lower than 1, or the one of the draft tokens (“in” and “the”, which were predicted by the models p1 and p2 during the prediction) is identical to a corresponding one of the answer tokens, the one of the draft tokens passes the verification (“finds a largest n such that (i) a prediction from model p1 of a next token for an input of the current input concatenated with the first through the (n−1)st tokens independently predicted by models p1 through pn-1 matches (ii) the independent prediction of the n-th token by model pn”, [0050], see also “In the verification substep 420, the base model p1 scores each of the independent predictions, conditioning on the previous independent predictions where applicable. In the example of FIG. 4, the highest probability prediction, or scoring prediction, for the third position is “car”. The scoring prediction “car” is predicted by p1 using an input 414, i.e., the words “I saw a dog ride in the”. The input 414 is the words of prefix 412 concatenated with the independently predicted “in” and “the”, which were predicted by the models p1 and p2 during the prediction substep 410. The scoring prediction “car” is different from the prediction by p3 of “bus” for the third position”, [0056], see also “In the acceptance substep 430, the system extends the prefix, ŷ, to include the predictions for the first and second positions, i.e., “in” and “the”, but not the prediction for the third position, i.e., “bus””, [0057]).
It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to combine the references to yield the predictable result of passing Lott’s verification when the one of the draft tokens is identical to a corresponding one of the answer tokens as disclosed by Shazeer because both references address the same problem of accelerating autoregressive/Transformer token generation by generating candidate tokens and verifying them in parallel. Substituting Lott’s token-verification using Shazeer’s exact match criterion would have been a predictable implementation of a known token-verification technique.
Lott in view of Shazeer discloses a memory, configured to store at least one of the plurality of answer tokens and at least one of the plurality of draft tokens (Lott, [0135], note the tokens have to be stored, at least temporarily, for processing).
Lott in view of Shazeer does not explicitly disclose storing KV caches of at least one of the plurality of answer tokens and at least one of the plurality of draft tokens.
In an analogous art similarly performing AI inference, Ramanujan discloses storing KV caches of tokens (“As requests are processed by an AI model using a GPU, a KV cache is used to store key vectors and value vectors calculated for particular tokens in a request in distinct KV blocks”, [0009]).
It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to combine the references to yield the predictable result of storing KV caches for Lott’s at least one of the plurality of answer tokens and at least one of the plurality of draft tokens because KV caching is a standard performance optimization in AI inference which increases GPU efficiency by avoiding repetitive calculations (see Ramanujan, [0009]).
Claim 17:
Lott in view of Shazeer and Ramanujan discloses the system according to claim 16, wherein the ratio of the first probability of the one of the draft tokens to the second probability of the one of the draft tokens is referred as … (further narrowing a limitation claimed in the alternative and not mapped by the references).
Claim 19:
Lott in view of Shazeer and Ramanujan the system according to claim 16, wherein the processor comprises a CPU and a NPU ([0122]-[0123]), the system configured to receive an input prompt, generate sequentially a plurality of draft tokens based on the input prompt by running a draft model (“generating a response to an input query using a generative artificial intelligence model. The method generally includes generating, based on an input query and a first generative model, a plurality of sets of tokens, each set of tokens in the plurality of sets of tokens corresponding to a candidate response to the input query”, [0005], see also “generating responses to a query input into a large language model on a group basis using speculative decoding techniques. Generally, the draft model can generate one or more sets of tokens as candidate responses to the query. The target model, in turn, can perform sampling rejection on a per-set basis. Tokens within a group can be selected based on conditional probabilities of tokens within the group”, [0027] and “the draft model may speculatively generate n tokens autoregressively”, [0030], note autoregressive models generate tokens sequentially, see “a single token may be generated each time an autoregressive model is executed, which means that N inferences may be performed to generate a sequence of N tokens”, [0029]), and generate a plurality of answer tokens in parallel based on the plurality of draft tokens by running a target model (“The plurality of sets of tokens are output to a second generative model for verification. An indication of a selected set of tokens from the plurality of sets of tokens is received from the second generative model based on the input query and the plurality of sets of tokens. The selected set of tokens is output as a response to the input query”, [0005], see also “The target model, in turn, can perform sampling rejection on a per-set basis. Tokens within a group can be selected based on conditional probabilities of tokens within the group”, [0027], see also “The target model takes the generated n tokens and processes the n tokens in parallel to generate probability distributions for each of the n tokens”, [0031]); the system configured to verify at least one of the draft tokens (“The plurality of sets of tokens are output to a second generative model for verification”, [0005]).
Lott in view of Shazeer and Ramanujan does not explicitly disclose the NPU performing the receiving of the input prompt and the generation of draft and answer tokens, the CPU performing the verification and notifying the NPU which draft tokens are accepted.
Lott discloses that an “NPU, such as the NPU 1108, is generally a specialized circuit configured for implementing control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), and the like” ([0124]) and “NPUs designed to accelerate inference are generally configured to operate on complete models. Such NPUs may thus be configured to input a new piece of data and rapidly process this new piece through an already trained model to generate a model output (e.g., an inference)” ([0128]).
It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to perform the receiving of the input prompt and the generation of draft and answer tokens using the NPU and performing the verification and notifying the NPU which draft tokens are accepted using the CPU in order to use the different processing units based on their functionality strengths and purpose.
Claim 20:
Lott in view of Shazeer and Ramanujan system according to claim 16, wherein the plurality of answer tokens comprise γ answer tokens, wherein an ith answer token is generated based on the first to i−1th draft tokens, and the one additional token is generated based on the first to Yth draft tokens (Shazeer, [0046] and Fig. 2B, last p1 output).
Claim 8:
Lott in view of Shazeer discloses the method according to claim 5, further comprising: if all of the draft tokens are accepted, the draft model receives the last draft token in the plurality of draft tokens and stores the last draft token (Lott, “If all groups of tokens are accepted by the target model, a final token may be sampled from a final tree leaf distribution (e.g., a probability distribution generated over the leaf nodes in the tree data structure 110 corresponding to a last accepted or verified token for any partial path through the tree data structure 110), which may allow for the target model to generate an additional token relative to the token(s) speculatively generated by the draft model. The selected set of tokens 112 may then be communicated back to the draft model”, [0043], note tokens have to be stored, at least temporarily, for processing).
Lott in view of Shazeer does not explicitly disclose generating KV cache of the last draft token.
In an analogous art similarly performing AI inference, Ramanujan discloses storing KV caches of tokens (“As requests are processed by an AI model using a GPU, a KV cache is used to store key vectors and value vectors calculated for particular tokens in a request in distinct KV blocks”, [0009]).
It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to combine the references to yield the predictable result of storing KV caches for Lott’s last draft token because KV caching is a standard performance optimization in AI inference which increases GPU efficiency by avoiding repetitive calculations (see Ramanujan, [0009]).
Claim 10:
Lott in view of Shazeer discloses the method according to claim 1, further comprising: if not all of the draft tokens are accepted according to the verification result, rejecting unaccepted draft token (Lott, “The target model can then verify the tokens generated by the draft model by comparing distributions from the draft model and target model to determine whether a token is accepted or rejected … the token may be rejected”, [0032], see also Shazeer, [0057] and Fig. 4, note that neither the draft token “bus” nor its corresponding answer token “car” are accepted, i.e. they are both rejected/removed).
Lott in view of Shazeer does not explicitly disclose KV cache of an unaccepted draft token is removed/rejected.
In an analogous art similarly performing AI inference, Ramanujan discloses KV caches of tokens are well known (“As requests are processed by an AI model using a GPU, a KV cache is used to store key vectors and value vectors calculated for particular tokens in a request in distinct KV blocks”, [0009]).
It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to combine the references to yield the predictable result of removing/rejecting Lott’s unaccepted draft token because KV caching is a standard performance optimization in AI inference which increases GPU efficiency by avoiding repetitive calculations (see Ramanujan, [0009]).
Claim 11:
Lott in view of Shazeer discloses the method according to claim 1, further comprising: if not all of the draft tokens are accepted according to the verification result, the answer token corresponding to an unaccepted draft token is removed (Lott, “The target model can then verify the tokens generated by the draft model by comparing distributions from the draft model and target model to determine whether a token is accepted or rejected … the token may be rejected”, [0032], see also Shazeer, [0057] and Fig. 4, note that neither the draft token “bus” nor its corresponding answer token “car” are accepted, i.e. they are both rejected/removed).
Lott in view of Shazeer does not explicitly disclose KV cache of the answer token corresponding to an unaccepted draft token is removed.
In an analogous art similarly performing AI inference, Ramanujan discloses KV caches of tokens are well known (“As requests are processed by an AI model using a GPU, a KV cache is used to store key vectors and value vectors calculated for particular tokens in a request in distinct KV blocks”, [0009]).
It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to combine the references to yield the predictable result of removing/rejecting Lott’s answer token corresponding to an unaccepted draft token because KV caching is a standard performance optimization in AI inference which increases GPU efficiency by avoiding repetitive calculations (see Ramanujan, [0009]).
Allowable Subject Matter
Claims 13 and 18 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter: the prior art of record, individually or in combination, does not disclose the verification pass/fail as claimed in claims 13 and 18.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chen et al. ("Accelerating large language model decoding with speculative sampling." arXiv preprint arXiv:2302.01318 (2023)) discloses generating a short draft of length 𝐾. This can be attained with either a parallel model or by calling a faster, auto-regressive model 𝐾 times. This model is referred as a draft model. The draft model is then scored using a larger, more powerful model. This model is referred as the target model. A subset of the 𝐾 draft tokens is then accepted from left to right using a modified rejection sampling scheme.
Santhanam et al. (US 2025/0021761) discloses a system for generating a response to a query input into a generative artificial intelligence model. An example method generally includes generating, based on an input query and a first generative artificial intelligence model, a sequence of tokens corresponding to a candidate response to the input query. The sequence of tokens and the input query are output to a second generative artificial intelligence model for verification. One or more first guidance signals for the generated sequence of tokens are received from the second generative artificial intelligence model. The candidate response to the input query is revised based on the generated sequence of tokens and the one or more first guidance signals, and the revised candidate response is output as a response to the received input query.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAMUEL G NEWAY whose telephone number is (571)270-1058. The examiner can normally be reached Monday-Friday 9:00am-5:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SAMUEL G NEWAY/Primary Examiner, Art Unit 2657