Prosecution Insights
Last updated: August 17, 2026
Application No. 18/436,474

SYSTEMS AND METHODS FOR RESPONDING TO LATENCY IN OUTPUT FROM A GENERATIVE MODEL

Non-Final OA §103§112
Filed
Feb 08, 2024
Examiner
JABLON, ASHER H.
Art Unit
Tech Center
Assignee
Shopify Inc.
OA Round
1 (Non-Final)
43%
Grant Probability
Moderate
1-2
OA Rounds
1y 10m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 43% of resolved cases
43%
Career Allowance Rate
40 granted / 94 resolved
-17.4% vs TC avg
Strong +44% interview lift
Without
With
+44.5%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
24 currently pending
Career history
121
Total Applications
across all art units

Statute-Specific Performance

§101
25.0%
-15.0% vs TC avg
§103
37.2%
-2.8% vs TC avg
§102
9.6%
-30.4% vs TC avg
§112
26.1%
-13.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 94 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 6 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. In claim 6, line 6, it is unclear what “time to first symbol” means, and it is unclear whether this limitation is different from line 4. Examiner treats this limitation as “time taken to receive a first symbol”. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-7 and 11-20 are rejected under 35 U.S.C. 103 as being unpatentable over Barros (US 20250190466 A1) in view of Soldati et al. (US 20250071029 A1). Regarding claim 1, Barros teaches: A computer-implemented method comprising: transmitting a first input prompt to a first generative model; ([0086], lines 1-2, [0088], [0089], lines 1-6) receiving first symbols output from the first generative model responsive to the first input prompt; ([0089], lines 8-13) responsive to the latency being within a particular range, transmitting a second input prompt to a second generative model, the second input prompt based on at least some of the first symbols received from the first generative model; ([0012] and [0087], lines 9-17 disclose that the latency of the first generative model is less than the latency of the second generative model because the first generative has a smaller number of parameters. [0093]-[0094] disclose generating a text prompt for input to the second generative model.) receiving second symbols output from the second generative model responsive to the second input prompt; and ([0095]) providing output based on the first symbols and the second symbols. ([0089], lines 6-8 and [0099], lines 1-4) However, Barros does not explicitly teach: measuring a latency But Soldati teaches: measuring a latency ([0044], lines 6-8 and [0256]) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have applied Soldati’s verification that a model’s execution time is below a specified threshold to Baros’ first generative model. A motivation for the combination is to track execution time of the models, because shorter execution time gives users a better experience. Regarding claim 2, the combination of Barros and Soldati teaches: The computer-implemented method of claim 1, Barros teaches: wherein the first generative model is a first large language model (LLM), and the second generative model is a second LLM. ([0012], lines 1-5) Regarding claim 3, the combination of Barros and Soldati teaches: The computer-implemented method of claim 1, Barros teaches: wherein only a partially-completed response to the first input prompt has been received from the first generative model when the second input prompt is transmitted to the second generative model, the partially-completed response based on the first symbols, and wherein a remaining portion of the response to the first input prompt is based on the second symbols output from the second generative model. ([0093]-[0094]) Regarding claim 4, the combination of Barros and Soldati teaches: The computer-implemented method of claim 3, Barros teaches: wherein providing output based on the first symbols and the second symbols comprises: providing, for output on a display of a device, the partially-completed response based on the first symbols; and ([0089], lines 6-8) responsive to receiving at least a portion of the second symbols, providing for output on the display, the remaining portion of the response based on the second symbols, the remaining portion provided for output adjacent to and following the partially-completed response. ([0095] and [0099], lines 1-4) Regarding claim 5, the combination of Barros and Soldati teaches: The computer-implemented method of claim 1, Barros teaches: wherein the latency is However, Barros does not explicitly teach: the latency is measured But Soldati teaches: the latency is measured based on an amount of time to receive a [model output] ([0044], lines 6-8 and [0256]) A motivation for the combination is the same as the motivation given for claim 1. Regarding claim 6, the combination of Barros and Soldati teaches: The computer-implemented method of claim 5, Barros teaches: wherein 13 teaches the latency is based on time to generate text, which corresponds to “time to receive a symbol”.) However, Barros does not explicitly teach: wherein measuring the latency comprises measuring But Soldati teaches: wherein measuring the latency comprises measuring at least one of: … time to receive a [model output] ([0044], lines 6-8 and [0256]) A motivation for the combination is the same as the motivation given for claim 1. Regarding claim 7, the combination of Barros and Soldati teaches: The computer-implemented method of claim 1, Barros teaches: wherein the second input prompt is further based on the first input prompt. ([0093]-[0094) Regarding claim 11, the combination of Barros and Soldati teaches: The computer-implemented method of claim 1, Barros teaches: wherein the first generative model and the second generative model are at least one of: a same model; different instances of the same model; ([0019], lines 1-4 teaches the smaller LLM is a pruned version of the larger LLM. The smaller LLM is a different instance of the complete larger LLM.) a same architecture; fine-tuned in a same way; or have a same configuration setting. Regarding claim 12, the combination of Barros and Soldati teaches: The computer-implemented method of claim 1, Barros teaches: wherein the first generative model is [not] hosted by a particular software-as-a-service (SaaS) provider and the second generative model is not hosted by the particular SaaS provider. ([0053], lines 10-12 discloses both the smaller and larger LLMs can be local at the client device.) Barros at [0053], lines 8-15 teaches implementations wherein the smaller LLM is at the client device while the larger LLM is remote to the client devices, both the smaller and larger LLMs are local at the client device, and both the smaller and larger LLMs are remote to the client device. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have implemented Barros and Soldati’s system wherein the larger LLM is local at the client device and the smaller LLM is remote to the client device. This modification would give the system a larger selection of models to use than the models stored locally. Regarding claim 13, the combination of Barros and Soldati teaches: The computer-implemented method of claim 1, Barros teaches: wherein prior to transmitting the second input prompt to the second generative model, the method further comprises transmitting a prompt to the second generative model and generative model on training instances, and [0128] on page 16, col. 2, lines 6-9 indicates that training happens before deployment.) However, Barros does not explicitly teach: measuring the latency associated with receiving symbols from the second generative model in response to the prompt. But Soldati teaches: measuring the latency associated with receiving [a model output] ([0044], lines 6-8 and [0256]) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have applied Soldati’s verification that a model’s execution time is below a specified threshold to Baros’ second generative model. A motivation for the combination is to track execution time of the models, because shorter execution time gives users a better experience. Claim 14 recites a system which implements the same features as the method of claim 1 and is therefore rejected for at least the same reasons. Regarding claim 14, Barros teaches: A system comprising: at least one processor; and a memory storing processor-executable instructions that, when executed, cause the at least one processor to: ([0123] and [0124], lines 1-7) Claims 15-19 each recites a system which implements the same features as the method of claims 3-5, 7, and 11, respectively, and are therefore rejected for at least the same reasons. Claim 20 recites a product which implements the same features as the method of claim 1 and is therefore rejected for at least the same reasons. Regarding claim 20, Barros teaches: A non-transitory computer readable medium having stored thereon computer-executable instructions that, when executed by a computer, cause the computer to perform operations comprising: ([0123] and [0124], lines 1-7) Claims 8-10 are rejected under 35 U.S.C. 103 as being unpatentable over Barros (US 20250190466 A1) in view of Soldati et al. (US 20250071029 A1) and Mukherjee et al. (US 20240354436 A1). Regarding claim 8, the combination of Barros and Soldati teaches: The computer-implemented method of claim 7, Barros teaches: wherein the second input prompt is further based on content corresponding to an exchange between a device and the first generative model [including] However, Barros and Soldati do not explicitly teach: an exchange between a device and the first generative model prior to the first input prompt. But Mukherjee teaches: content corresponding to an exchange between a device and the first generative model prior to the first input prompt. ([0046] and [0072], lines 1-8 discloses exchanges between a user 150 via a user computing device and an LLM. A prompt for a LLM for responding to a user query is based on context, which includes previous user queries and responses from the LLM.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have applied Mukherjee’s technique (i.e., generating a prompt in response to a current user query and based on previous user queries and responses from the LLM) for generating a first response from the first generative model to the combination of Barros and Soldati. A motivation for the combination is to assist the LLM in generating output that is less prone to hallucination and more likely to meet the expectation of the user. (Mukherjee, [0046]) Regarding claim 9, the combination of Barros, Soldati, and Mukherjee teaches: The computer-implemented method of claim 8, However, Barros and Soldati do not explicitly teach: wherein the content provides a summary of the exchange. But Mukherjee teaches: wherein the content provides a summary of the exchange. ([0047]) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have applied Mukherjee’s technique of summarizing the context to the combination of Barros, Soldati, and Mukherjee. A motivation for the combination is to enable the system to provide the prompt for the LLM that is detailed enough without exceeding a size limit of a prompt window. (Mukherjee, [0047]) Regarding claim 10, the combination of Barros, Soldati, and Mukherjee teaches: The computer-implemented method of claim 9, Barros teaches: wherein the second input prompt is based on However, Barros and Soldati do not explicitly teach: wherein the second input prompt is based on both the summary of the exchange and input prompts and symbols transmitted between the device and the first generative model subsequent to the exchange But Mukherjee teaches: the summary of the exchange and input prompts and symbols transmitted between the device and the first generative model subsequent to the exchange In the combination of Barros, Soldati, and Mukherjee, the second prompt would be based on the first prompt, which includes Mukherjee’s summary of the conversation history. A motivation for the combination is the same as the motivations given for claims 8 and 9. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Asher H. Jablon whose telephone number is (571)270-7648. The examiner can normally be reached Monday - Friday, 9:00 am - 6:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Al Kawsar can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /A.H.J./Examiner, Art Unit 2127 /ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127
Read full office action

Prosecution Timeline

Feb 08, 2024
Application Filed
Jul 24, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12675727
METHOD AND SYSTEM FOR DETERMINING POLICIES, RULES, AND AGENT CHARACTERISTICS, FOR AUTOMATING AGENTS, AND PROTECTION
5y 10m to grant Granted Jul 07, 2026
Patent 12643559
NETWORK FOR DETECTING EDGE CASES FOR USE IN TRAINING AUTONOMOUS VEHICLE CONTROL SYSTEMS
1y 9m to grant Granted Jun 02, 2026
Patent 12626141
AUTOMATED GENERATION OF MACHINE LEARNING MODELS
3y 5m to grant Granted May 12, 2026
Patent 12614076
NEURAL NETWORK OPTIMIZATION DEVICE FOR EDGE DEVICE MEETING ON-DEMAND INSTRUCTION AND METHOD USING THE SAME
1y 9m to grant Granted Apr 28, 2026
Patent 12572794
SYSTEM AND METHOD FOR AUTOMATED OPTIMAZATION OF A NEURAL NETWORK MODEL
5y 4m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
43%
Grant Probability
87%
With Interview (+44.5%)
4y 4m (~1y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 94 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month