Prosecution Insights
Last updated: October 02, 2026
Application No. 19/009,733

Training method and apparatus for large models

Non-Final OA §102§103§112
Filed
Jan 03, 2025
Priority
Jan 04, 2024 — CN 202410010377.9
Examiner
SHAIKH, ZEESHAN MAHMOOD
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Alipay.com Co., Ltd.
OA Round
1 (Non-Final)
59%
Grant Probability
Moderate
1-2
OA Rounds
1y 4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 59% of resolved cases
59%
Career Allowance Rate
27 granted / 46 resolved
-3.3% vs TC avg
Strong +44% interview lift
Without
With
+44.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
20 currently pending
Career history
75
Total Applications
across all art units

Statute-Specific Performance

§101
26.2%
-13.8% vs TC avg
§103
47.6%
+7.6% vs TC avg
§102
16.5%
-23.5% vs TC avg
§112
4.3%
-35.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 46 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 1/3/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Specification The disclosure is objected to because of the following informalities: The examiner objects to the specification for not clearly defining “a large model”. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-10 and 12-21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Independent claims 1, 12, and 13 recite, “a large model”. The specification defines the large model in paragraph [0002] stating, “a large model refers to a model with a large quantity of parameters, for example, a deep neural network with more than 1 billion parameters, which can process a huge amount of data, complete various complex tasks, such as natural language processing, computer vision, and speech recognition”. “Large model” is generic and not a specific term of art such as a large language model (LLM) or a neural network. The examiner views large as a relative term. It is not clear if a model is large based on parameters or the number of layers. Is a model with a billion parameters considered a small model? The specification provides an example but not a specific definition. Therefore, the examiner views the claims as being indefinite. Claims 2-10 and 14-21 are also rejected because they are dependent on indefinite based claims 1 and 13. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-7 and 12-19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Lin et al. US 20230145452 A1 (hereinafter Lin). Regarding independent claims 1, 12, and 13, Lin teaches a training method for large models, wherein a large model comprises a first quantity of first network layers with a same first structure (FIG. 3, 302, 303), and the method comprises / a non-transitory computer-readable storage medium comprising instructions stored therein that, when executed by a processor of a computing device, causes the computing device to / a computing device, comprising a memory and a processor, wherein the memory stores executable instructions that, in response to execution by the processor (FIG. 1, 102-103), causes the computing device to: performing preliminary training on the large model under a first constraint condition, wherein the first constraint condition imposes a limitation that in a preliminary training process, different first network layers use same parameters ([0016] “The present disclosure plans to divide ultra-large-scale training (including pre-training) into two stages for execution. In the first stage, a model only uses a small number of weight parameters (referred to as weights, weight coefficients, or weight parameters in some articles) to perform efficient training to realize model convergence. In the second stage, training is continuously performed according to the model obtained through the first stage, so the model may converge at a relatively low loss level, thereby greatly reducing training steps required for a large model to converge. In this condition, training of the second stage may adopt a CPU offload mode”); and when the limitation imposed by the first constraint condition is removed, performing further training on the large model obtained after the preliminary training ([0019] “Before the pre-trained model is applicable to a real scene, incremental training may be further performed on the pre-trained model by using sample data acquired in the real scene, to finely adjust the weight parameter in the pre-trained model through the incremental training, so as to obtain a “completed” model applicable to the real scene, which is generally referred to as a trained model”). Regarding claims 2 and 14, Lin teaches all of the limitations of claim 1 and 13, upon which claims 2 and 14 depend. Additionally, Lin teaches wherein the first structure comprises a first network part and a second network part (FIG. 4, S01, S03; [0041] “In step S01, a first model is trained to obtain a parameter set of the trained first model”; [0043] “In step S03, the second model is trained to realize model convergence”); the further training comprises a first sub-training with a second constraint condition and a second sub-training in which the second constraint condition is removed, wherein the first sub-training and the second sub-training are successively performed; and the second constraint condition imposes a limitation that first network parts of different first network layers use same parameters in a sub-training process ([0052] “after each round of training based on step S01 is finished, whether the error between the result and the expected result of the first model satisfies a set condition (for example, a decrease amplitude of the error is less than a set standard) is determined. If so, step S01 is ended and step S02 starts to be executed. Otherwise, a next round of training is performed based on step S01”, examiner interprets the condition as the constraint). Regarding claims 3 and 15, Lin teaches all of the limitations of claim 2 and 14, upon which claims 3 and 15 depend. Additionally, Lin teaches wherein the large model is specifically a multimodal large model applicable to a picture mode (FIG. 2, 202, [0032] “the dedicated processing unit 202 is a graphics processing unit (GPU). In another example, the dedicated processing unit 202 is a GPU. The GPU is a microprocessor dedicated to image and graphics related operation works”) and a text mode ([0031] “the training samples are not limited to multi-modality or single-modality. For example, each piece of visual information, text information, and audio information is referred to as one modality”), the first network part comprises a self-attention sublayer, and the second network part comprises a first feed-forward neural network sublayer corresponding to the picture mode and a second feed-forward neural network sublayer corresponding to the text mode (FIG. 3, 3011, 3013). Regarding claims 4 and 16, Lin teaches all of the limitations of claim 1 and 13, upon which claims 4 and 16 depend. Additionally, Lin teaches wherein the large model further comprises a second quantity of second network layers with a same second structure ([0071-0073] “copying the parameter set for a plurality of times as weight parameters of a plurality of second layers of a second model; and training the second model to realize model convergence, wherein the first model and the second model have a same computation graph, and the number of the plurality of second layers is equal to or greater than the number of the plurality of first layers”); and the first constraint condition imposes a further limitation that in the preliminary training process, different second network layers use same parameters ([0078] “after training the first model and before copying the parameter set for the plurality of times, determining whether an error between a result and an expected result of the first model satisfies a set condition”). Regarding claim 5 and 17, Lin teaches all of the limitations of claim 4 and 16, upon which claims 5 and 17 depend. Additionally, Lin teaches wherein the second structure comprises a third network part and a fourth network part; the further training comprises a first sub-training with a second constraint condition and a second sub-training in which the second constraint condition is removed, wherein first sub-training and the second sub-training are successively performed ([0045] “The neural network model structure includes a plurality of operators (or symbolic expressions of the operators) and a connection relationship thereof, which may be shown in a patterning manner, so that the neural network model structure is referred to as a static computation graph”, examiner interprets operators as the network parts); and the second constraint condition imposes a limitation that third network parts of different second network layers use same parameters in a sub-training process ([0052] “after each round of training based on step S01 is finished, whether the error between the result and the expected result of the first model satisfies a set condition (for example, a decrease amplitude of the error is less than a set standard) is determined. If so, step S01 is ended and step S02 starts to be executed. Otherwise, a next round of training is performed based on step S01”). Regarding claim 6 and 18, Lin teaches all of the limitations of claim 5 and 17, upon which claims 6 and 18 depend. Additionally, Lin teaches wherein the first structure comprises a first network part and a second network part (FIG. 4, S01, S03; [0041] “In step S01, a first model is trained to obtain a parameter set of the trained first model”; [0043] In step S03, the second model is trained to realize model convergence); and the second constraint condition imposes a further limitation that first network parts of different first network layers use same parameters in a sub-training process ([0052] “after each round of training based on step S01 is finished, whether the error between the result and the expected result of the first model satisfies a set condition (for example, a decrease amplitude of the error is less than a set standard) is determined. If so, step S01 is ended and step S02 starts to be executed. Otherwise, a next round of training is performed based on step S01”, examiner interprets the condition as the constraint;). Regarding claim 7 and 19, Lin teaches all of the limitations of claim 6 and 18, upon which claims 7 and 19 depend. Additionally, Lin teaches wherein the large model is specifically a multimodal large model applicable to a picture mode and a text mode, the first network part comprises a self-attention sublayer, and the second network part comprises a first feed-forward neural network sublayer corresponding to the picture mode and a second feed-forward neural network sublayer corresponding to the text mode (FIG. 2, 202, [0032] “the dedicated processing unit 202 is a graphics processing unit (GPU). In another example, the dedicated processing unit 202 is a GPU. The GPU is a microprocessor dedicated to image and graphics related operation works”, [0031] “the training samples are not limited to multi-modality or single-modality. For example, each piece of visual information, text information, and audio information is referred to as one modality”, FIG. 3, 3011, 3013); the third network part is a self-attention sublayer shared by two modes; and the fourth network part comprises a third feed-forward neural network sublayer shared by the two modes ([0035] “The processing units are shown in the figure, which include a multi-head self-attention mechanism layer 3011, a summation and normalization layer 3012, a feedforward neural network layer 3013, and a summation and normalization layer 3014”; [0036] Examiner interprets functions as the modes). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 8-10 and 20-12 are rejected under 35 U.S.C. 103 as being unpatentable over Lin in view of Niu et al. US 20220292269 A1 (hereinafter Niu). Regarding claims 8 and 20 Lin teaches all of the limitations of claims 1 and 13, upon which claims 8 and 20 depend. Lin fails to teach wherein the large model is a multimodal large model applicable to a picture mode and a text mode, an input of the large model comprises a first initial vector of the picture mode and a second initial vector of the text mode, and an output of the large model comprises a first fusion vector of the picture mode and a second fusion vector of the text mode; the first initial vector comprises a picture embedding vector of a sample picture and block embedding vectors respectively corresponding to a plurality of image blocks in the sample picture, and the second initial vector comprises a sentence embedding vector of a sample sentence and word embedding vectors respectively corresponding to a plurality of segments in the sample sentence; and the first fusion vector comprises a picture fusion vector of the sample picture and block fusion vectors respectively corresponding to the plurality of image blocks, and the second fusion vector comprises a sentence fusion vector of the sample sentence and word fusion vectors respectively corresponding to the plurality of segments. However, Niu teaches wherein the large model is a multimodal large model applicable to a picture mode and a text mode, an input of the large model comprises a first initial vector of the picture mode and a second initial vector of the text mode, and an output of the large model comprises a first fusion vector of the picture mode and a second fusion vector of the text mode; the first initial vector comprises a picture embedding vector of a sample picture and block embedding vectors respectively corresponding to a plurality of image blocks in the sample picture, and the second initial vector comprises a sentence embedding vector of a sample sentence and word embedding vectors respectively corresponding to a plurality of segments in the sample sentence; and the first fusion vector comprises a picture fusion vector of the sample picture and block fusion vectors respectively corresponding to the plurality of image blocks, and the second fusion vector comprises a sentence fusion vector of the sample sentence and word fusion vectors respectively corresponding to the plurality of segments ([0056] “The similarity between the image and the text in the image-text pair obtained by the rewriting extension is determined by: stitching the image and the text, and mapping a vector representation of the stitched language material obtained by the pre-trained model into a similarity value. This way of calculating the similarity is referred to as a “single-flow” way. In this way, the sequence of the image and the sequence of the text are stitched and then input into the pre-trained model. The pre-trained model obtains an overall vector representation for the stitched sequence”, examiner interprets the overall vector as the fusion vector.) Lin in view of Niu are considered to be analogous to the claimed invention because both are the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the techniques of training a model of Lin with the technique of generating fusion text and image vectors taught by Niu in order to improve deep learning technologies in the field of artificial intelligence technologies (see Niu [0002]). Regarding claims 9 and 21, Lin in view of Niu teaches all of the limitations of claim 8 and 20, upon which claims 9 and 21 depend. Additionally, Niu teaches wherein the preliminary training and/or the further training comprise/comprises the following training manner: adjusting a model parameter by maximizing a score of similarity between a sample picture and a sample sentence comprised in a positive sample pair and minimizing a score of similarity between a sample picture and a sample sentence comprised in a negative sample pair, wherein a similarity score is determined based on vector similarity between a picture fusion vector of the sample picture and a sentence fission vector of the sample sentence ([0056] “The pre-trained model obtains an overall vector representation for the stitched sequence, and the similarity value is obtained after the vector representation is mapped (for example, Softmax)”; [0057] The similarity between the image and the text in the image-text pair obtained by the retrieval extension is determined by: calculating similarity between a vector representation of the image obtained by the pre-trained model and a vector representation of the text obtained by the pre-trained model). Regarding claim 10, Lin in view of Niu teaches all of the limitations of claim 8, upon which claim 10 depend. Additionally, Niu teaches wherein the preliminary training and/or the further training comprise/comprises the following training manner: randomly masking block embedding vectors corresponding to a part of image blocks in the first initial vector, or randomly masking word embedding vectors corresponding to a part of segments in the second initial vector, predicting a masked image block or segment through an output of the large model, and adjusting a model parameter based on a predicted masked object and an actual masked object ([0062] “A region is randomly selected from the image to be masked, and the masked region is reconstructed using a vector representation of an unmasked region by pre-trained model, with a training target of minimizing a difference between the reconstructed region and the masked region in the image”). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Baughman et al. (US 20250200327 A1) teaches a method, system, and computer program product for adaptatively training a large language model based on predicted changes in computing resources. A processor may identify data elements of computing resources for training a large language model. The identified data elements include a configured model topology. The processor may generate a forecast vector capturing a predicted change in the configured model topology over a time period. The processor may execute, using an optimization algorithm and the forecast vector, a series of optimization experiments using each training type of a plurality of training types to determine an optimal computing resource supply and demand over the time period. The processor may determine, based on the series of optimization experiments, an accuracy penalty value associated with each training type. The processor may train the large language model using a first training type that has a lowest accuracy penalty. Zhao (US 20210133535 A1) teaches techniques for auto composing using a transformer-based language model having a parameter sharing decoder pair (PSDP) that reduces a number of parameters of the model and maintains the capability of generating understandable and reasonable compositions. In one particular aspect, a method is provided that includes obtaining a full encoder sequence, and inputting the full encoder sequence into a transformer model having a PSDP. The PSDP includes: a first decoder having parameters that are shared across all N layers of the first decoder; and a second decoder having parameters that are shared across all N layers of the first decoder. The parameters of the first decoder are different from the parameters of the second decoder. The method further includes using the transformer model to predict sequence elements based on the full encoder sequence, generate an output sequence comprising the sequence elements, and output an output sequence different from the full encoder sequence. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZEESHAN SHAIKH whose telephone number is (703)756-1730. The examiner can normally be reached Monday-Friday 7:30AM-5:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ZEESHAN MAHMOOD SHAIKH/Examiner, Art Unit 2658 /RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Jan 03, 2025
Application Filed
Aug 24, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748800
LOW-RESOURCE, MULTI-LINGUAL TRANSFORMER MODELS
4y 3m to grant Granted Sep 29, 2026
Patent 12737562
SYSTEMS AND METHODS FOR TRANSLATION EVALUATION
4y 4m to grant Granted Sep 15, 2026
Patent 12694235
Adaptable Transformer Models via Key Term Replacement
3y 10m to grant Granted Jul 28, 2026
Patent 12646510
DETERMINING WHETHER SPEECH INPUT IS INTENDED FOR A DIGITAL ASSISTANT
3y 8m to grant Granted Jun 02, 2026
Patent 12633299
LINEAR PREDICTION CODING PARAMETER CODING METHOD AND CODING APPARATUS
3y 6m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
59%
Grant Probability
99%
With Interview (+44.3%)
3y 1m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 46 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month