Prosecution Insights
Last updated: October 02, 2026
Application No. 18/418,624

MOE MODEL FROM BLOCK SPARSE COMPUTATION'S POINT OF VIEW

Non-Final OA §103
Filed
Jan 22, 2024
Examiner
TRAN, DAVID HOANG
Art Unit
Tech Center
Assignee
Tencent Technology (Shenzhen) Company Limited
OA Round
1 (Non-Final)
23%
Grant Probability
At Risk
1-2
OA Rounds
1y 9m
Est. Remaining
46%
With Interview

Examiner Intelligence

Grants only 23% of cases
23%
Career Allowance Rate
5 granted / 22 resolved
-37.3% vs TC avg
Strong +23% interview lift
Without
With
+23.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
26 currently pending
Career history
57
Total Applications
across all art units

Statute-Specific Performance

§101
28.1%
-11.9% vs TC avg
§103
51.2%
+11.2% vs TC avg
§102
7.9%
-32.1% vs TC avg
§112
11.8%
-28.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 22 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Drawings The drawings have been received on 01/22/2024. These drawings are accepted. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-9, 11-18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (MoEfication: Transformer Feed-forward Layers are Mixtures of Experts); hereinafter Zhang in view of Nie et al. (EvoMoE: An Evolutional Mixture-of-Experts Training Framework via Dense-To-Sparse Gate); hereinafter Nie Claim 1 is rejected under Zhang and Nie. Regarding claim 1, Zhang teaches a method performed by at least one processor, the method comprising: forming one or more routers of a mixture of experts (MoE) model by computing router parameters based on a plurality of MoE weights; (Zhang [page 4, 3.3. Expert Selection]: “we introduce how to create a router for expert selection … Similarity Selection. To utilize the parameter information, we average all columns of W 1 i and use it as the expert representation. Given an input x, we calculate the cosine similarity between the expert representation and x as si.”) implementing at least one of the routers of the MoE model to compute one or more outputs by using a number of the MoE weights less than that of the plurality of MoE weights; and (Zhang [page 4, 3.3 Expert Selection]: “In this subsection, we introduce how to create a router for expert selection. An MoEfied FFN processed an input x by F m x =   ∑ i ∈ S o x W 1 i + b 1 i W 2 i + b 2 i , (7) where S is the set of the selected experts. If all experts are selected, we have F m x = F ( x ) . Considering that o x W 1 i + b 1 i W 2 i equals to 0 for most experts, we try to select n experts, where n < k,) deriving an MoE expert [based on iteratively updating the router parameters] according to the one or more outputs. (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores”) Zhang does not appear to explicitly teach based on iteratively updating the router parameters However, Nie teaches based on iteratively updating the router parameters (Nie [page 6]: “Algorithm 1 illustrates the MoE training process in our EvoMoE framework. First, input tokens are processed by the shared expert e0 (line 1-2). Then EvoMoE switches the training into standard MoE models’ training, by adding a gate network at each MoE layer and diversifying all experts from the shared expert (line 4-5). After this expert-diversify phase, EvoMoE steps into the gate-sparsify phase, where it schedules the gate temperature coefficients and then obtains the token to-expert routing relation from DTS-gate (line 7-9). Tokens will be dispatched to corresponding experts and aggregated together by weighted sum operating (line 10-12).”; and [page 5, A. Problem Formulation]: “Given an input token xs, a series of experts {e1,...,eN} and a learn-able gate with parameter Wg, Func is adopted by the gate network to determine the targeted experts for it, i.e., the token-to-expert assignment, formulated in Equation 6. g(xs)is a 1×N vector, which represents the scores of xs with respect to experts. Meanwhile, each expert will process the input token separately as ei(xs) and combine their output as Equation 7.”; Note: See the iterative process of Algorithm 1.) It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the iterative mixture-of-expert framework of Nie to reduce computation cost and improve model quality (Nie, page 2). Zhang and Nie are analogous art because they both concern improving mixture-of-expert models. Claim 2 is rejected under Zhang and Nie with the incorporation of claim 1. Regarding claim 2, Zhang does not appear to explicitly teach wherein a first input layer of a first iteration of iteratively updating the router parameters comprises N=1 where N is a total number of the one or more routers. However, Nie teaches wherein a first input layer of a first iteration of iteratively updating the router parameters comprises N=1 where N is a total number of the one or more routers. (Nie [page 6]: “we train one shared-expert instead of N individual experts in the early stage (illustrated as the left of Figure 1). Because all experts within the same MoE layer share weights, the model is equal to its corresponding non MoE model as a small dense model.”; Note: See Algorithm 1 for the iteration lines 1-5) It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the iterative mixture-of-expert framework of Nie to reduce computation cost and improve model quality (Nie, page 2). Zhang and Nie are analogous art because they both concern improving mixture-of-expert models. Claim 3 is rejected under Zhang and Nie with the incorporation of claim 1. Regarding claim 3, Zhang teaches wherein a total number of the plurality of MoE weights is n, (Zhang [page 3, 3.2 Expert Construction]: “the number of experts is k”; and [page 4, 3.1 Overall Framework]: “where W 1 and W 2 are the weight matrices”) wherein a total number of the one or more outputs of the one or more routers is K and (Zhang [page 3, 3.2 Expert Construction]: “the number of experts is k”) a total number of the MoE weights less than that of the plurality of MoE weights is k, and (Zhang [page 4, 3.3 Expert Selection]: “An MoEfied FFN processed an input x by F m x =   ∑ i ∈ S o x W 1 i + b 1 i W 2 i + b 2 i , (7) where S is the set of the selected experts. If all experts are selected, we have F m x = F ( x ) . Considering that o x W 1 i + b 1 i W 2 i equals to 0 for most experts, we try to select n experts, where n < k,) Zhang does not appear to explicitly teach wherein updating the router parameters according to the one or more outputs comprises updating the router parameters by replacing the total number of the plurality of MoE weights n total number of the MoE weights k less than that of the plurality of MoE weights n. However, Nie teaches wherein updating the router parameters according to the one or more outputs comprises updating the router parameters by replacing the total number of the plurality of MoE weights n total number of the MoE weights k less than that of the plurality of MoE weights n. (See Algorithm 2 of Nie to see that line 2 computes the output scores, lines 3-5 are the iterative process and lines 6-8 are selecting 1 expert which is k less than the plurality of MoE weights of n.) It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the iterative mixture-of-expert framework of Nie to reduce computation cost and improve model quality (Nie, page 2). Zhang and Nie are analogous art because they both concern improving mixture-of-expert models. Claim 4 is rejected under Zhang and Nie with the incorporation of claim 1. Regarding claim 4, Zhang teaches deriving the MoE expert further comprises estimating an output block index k by computing which of a plurality of blocks has a highest sum of activations. (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.) Claim 5 is rejected under Zhang and Nie with the incorporation of claim 1. Regarding claim 5, Zhang teaches wherein computing which of the plurality of blocks has the highest sum of activations is based on each of the plurality of blocks accumulated activations, (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.) a nonlinear activation function, (Zhang [page 4, 3.1 Overall Framework]: “σ(·) is a non-linear activation function”) the MoE weights, (Zhang [page 4, 3.1 Overall Framework]: “where W 1 and W 2 are the weight matrices”) an input value of an input block, and (Zhang [page 4, 3.1 Overall Framework]: “for a given input representation x”) a bias corresponding to the one or more outputs. (Zhang [page 4, 3.1 Overall Framework]: " b 1 and b 2 are the bias vectors”) Claim 6 is rejected under Zhang and Nie with the incorporation of claim 1. Regarding claim 6, Zhang teaches wherein each of the plurality of blocks accumulated activations is Sn k determined by snk = ∑j f (∑i wnk,ij xn,i + bnk, j) (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.) where f(.) represents the nonlinear activation function, (Zhang [page 4, 3.1 Overall Framework]: “σ(·) is a non-linear activation function”) wnk,ij represents the MoE weights, (Zhang [page 4, 3.1 Overall Framework]: “where W 1 and W 2 are the weight matrices”) xn,i represents the input value of the input block, and (Zhang [page 4, 3.1 Overall Framework]: “for a given input representation x”) bnk, j represents the bias. (Zhang [page 4, 3.1 Overall Framework]: " b 1 and b 2 are the bias vectors”) Claim 7 is rejected under Zhang and Nie with the incorporation of claim 1. Regarding claim 7, Zhang teaches wherein the MoE expert is derived further based on a sum of activations of the MoE expert. (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.) Claim 8 is rejected under Zhang and Nie with the incorporation of claim 1. Regarding claim 8, Zhang teaches wherein the MoE expert is derived as one of a plurality of blocks of the MoE that is estimated as having a highest sum of activations among the plurality of blocks. (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.) Claim 9 is rejected under Zhang and Nie with the incorporation of claim 1. Regarding claim 19, Zhang teaches wherein the MoE expert is implemented in a large language model and (Zhang [page 4, 4.1 Experimental Setups]: “Models and Hyperparameter. We use four variants of T5 (Raffel et al., 2020), which are the 60-million-parameter T5-Small, the 200-millionparameter T5-Base, the 700-million-parameter T5-Large, and the 3-billion-parameter T5-XLarge.”; Note: T5 (Text-to-Text Transfer Transformer) is a large language model.) Claim 11 is rejected under Zhang and Nie. Regarding claim 11, Zhang teaches an apparatus comprising: at least one memory configured to store computer program code; at least one processor configured to access the computer program code and operate as instructed by the computer program code, the computer program code including: (Zhang [page 8]: “MoEfication can significantly reduce the total FLOPS, such as 2x speedup in the ratio of 25%. Meanwhile, the speedup on CPU is close to that on FLOPS. Considering that CPU is widely used for model inference in real-world scenarios, MoEfication is practical for the acceleration of various NLP applications.”) The remainder of claim 11 is claim 1 in the form of an apparatus and is rejected for the same reasons as claim 1 stated above. Dependent claim 12 is claim 2 in the form of an apparatus and is rejected for the same reasons as claim 2 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above. Dependent claim 13 is claim 3 in the form of an apparatus and is rejected for the same reasons as claim 3 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above. Dependent claim 14 is claim 4 in the form of an apparatus and is rejected for the same reasons as claim 4 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above. Dependent claim 15 is claim 5 in the form of an apparatus and is rejected for the same reasons as claim 5 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above. Dependent claim 16 is claim 6 in the form of an apparatus and is rejected for the same reasons as claim 6 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above. Dependent claim 17 is claim 7 in the form of an apparatus and is rejected for the same reasons as claim 7 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above. Dependent claim 18 is claim 8 in the form of an apparatus and is rejected for the same reasons as claim 8 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above. Claim 20 is rejected under Zhang and Nie. Regarding claim 20, Zhang teaches a non-transitory computer readable medium storing a program causing a computer to: (Zhang [page 8]: “MoEfication can significantly reduce the total FLOPS, such as 2x speedup in the ratio of 25%. Meanwhile, the speedup on CPU is close to that on FLOPS. Considering that CPU is widely used for model inference in real-world scenarios, MoEfication is practical for the acceleration of various NLP applications.”) The remainder of claim 20 is claim 1 in the form of a non-transitory computer readable medium and is rejected for the same reasons as claim 1 stated above. Claims 10 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang and Nie in view of Riquelme et al. (Scaling Vision with Sparse Mixture of Experts); hereinafter Riquelme Claim 10 is rejected under Zhang, Nie and Riquelme with the incorporation of claim 1. Regarding claim 10, Zhang does not appear to explicitly teach wherein the MoE expert is implemented in an image model. However, Riquelme teaches wherein the MoE expert is implemented in an image model. (Riquelme [page 1, 1 Introduction]: “We introduce the Vision MoE(V-MoE), a sparse variant of the recent Vision Transformer (ViT) architecture [20] for image classification. The V-MoE replaces a subset of the dense feedforward layers in ViT with sparse MoE layers, where each image patch is “routed” to a subset of “experts” (MLPs).”) It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the image model of Riquelme for smooth performance and reduced training time (Riquelme, page 1, 1 Introduction). Zhang and Riquelme are analogous art because they both concern improving mixture-of-expert models. Claim 19 is rejected under Zhang, Nie and Riquelme with the incorporation of claim 11. Regarding claim 19, Zhang teaches wherein the MoE expert is implemented in a large language model and (Zhang [page 4, 4.1 Experimental Setups]: “Models and Hyperparameter. We use four variants of T5 (Raffel et al., 2020), which are the 60-million-parameter T5-Small, the 200-millionparameter T5-Base, the 700-million-parameter T5-Large, and the 3-billion-parameter T5-XLarge.”; Note: T5 (Text-to-Text Transfer Transformer) is a large language model.) Zhang does not appear to explicitly teach an image model. However, Riquelme teaches an image model. (Riquelme [page 1, 1 Introduction]: “We introduce the Vision MoE(V-MoE), a sparse variant of the recent Vision Transformer (ViT) architecture [20] for image classification. The V-MoE replaces a subset of the dense feedforward layers in ViT with sparse MoE layers, where each image patch is “routed” to a subset of “experts” (MLPs).”) It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the image model of Riquelme for smooth performance and reduced training time (Riquelme, page 1, 1 Introduction). Zhang and Riquelme are analogous art because they both concern improving mixture-of-expert models. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID H TRAN whose telephone number is (703)756-1525. The examiner can normally be reached M-F 9:30 am - 5:30 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DAVID H TRAN/Examiner, Art Unit 2147 /ERIC NILSSON/Primary Examiner, Art Unit 2151
Read full office action

Prosecution Timeline

Jan 22, 2024
Application Filed
Sep 17, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12632724
CANONICALIZATION OF DATA WITHIN OPEN KNOWLEDGE GRAPHS
4y 8m to grant Granted May 19, 2026
Patent 12579404
PROCESSOR FOR NEURAL NETWORK, PROCESSING METHOD FOR NEURAL NETWORK, AND NON-TRANSITORY COMPUTER READABLE STORAGE MEDIUM
4y 2m to grant Granted Mar 17, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
23%
Grant Probability
46%
With Interview (+23.0%)
4y 5m (~1y 9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 22 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month