Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings have been received on 01/22/2024. These drawings are accepted.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-9, 11-18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (MoEfication: Transformer Feed-forward Layers are Mixtures of Experts); hereinafter Zhang in view of Nie et al. (EvoMoE: An Evolutional Mixture-of-Experts Training Framework via Dense-To-Sparse Gate); hereinafter Nie
Claim 1 is rejected under Zhang and Nie.
Regarding claim 1, Zhang teaches a method performed by at least one processor, the method comprising:
forming one or more routers of a mixture of experts (MoE) model by computing router parameters based on a plurality of MoE weights; (Zhang [page 4, 3.3. Expert Selection]: “we introduce how to create a router for expert selection … Similarity Selection. To utilize the parameter information, we average all columns of
W
1
i
and use it as the expert representation. Given an input x, we calculate the cosine similarity between the expert representation and x as si.”)
implementing at least one of the routers of the MoE model to compute one or more outputs by using a number of the MoE weights less than that of the plurality of MoE weights; and (Zhang [page 4, 3.3 Expert Selection]: “In this subsection, we introduce how to create a router for expert selection. An MoEfied FFN processed an input x by
F
m
x
=
∑
i
∈
S
o
x
W
1
i
+
b
1
i
W
2
i
+
b
2
i
, (7) where S is the set of the selected experts. If all experts are selected, we have
F
m
x
=
F
(
x
)
. Considering that o
x
W
1
i
+
b
1
i
W
2
i
equals to 0 for most experts, we try to select n experts, where n < k,)
deriving an MoE expert [based on iteratively updating the router parameters] according to the one or more outputs. (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores”)
Zhang does not appear to explicitly teach based on iteratively updating the router parameters
However, Nie teaches based on iteratively updating the router parameters (Nie [page 6]: “Algorithm 1 illustrates the MoE training process in our EvoMoE framework. First, input tokens are processed by the shared expert e0 (line 1-2). Then EvoMoE switches the training into standard MoE models’ training, by adding a gate network at each MoE layer and diversifying all experts from the shared expert (line 4-5). After this expert-diversify phase, EvoMoE steps into the gate-sparsify phase, where it schedules the gate temperature coefficients and then obtains the token to-expert routing relation from DTS-gate (line 7-9). Tokens will be dispatched to corresponding experts and aggregated together by weighted sum operating (line 10-12).”; and [page 5, A. Problem Formulation]: “Given an input token xs, a series of experts {e1,...,eN} and a learn-able gate with parameter Wg, Func is adopted by the gate network to determine the targeted experts for it, i.e., the token-to-expert assignment, formulated in Equation 6. g(xs)is a 1×N vector, which represents the scores of xs with respect to experts. Meanwhile, each expert will process the input token separately as ei(xs) and combine their output as Equation 7.”; Note: See the iterative process of Algorithm 1.)
It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the iterative mixture-of-expert framework of Nie to reduce computation cost and improve model quality (Nie, page 2). Zhang and Nie are analogous art because they both concern improving mixture-of-expert models.
Claim 2 is rejected under Zhang and Nie with the incorporation of claim 1.
Regarding claim 2, Zhang does not appear to explicitly teach wherein a first input layer of a first iteration of iteratively updating the router parameters comprises N=1 where N is a total number of the one or more routers.
However, Nie teaches wherein a first input layer of a first iteration of iteratively updating the router parameters comprises N=1 where N is a total number of the one or more routers. (Nie [page 6]: “we train one shared-expert instead of N individual experts in the early stage (illustrated as the left of Figure 1). Because all experts within the same MoE layer share weights, the model is equal to its corresponding non MoE model as a small dense model.”; Note: See Algorithm 1 for the iteration lines 1-5)
It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the iterative mixture-of-expert framework of Nie to reduce computation cost and improve model quality (Nie, page 2). Zhang and Nie are analogous art because they both concern improving mixture-of-expert models.
Claim 3 is rejected under Zhang and Nie with the incorporation of claim 1.
Regarding claim 3, Zhang teaches wherein a total number of the plurality of MoE weights is n, (Zhang [page 3, 3.2 Expert Construction]: “the number of experts is k”; and [page 4, 3.1 Overall Framework]: “where
W
1
and
W
2
are the weight matrices”)
wherein a total number of the one or more outputs of the one or more routers is K and (Zhang [page 3, 3.2 Expert Construction]: “the number of experts is k”)
a total number of the MoE weights less than that of the plurality of MoE weights is k, and (Zhang [page 4, 3.3 Expert Selection]: “An MoEfied FFN processed an input x by
F
m
x
=
∑
i
∈
S
o
x
W
1
i
+
b
1
i
W
2
i
+
b
2
i
, (7) where S is the set of the selected experts. If all experts are selected, we have
F
m
x
=
F
(
x
)
. Considering that o
x
W
1
i
+
b
1
i
W
2
i
equals to 0 for most experts, we try to select n experts, where n < k,)
Zhang does not appear to explicitly teach wherein updating the router parameters according to the one or more outputs comprises updating the router parameters by replacing the total number of the plurality of MoE weights n total number of the MoE weights k less than that of the plurality of MoE weights n.
However, Nie teaches wherein updating the router parameters according to the one or more outputs comprises updating the router parameters by replacing the total number of the plurality of MoE weights n total number of the MoE weights k less than that of the plurality of MoE weights n. (See Algorithm 2 of Nie to see that line 2 computes the output scores, lines 3-5 are the iterative process and lines 6-8 are selecting 1 expert which is k less than the plurality of MoE weights of n.)
It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the iterative mixture-of-expert framework of Nie to reduce computation cost and improve model quality (Nie, page 2). Zhang and Nie are analogous art because they both concern improving mixture-of-expert models.
Claim 4 is rejected under Zhang and Nie with the incorporation of claim 1.
Regarding claim 4, Zhang teaches deriving the MoE expert further comprises estimating an output block index k by computing which of a plurality of blocks has a highest sum of activations. (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.)
Claim 5 is rejected under Zhang and Nie with the incorporation of claim 1.
Regarding claim 5, Zhang teaches wherein computing which of the plurality of blocks has the highest sum of activations is based on each of the plurality of blocks accumulated activations, (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.)
a nonlinear activation function, (Zhang [page 4, 3.1 Overall Framework]: “σ(·) is a non-linear activation function”)
the MoE weights, (Zhang [page 4, 3.1 Overall Framework]: “where
W
1
and
W
2
are the weight matrices”)
an input value of an input block, and (Zhang [page 4, 3.1 Overall Framework]: “for a given input representation x”)
a bias corresponding to the one or more outputs. (Zhang [page 4, 3.1 Overall Framework]:
"
b
1
and
b
2
are the bias vectors”)
Claim 6 is rejected under Zhang and Nie with the incorporation of claim 1.
Regarding claim 6, Zhang teaches wherein each of the plurality of blocks accumulated activations is Sn k determined by snk = ∑j f (∑i wnk,ij xn,i + bnk, j) (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.)
where f(.) represents the nonlinear activation function, (Zhang [page 4, 3.1 Overall Framework]: “σ(·) is a non-linear activation function”)
wnk,ij represents the MoE weights, (Zhang [page 4, 3.1 Overall Framework]: “where
W
1
and
W
2
are the weight matrices”)
xn,i represents the input value of the input block, and (Zhang [page 4, 3.1 Overall Framework]: “for a given input representation x”)
bnk, j represents the bias. (Zhang [page 4, 3.1 Overall Framework]:
"
b
1
and
b
2
are the bias vectors”)
Claim 7 is rejected under Zhang and Nie with the incorporation of claim 1.
Regarding claim 7, Zhang teaches wherein the MoE expert is derived further based on a sum of activations of the MoE expert. (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.)
Claim 8 is rejected under Zhang and Nie with the incorporation of claim 1.
Regarding claim 8, Zhang teaches wherein the MoE expert is derived as one of a plurality of blocks of the MoE that is estimated as having a highest sum of activations among the plurality of blocks. (Zhang [page 4, Expert Selection]: “The selection methods will assign a score si to each expert for the given input x and select the experts with the n highest scores by (8) … We calculate the sum of positive values in each expert as si and select experts using Equation 8.)
Claim 9 is rejected under Zhang and Nie with the incorporation of claim 1.
Regarding claim 19, Zhang teaches wherein the MoE expert is implemented in a large language model and (Zhang [page 4, 4.1 Experimental Setups]: “Models and Hyperparameter. We use four variants of T5 (Raffel et al., 2020), which are the 60-million-parameter T5-Small, the 200-millionparameter T5-Base, the 700-million-parameter T5-Large, and the 3-billion-parameter T5-XLarge.”; Note: T5 (Text-to-Text Transfer Transformer) is a large language model.)
Claim 11 is rejected under Zhang and Nie.
Regarding claim 11, Zhang teaches an apparatus comprising:
at least one memory configured to store computer program code;
at least one processor configured to access the computer program code and operate as instructed by the computer program code, the computer program code including: (Zhang [page 8]: “MoEfication can significantly reduce the total FLOPS, such as 2x speedup in the ratio of 25%. Meanwhile, the speedup on CPU is close to that on FLOPS. Considering that CPU is widely used for model inference in real-world scenarios, MoEfication is practical for the acceleration of various NLP applications.”)
The remainder of claim 11 is claim 1 in the form of an apparatus and is rejected for the same reasons as claim 1 stated above.
Dependent claim 12 is claim 2 in the form of an apparatus and is rejected for the same reasons as claim 2 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above.
Dependent claim 13 is claim 3 in the form of an apparatus and is rejected for the same reasons as claim 3 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above.
Dependent claim 14 is claim 4 in the form of an apparatus and is rejected for the same reasons as claim 4 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above.
Dependent claim 15 is claim 5 in the form of an apparatus and is rejected for the same reasons as claim 5 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above.
Dependent claim 16 is claim 6 in the form of an apparatus and is rejected for the same reasons as claim 6 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above.
Dependent claim 17 is claim 7 in the form of an apparatus and is rejected for the same reasons as claim 7 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above.
Dependent claim 18 is claim 8 in the form of an apparatus and is rejected for the same reasons as claim 8 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 11, see the rejection of claim 11 above.
Claim 20 is rejected under Zhang and Nie.
Regarding claim 20, Zhang teaches a non-transitory computer readable medium storing a program causing a computer to: (Zhang [page 8]: “MoEfication can significantly reduce the total FLOPS, such as 2x speedup in the ratio of 25%. Meanwhile, the speedup on CPU is close to that on FLOPS. Considering that CPU is widely used for model inference in real-world scenarios, MoEfication is practical for the acceleration of various NLP applications.”)
The remainder of claim 20 is claim 1 in the form of a non-transitory computer readable medium and is rejected for the same reasons as claim 1 stated above.
Claims 10 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang and Nie in view of Riquelme et al. (Scaling Vision with Sparse Mixture of Experts); hereinafter Riquelme
Claim 10 is rejected under Zhang, Nie and Riquelme with the incorporation of claim 1.
Regarding claim 10, Zhang does not appear to explicitly teach wherein the MoE expert is implemented in an image model.
However, Riquelme teaches wherein the MoE expert is implemented in an image model. (Riquelme [page 1, 1 Introduction]: “We introduce the Vision MoE(V-MoE), a sparse variant of the recent Vision Transformer (ViT) architecture [20] for image classification. The V-MoE replaces a subset of the dense feedforward layers in ViT with sparse MoE layers, where each image patch is “routed” to a subset of “experts” (MLPs).”)
It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the image model of Riquelme for smooth performance and reduced training time (Riquelme, page 1, 1 Introduction). Zhang and Riquelme are analogous art because they both concern improving mixture-of-expert models.
Claim 19 is rejected under Zhang, Nie and Riquelme with the incorporation of claim 11.
Regarding claim 19, Zhang teaches wherein the MoE expert is implemented in a large language model and (Zhang [page 4, 4.1 Experimental Setups]: “Models and Hyperparameter. We use four variants of T5 (Raffel et al., 2020), which are the 60-million-parameter T5-Small, the 200-millionparameter T5-Base, the 700-million-parameter T5-Large, and the 3-billion-parameter T5-XLarge.”; Note: T5 (Text-to-Text Transfer Transformer) is a large language model.)
Zhang does not appear to explicitly teach an image model.
However, Riquelme teaches an image model. (Riquelme [page 1, 1 Introduction]: “We introduce the Vision MoE(V-MoE), a sparse variant of the recent Vision Transformer (ViT) architecture [20] for image classification. The V-MoE replaces a subset of the dense feedforward layers in ViT with sparse MoE layers, where each image patch is “routed” to a subset of “experts” (MLPs).”)
It would have been obvious before the effective filing date to combine the creation of a router for expert selection of Zhang with the image model of Riquelme for smooth performance and reduced training time (Riquelme, page 1, 1 Introduction). Zhang and Riquelme are analogous art because they both concern improving mixture-of-expert models.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID H TRAN whose telephone number is (703)756-1525. The examiner can normally be reached M-F 9:30 am - 5:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DAVID H TRAN/Examiner, Art Unit 2147
/ERIC NILSSON/Primary Examiner, Art Unit 2151