CTNF 18/372,661 CTNF 80413 DETAILED ACTION 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. This non-final office action is in response to the application filed 25 September 2023. Claims 1-20 are pending. Claims 1, 15, and 16 are independent claims. Information Disclosure Statement The information disclosure statements (IDS) submitted on 25 September 2023 and 19 October 2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner. Drawings The examiner accepts the drawings filed 25 September 2023. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: According to Step 1 of the two Step analysis, claims 1-14 are directed toward a method (process). Claim 15 is directed toward a computer program product (manufacture). Claims 16-20 are directed toward a system (machine). Therefore, each of these claims falls within one of the four statutory categories. Claim 1 : Step 2A, Prong 1: With respect to claim 1, the claim recites: linearly transforming the parameters of the first transformer using a combination of a width-growth operator and a depth-growth operator, wherein the linear transformation produces a set of new parameters, the set corresponding to the size dimensions of the second transformer (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth operation and a depth-growth operator, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: The judicial exception is not integrated into a practical application. With respect to claim 1, the claim recites the additional elements: accessing parameters of a first transformer receiving size dimensions of a second transformer that is to be trained and is larger than the first transformer These additional elements amount to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. Further, the claim recites the additional element: initializing the second transformer with the set of new parameters This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do no integrate the judicial exception into a practical application. Step 2B: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. With respect to claim 1, the claim recites the additional elements: accessing parameters of a first transformer receiving size dimensions of a second transformer that is to be trained and is larger than the first transformer These additional elements amount to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. Further, the claim recites the additional element: initializing the second transformer with the set of new parameters This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 2 : With respect to claim 2, the claim depends upon claim 1. The analysis of claim 1 is incorporated herein by reference. Step 2A, Prong 1: Claim 2 does not contain any additional limitations analyzed under Step 2A, Prong 1. Step 2A, Prong 2: The judicial exception is not integrated into a practical application. With respect to claim 2, the claim recites the additional elements: training the initialized second transformer with training data to produce a trained second transformer; and performing inferencing via the trained second transformer This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do no integrate the judicial exception into a practical application. Step 2B: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. With respect to claim 2, the claim recites the additional elements: training the initialized second transformer with training data to produce a trained second transformer; and performing inferencing via the trained second transformer This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 3 : With respect to claim 3, the claim depends upon claim 2. The analysis of claim 2 is incorporated herein by reference. Step 2A, Prong 1: Claim 3 does not contain any additional limitations analyzed under Step 2A, Prong 1. Step 2A, Prong 2: The judicial exception is not integrated into a practical application. With respect to claim 3, the claim recites the additional elements: wherein the inferencing comprises performing natural language processing to control a device via a network interface This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do no integrate the judicial exception into a practical application. Step 2B: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. With respect to claim 3, the claim recites the additional elements: wherein the inferencing comprises performing natural language processing to control a device via a network interface This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 4 : With respect to claim 4, the claim depends upon claim 1. The analysis of claim 1 is incorporated herein by reference. Step 2A, Prong 1: Claim 4 recites the elements: wherein the combination of the width-growth operator and the depth-growth operator is a multiplication (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth multiplication operation and a depth-growth multiplication operator, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: Claim 4 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 4 does not contain any additional limitations analyzed under Step 2B. Claim 5 : With respect to claim 5, the claim depends upon claim 1. The analysis of claim 1 is incorporated herein by reference. Step 2A, Prong 1: Claim 5 recites the elements: wherein the width-growth operator comprises a block-diagonal matrix and the depth-growth operator comprises an array of diagonal matrices (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth block-diagonal matrix operation and a depth-growth array of diagonal matrices operation, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: Claim 5 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 5 does not contain any additional limitations analyzed under Step 2B. Claim 6 : With respect to claim 6, the claim depends upon claim 5. The analysis of claim 5 is incorporated herein by reference. Step 2A, Prong 1: Claim 6 recites the elements: wherein both the array of diagonal matrices and the block-diagonal matrix are sparse (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth sparse block-diagonal matrix operation and a depth-growth sparse diagonal matrices operator, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: Claim 6 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 6 does not contain any additional limitations analyzed under Step 2B. Claim 7 : With respect to claim 7, the claim depends upon claim 1. The analysis of claim 1 is incorporated herein by reference. Step 2A, Prong 1: Claim 7 recites the elements: wherein the depth-growth operator linearly combines all layers of the first transformer and, via a factorization, groups the parameters of the first transformer by the layers (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth operator and a depth-growth operator that linearly combines all layers of the first transformer, and via a factorization, groups the parameters of the first transformer by layers, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: Claim 7 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 7 does not contain any additional limitations analyzed under Step 2B. Claim 8 : With respect to claim 8, the claim depends upon claim 1. The analysis of claim 1 is incorporated herein by reference. Step 2A, Prong 1: Claim 8 recites the elements: applying a Kronecker factorization to the width-growth operator and to the depth-growth operator to reduce a respective number of learnable parameters of the width-growth operator and the depth-growth operator (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation of applying a Kronecker factorization to the width-growth and depth-growth parameters to reduce the number of learnable parameters for the width-growth and depth-growth operators) Step 2A, Prong 2: Claim 8 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 8 does not contain any additional limitations analyzed under Step 2B. Claim 9 : With respect to claim 9, the claim depends upon claim 8. The analysis of claim 8 is incorporated herein by reference. Step 2A, Prong 1: Claim 9 recites the elements: wherein, for the depth-growth operator, an entire layer is treated as a single group, a new layer is constructed by combining existing layers, and parameters for all neurons within a single layer are tied in a same layer (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth operation and a depth-growth multiplication operator, wherein for the depth-growth operator, an entire layer is treated as a single group, a new layers is constructed by combining existing layers, and parameters for all neurons within a single layer are tied in the same layer, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: Claim 9 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 9 does not contain any additional limitations analyzed under Step 2B. Claim 10 : With respect to claim 10, the claim depends upon claim 8. The analysis of claim 8 is incorporated herein by reference. Step 2A, Prong 1: Claim 10 recites the elements: wherein, for the width-growth operator, the parameters of the first transformer are grouped by neurons (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth operation, wherein for the width-growth operator, the parameters for the first transformer are grouped by neurons, and a depth-growth operator, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: Claim 10 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 10 does not contain any additional limitations analyzed under Step 2B. Claim 11 : With respect to claim 11, the claim depends upon claim 1. The analysis of claim 1 is incorporated herein by reference. Step 2A, Prong 1: Claim 11 recites the elements: wherein the linearly transforming the parameters of the first transformer comprise linearly transforming parameters of an embedded layer of the first transformer to produce extended parameters for an embedding layer of the second transformer (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation parameters of the first transformer to produce a set of extended parameters for an embedding layer of the second transformer) Step 2A, Prong 2: Claim 11 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 11 does not contain any additional limitations analyzed under Step 2B. Claim 12 : With respect to claim 12, the claim depends upon claim 1. The analysis of claim 1 is incorporated herein by reference. Step 2A, Prong 1: Claim 12 recites the elements: wherein the linearly transforming the parameters of the first transformer comprises using an embedding layer matrix to parameterize at least one of an attention layer and a feedforward layer of the second transformer (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation to parameterize at least one of an attention layer and a feedforward layer of the second transformer by linearly transforming the parameters of the first transformer using an embedding layer matrix) Step 2A, Prong 2: Claim 12 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 12 does not contain any additional limitations analyzed under Step 2B. Claim 13 : With respect to claim 13, the claim depends upon claim 1. The analysis of claim 1 is incorporated herein by reference. Step 2A, Prong 1: Claim 13 recites the elements: wherein the combination of the width-grown operator and the depth-growth operator is learned via steps of stochastic gradient descent (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an observation using stochastic gradient descent to learn the combination of the width-growth operator and the depth-growth operator) Step 2A, Prong 2: Claim 13 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 13 does not contain any additional limitations analyzed under Step 2B. Claim 14 : With respect to claim 14, the claim depends upon claim 1. The analysis of claim 1 is incorporated herein by reference. Step 2A, Prong 1: Claim 14 recites the elements: performing a technique selected from the group consisting of layer dropping, token dropping or staged training (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an judgement and performing one of layer dropping, token dropping, or staged training) Step 2A, Prong 2: Claim 14 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 14 does not contain any additional limitations analyzed under Step 2B. Claim 15 : Step 2A, Prong 1: With respect to claim 15, the claim recites: linearly transforming the parameters of the first transformer using a combination of a width-growth operator and a depth-growth operator, wherein the linear transformation produces a set of new parameters, the set corresponding to the size dimensions of the second transformer (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth operation and a depth-growth operator, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: The judicial exception is not integrated into a practical application. With respect to claim 15, the claim recites the additional element: one or more tangible computer-readable storage media and program instructions stored on at least one of the one or more tangible computer-readable storage medium, the program instructions executable by a processor The claimed element is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)). Additionally, the claim recites the additional elements: accessing parameters of a first transformer receiving size dimensions of a second transformer that is to be trained and is larger than the first transformer These additional elements amount to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. Further, the claim recites the additional element: initializing the second transformer with the set of new parameters This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do no integrate the judicial exception into a practical application. Step 2B: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. With respect to claim 15, the claim recites the additional element: one or more tangible computer-readable storage media and program instructions stored on at least one of the one or more tangible computer-readable storage medium, the program instructions executable by a processor The claimed element is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)). Additionally, the claim recites the additional elements: accessing parameters of a first transformer receiving size dimensions of a second transformer that is to be trained and is larger than the first transformer These additional elements amount to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. Further, the claim recites the additional element: initializing the second transformer with the set of new parameters This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 16 : Step 2A, Prong 1: With respect to claim 16, the claim recites: linearly transforming the parameters of the first transformer using a combination of a width-growth operator and a depth-growth operator, wherein the linear transformation produces a set of new parameters, the set corresponding to the size dimensions of the second transformer (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth operation and a depth-growth operator, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: The judicial exception is not integrated into a practical application. With respect to claim 16, the claim recites the additional element: a memory at least one processor, coupled to said memory, and operative to perform operations The claimed elements are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Additionally, the claim recites the additional elements: accessing parameters of a first transformer receiving size dimensions of a second transformer that is to be trained and is larger than the first transformer These additional elements amount to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. Further, the claim recites the additional element: initializing the second transformer with the set of new parameters This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do no integrate the judicial exception into a practical application. Step 2B: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. With respect to claim 16, the claim recites the additional element: a memory at least one processor, coupled to said memory, and operative to perform operations The claimed elements are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Additionally, the claim recites the additional elements: accessing parameters of a first transformer receiving size dimensions of a second transformer that is to be trained and is larger than the first transformer These additional elements amount to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. Further, the claim recites the additional element: initializing the second transformer with the set of new parameters This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 17 : With respect to claim 17, the claim depends upon claim 16. The analysis of claim 16 is incorporated herein by reference. Step 2A, Prong 1: Claim 17 does not contain any additional limitations analyzed under Step 2A, Prong 1. Step 2A, Prong 2: The judicial exception is not integrated into a practical application. With respect to claim 17, the claim recites the additional elements: training the initialized second transformer with training data to produce a trained second transformer; and performing inferencing via the trained second transformer This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do no integrate the judicial exception into a practical application. Step 2B: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. With respect to claim 17, the claim recites the additional elements: training the initialized second transformer with training data to produce a trained second transformer; and performing inferencing via the trained second transformer This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 18 : With respect to claim 18, the claim depends upon claim 17. The analysis of claim 17 is incorporated herein by reference. Step 2A, Prong 1: Claim 18 does not contain any additional limitations analyzed under Step 2A, Prong 1. Step 2A, Prong 2: The judicial exception is not integrated into a practical application. With respect to claim 18, the claim recites the additional elements: wherein the inferencing comprises performing natural language processing to control a device via a network interface This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do no integrate the judicial exception into a practical application. Step 2B: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. With respect to claim 18, the claim recites the additional elements: wherein the inferencing comprises performing natural language processing to control a device via a network interface This element is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 19 : With respect to claim 19, the claim depends upon claim 16. The analysis of claim 16 is incorporated herein by reference. Step 2A, Prong 1: Claim 19 recites the elements: wherein the depth-growth operator linearly combines all layers of the first transformer and, via a factorization, groups the parameters of the first transformer by the layers (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation on a set of parameters to linearly transform the parameters, based upon a width-growth operator and a depth-growth operator that linearly combines all layers of the first transformer, and via a factorization, groups the parameters of the first transformer by layers, to produce new parameters corresponding to a specified size dimension) Step 2A, Prong 2: Claim 19 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 19 does not contain any additional limitations analyzed under Step 2B. Claim 20 : With respect to claim 20, the claim depends upon claim 16. The analysis of claim 16 is incorporated herein by reference. Step 2A, Prong 1: Claim 20 recites the elements: applying a Kronecker factorization to the width-growth operator and to the depth-growth operator to reduce a respective number of learnable parameters of the width-growth operator and the depth-growth operator (mental process: As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses performing an evaluation of applying a Kronecker factorization to the width-growth and depth-growth parameters to reduce the number of learnable parameters for the width-growth and depth-growth operators) Step 2A, Prong 2: Claim 20 does not contain any additional limitations analyzed under Step 2A, Prong 2. Step 2B: Claim 20 does not contain any additional limitations analyzed under Step 2B. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-20-02-aia AIA This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. 07-21-aia AIA Claim s 1-2, 4, 11, and 14-17 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. ( bert2BERT: Toward Reusable Pretrained Language Models , 14 October 2021, provided by application on IDS filed 19 October 2024, hereafter Chen) and further in view of Liu et al. (US 2020/0372115, published 26 November 2020, hereafter Liu) . As per independent claim 1, Chen discloses a method for machine learning training comprising: accessing parameters of a first transformer (Section 3.1: Here, a transformer-based Pre-trained Language Model (PLM), such as BERT (Section 1). A target model is trained by initializing the target model T with the parameters of the existing model S . The examiner interprets the parameters of the existing model S as the parameters of the first transformer) receiving size dimensions of a second transformer that is to be trained and is larger than the first transformer (Section 3.1: Here, the existing model S has dimensions S(L S ,D S ) and the target model T , which the examiner interprets as the second transformer, has dimensions T(L T ,D T ) , wherein L S ≤ L T and D S ≤ D T ) transform the parameters of the first transforming using a combination of a width-growth operator and a depth-growth operator (Section 3.2: Here, bert2BERT initializes the target model T with parameters of the existing model S by a width-wise expansion and depth-wise expansion), wherein the transformation produces a set of new parameters, the set corresponding to the size dimensions of the second transformer (Section 3.2: Here, the width-wise expansion is decomposed into expansion of the set of matrices or vectors via in-dimension and out-dimension expansion via function preserving and advanced knowledge initialization. Additionally, depth-wise expansion is performed by iteratively stacking the widened model until the depth is equal to the target model (Section 3.3.3)) initializing the second transformer with the set of new parameters (Section 3.3.1: Here, the initialized target model T , receives a set of parameters from the source model and is initialized to function as the source model, meaning that given the same input, the initialized target model has the same output as the source model) Chen fails to specifically disclose a linear transformation. However, Liu, which is analogous to the claimed invention because it is directed toward transforming from a source space to a target space, discloses use of a linear transformation (paragraphs 0037, 0052, and 0057: Here, a transformation from a source space to a target space using a transformation matrix to perform linear transformation (paragraph 0037) to allow for alignment of multiple embeddings for training (paragraph 0057)). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Liu with Chen, with a reasonable expectation of success, as it would have allowed for performing linear transformation to map data between a source and target transformer based upon a transformation matrix (Liu: paragraphs 0037, 0052, and 0057). As per dependent claim 2, Chen and Liu disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Chen further discloses: training the initialized second transformer with training data to produce a trained second transformer (Sections 3.1 and 4.1: Here, parameters from the source model S (first transformer) are transferred to the target model T (second transformer). The second transformer is then trained. This transfer of parameters from the first transformer to the second transformer improves the convergence rate during training) performing inferencing via the trained second transformer (Section 5: Here, the training of the second transformer is applied to natural language processing to perform inferencing) As per dependent claim 4, Chen and Liu disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Chen discloses wherein the combination of the width-growth operator and the depth-growth operator is a multiplication (Section 3.2: Here, the width-wise expansion is decomposed into expansion of the set of matrices or vectors via in-dimension and out-dimension expansion via function preserving and advanced knowledge initialization. Additionally, depth-wise expansion is performed by iteratively stacking the widened model until the depth is equal to the target model (Section 3.3.3). This constitutes expanding a width and depth by a multiple). As per dependent claim 11, Chen and Liu disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Chen discloses wherein transforming the parameters of the first transformer comprises transforming parameters of an embedding layer of the first transformer to produce extended parameters for an embedding layer of the second transformer (Sections 2.1, 3.2, and 3.3.3: Here, an embedding layer is used in the BERT architecture for extending the depth-growth and width-growth parameters). As per dependent claim 14, Chen and Liu disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Chen discloses performing a technique selected from the group consisting of layer dropping (Section 2: Here, layer dropping is a training technique for efficient pre-training), token dropping, and staged training. With respect to independent claim 15, the claim recites the limitations substantially similar to those in claim 1. The analysis of claim 1 is incorporated herein by reference. Further, Liu discloses one or more tangible computer-readable storage media (Figure 2, item 220; paragraph 0011) and program instructions stored on at least one of the one or more tangible computer-readable storage media, the program instructions executable by a process (Figure 2, item 210; paragraph 0011). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Liu with Chen-Liu, with a reasonable expectation of success, as it would have allowed for implementing the method as a computer program product (Liu: paragraph 0011). With respect to independent claim 16, the claim recites the limitations substantially similar to those in claim 1. The analysis of claim 1 is incorporated herein by reference. Further, Liu discloses a memory (Figure 2, item 220; paragraph 0011) and at least one processor, coupled to the memory, and operative to perform operations (Figure 2, item 210; paragraph 0011). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Liu with Chen-Liu, with a reasonable expectation of success, as it would have allowed for implementing the method as a system (Liu: paragraph 0011). With respect to dependent claim 17, the claim recites the limitations substantially similar to those in claim 2. Claim 17 is rejected under similar rationale . 07-21-aia AIA Claim s 3 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen and Liu and further in view of Barut et al. (US 12045288, filed 24 September 2020, hereafter Barut) . As per dependent claim 3, Chen and Liu disclose the limitations similar to those in claim 2, and the same rejection is incorporated herein. Chen discloses wherein the inferencing comprises performing natural language processing (Section 5: Here, the training of the second transformer is applied to natural language processing to perform inferencing). Chen fails to specifically disclose performing natural language processing to control a device via a network interface. However, Barut, which is analogous to the claimed invention because it is directed toward performing natural language processing to control a device, discloses performing natural language processing to control a device via a network interface (column 31, lines 5-28: Here, a natural language processing system receives spoken or textual natural language input via a network interface to control a device to take one or more actions based upon the NLP input). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Barut with Chen-Liu, with a reasonable expectation of success, as it would have allowed for controlling a device based upon user input (Barut: column 31, lines 5-28). With respect to dependent claim 18, the claim recites the limitations substantially similar to those in claim 3. Claim 18 is rejected under similar rationale . 07-21-aia AIA Claim s 5-6 are rejected under 35 U.S.C. 103 as being unpatentable over Chen and Liu and further in view of Fok et al. (US 2022/0012575, published 13 January 2022, hereafter Fok) . As per dependent claim 5, Chen and Liu disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Chen fails to specifically disclose wherein the width-growth operator comprises a block-diagonal matrix and the depth-growth operator comprises an array of diagonal matrices. However, Fok, which is analogous to the claimed invention because it is directed toward scaling neural networks, discloses using a block diagonal matrix and the diagonal matrices (paragraphs 0068-0069 and 0075: Here, a neural network is scaled using a block diagonal matrix and diagonal matrices). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Fok with Chen-Liu, with a reasonable expectation of success, as it would have allowed for scaling a neural network using sparse diagonal matrices (Fok: paragraph 0068). As per dependent claim 6, Chen, Liu, and Fok disclose the limitations similar to those in claim 5, and the same rejection is incorporated herein. Fok discloses wherein the diagonal matrices and block-diagonal matrix are sparse (paragraph 0069). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Fok with Chen-Liu, with a reasonable expectation of success, as it would have allowed for saving memory space (Fok: paragraph 0055) . 07-21-aia AIA Claim s 7-10 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Chen and Liu and further in view of Grosse et al. ( A Kronecker-factored approximate Fisher matrix for convolution layers , 2016, hereafter Grosse) . As per dependent claim 7, Chen and Liu disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Chen fails to specifically disclose wherein the depth-growth operator linearly combines all layers of the first transformer and, via a factorization, groups the parameters of the first transformer by the layers. However, Grosse, which is analogous to the claimed invention because it is directed toward factorization, discloses linearly combines all layers of the first transformer and, via a factorization, groups the parameters of the first transformer by the layers (Section 3: Here, each block of the matrix combines all the parameters relevant to a layer into a single block. These blocks are then processed by the Kronecker factorization). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Grosse with Chen-Liu, with a reasonable expectation of success, as it would have allowed for approximating layers to improve approximations for training a machine learning model in order to improve speed (Grosse: Abstract). As per dependent claim 8, Chen and Liu disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. While Chen discloses a width-growth operator and a depth-growth operator (Sections 3.2 and 3.3.3), but Chen fails to specifically disclose applying a Kronecker factorization to the width-growth operation and to the depth-growth operator to reduce a respective number of learnable parameters of the width-growth operator and the depth-growth operator. However, Grosse, which is analogous to the claimed invention because it is directed toward Kronecker factorization, discloses applying a Kronecker factorization to the operators to reduce a respective number of learnable parameters (Section 3: Here, a Kronecker-factors approximate curvature is calculated for each block, with each block representing a layer. This reduces the number of learnable parameters). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Grosse with Chen-Liu, with a reasonable expectation of success, as it would have allowed for approximating layers to improve approximations for training a machine learning model in order to improve speed (Grosse: Abstract). As per dependent claim 9, Chen, Liu, and Grosse disclose the limitations similar to those in claim 8, and the same rejection is incorporated herein. Chen discloses a depth-growth operator (Section 3.3.3). However, Chen fails to disclose wherein, an entire layer is treated as a single group, a new layer is constructed by combining existing layers, and parameters for all neurons within a single layer are tied in a same layer. However, Grosse, which is analogous to the claimed invention because it is directed toward Kronecker factorization, discloses wherein, an entire layer is treated as a single group (Section 3: Here, each layer is represented as a single block), a new layer is constructed by combining existing layers (Section 3: Here, a factorization of each block is generated as a block), and parameters for all neurons within a single layer are tied in a same layer (Section 3: Here, the neural network, comprised of neurons, shares parameters for all neurons within a single level (Section 2)). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Grosse with Chen-Liu, with a reasonable expectation of success, as it would have allowed for approximating layers to improve approximations for training a machine learning model in order to improve speed (Grosse: Abstract). As per dependent claim 10, Chen, Liu, and Grosse disclose the limitations similar to those in claim 9, and the same rejection is incorporated herein. Chen discloses a width-growth operator (Section 3.2), but fails to specifically disclose wherein, the parameters of the first transformer are grouped by neurons. However, Gross, which is analogous to the claimed invention because it is directed toward grouping neurons of a neural network into a vector, discloses wherein, the parameters of the first transformer are grouped by neurons (Section 2: Here, all parameters of the neural network, and therefore, its neurons, are grouped into a vector). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Grosse with Chen-Liu, with a reasonable expectation of success, as it would have allowed for approximating layers to improve approximations for training a machine learning model in order to improve speed (Grosse: Abstract). With respect to dependent claims 19-20, the claims recites the limitations substantially similar to those in claims 7-8, respectively. Claims 19-20 are rejected under similar rationale . 07-21-aia AIA Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Chen and Liu and further in view of Zhou et al. (US 2025/0190756, provisional filed 25 May 2022, hereafter Zhou) and further in view of Xu (US 2025/0021800, filed 14 July 2023) . As per dependent claim 12, Chen and Liu disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Chen fails to discloses transforming the parameters of the first transformer using an embedding layer matrix to parameterize at least one of an attention layer and a feedforward layer of the second transformer. However, Zhou, which is analogous to the claimed invention because it is directed toward encoding features, discloses transforming the parameters using an embedding layer matrix to parameterize values (paragraph 0043: Here, a position embedding matrix is used to generate an output vector). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Zhou with Chen-Liu, with a reasonable expectation of success, as it would have allowed for generating model vectors based upon an embedding layer matrix (Zhou: paragraph 0043). Additionally, Xu, which is analogous to the claimed invention because it is directed toward efficient processing of neural networks, discloses generating an output matrix based upon an embedded layer (paragraph 0065) and providing the embedded sequence to a feed-forward layer of a neural network (paragraph 0082). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Xu with Chen-Liu-Zhou, with a reasonable expectation of success, as it would have allowed for normalization of layers (Xu: paragraph 0082) . 07-21-aia AIA Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Chen and Liu and further in view of Chen et al. (US 2020/0160836, published 21 May 2020, hereafter Li) . As per dependent claim 13, Chen and Liu disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Chen discloses a width-growth operator and a depth-growth operator (Sections 3.2 and 3.3.3). Chen fails to specifically disclose stochastic gradient descent. However, Li, which is analogous to the claimed invention because it is directed toward using parameters to train models, discloses the use of stochastic gradient descent (paragraph 0067). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Li with Chen-Liu, with a reasonable expectation of success, as it would have allowed for optimizing the training of the model (Li: paragraph 0067) . Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure : Auld et al. (US 2024/0311094): Discloses reuse of blocks across large language models (paragraph 0058) Kochman et al. (US 2024/0283820): Discloses reuse of engineered features across large language models (paragraph 0035) Sanz et al. (US 12008026): Discloses reuse of previously trained large language models (column 12, line 55- column 13, line 19) Any inquiry concerning this communication or earlier communications from the examiner should be directed to KYLE R STORK whose telephone number is (571)272-4130. The examiner can normally be reached 8am - 2pm; 4pm - 6pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at 571/272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KYLE R STORK/Primary Examiner, Art Unit 2128 Application/Control Number: 18/372,661 Page 2 Art Unit: 2128 Application/Control Number: 18/372,661 Page 3 Art Unit: 2128 Application/Control Number: 18/372,661 Page 4 Art Unit: 2128 Application/Control Number: 18/372,661 Page 5 Art Unit: 2128 Application/Control Number: 18/372,661 Page 6 Art Unit: 2128 Application/Control Number: 18/372,661 Page 7 Art Unit: 2128 Application/Control Number: 18/372,661 Page 8 Art Unit: 2128 Application/Control Number: 18/372,661 Page 9 Art Unit: 2128 Application/Control Number: 18/372,661 Page 10 Art Unit: 2128 Application/Control Number: 18/372,661 Page 11 Art Unit: 2128 Application/Control Number: 18/372,661 Page 12 Art Unit: 2128 Application/Control Number: 18/372,661 Page 13 Art Unit: 2128 Application/Control Number: 18/372,661 Page 14 Art Unit: 2128 Application/Control Number: 18/372,661 Page 15 Art Unit: 2128 Application/Control Number: 18/372,661 Page 16 Art Unit: 2128 Application/Control Number: 18/372,661 Page 17 Art Unit: 2128 Application/Control Number: 18/372,661 Page 18 Art Unit: 2128 Application/Control Number: 18/372,661 Page 19 Art Unit: 2128 Application/Control Number: 18/372,661 Page 20 Art Unit: 2128 Application/Control Number: 18/372,661 Page 21 Art Unit: 2128 Application/Control Number: 18/372,661 Page 22 Art Unit: 2128 Application/Control Number: 18/372,661 Page 23 Art Unit: 2128 Application/Control Number: 18/372,661 Page 24 Art Unit: 2128 Application/Control Number: 18/372,661 Page 25 Art Unit: 2128 Application/Control Number: 18/372,661 Page 26 Art Unit: 2128 Application/Control Number: 18/372,661 Page 27 Art Unit: 2128 Application/Control Number: 18/372,661 Page 28 Art Unit: 2128 Application/Control Number: 18/372,661 Page 29 Art Unit: 2128 Application/Control Number: 18/372,661 Page 30 Art Unit: 2128 Application/Control Number: 18/372,661 Page 31 Art Unit: 2128 Application/Control Number: 18/372,661 Page 32 Art Unit: 2128 Application/Control Number: 18/372,661 Page 33 Art Unit: 2128 Application/Control Number: 18/372,661 Page 34 Art Unit: 2128 Application/Control Number: 18/372,661 Page 35 Art Unit: 2128 Application/Control Number: 18/372,661 Page 36 Art Unit: 2128 Application/Control Number: 18/372,661 Page 37 Art Unit: 2128 Application/Control Number: 18/372,661 Page 38 Art Unit: 2128 Application/Control Number: 18/372,661 Page 39 Art Unit: 2128 Application/Control Number: 18/372,661 Page 40 Art Unit: 2128 Application/Control Number: 18/372,661 Page 41 Art Unit: 2128 Application/Control Number: 18/372,661 Page 42 Art Unit: 2128