DETAILED ACTION
This Office Action is sent in response to the Applicant’s Communication received on 03/23/2026 for application number 17/775,357. The Office hereby acknowledges receipt of the following and placed of record in file: Specification, Drawings, Abstract, Oath/Declaration, IDS, and Claims.
Claims 1, 3, 6 and 7 are currently amended.
Claim 2 is canceled.
Claims 1 and 3-7 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 03/26/2026 has been entered.
Response to Arguments
35 USC 112
On page 6 of the remarks section, the Applicant asserts Applicant amended claim 3 and respectfully requests withdrawal of the objection.
The Examiner respectfully disagrees. Only one of the two instances of the limitation lacking antecedent basis was appropriately amended. Therefore, the 35 USC 112 rejection is maintained.
35 USC 101
On page 6 of the remarks section, Applicant argues that without conceding to the appropriateness of the rejection, claim 1 is amended to recite, in part, "perform[ing] learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer." Applicant respectfully submits that the claimed features cannot be practically performed in the human mind, even with a paper and pencil. For example, a human mind cannot practically perform learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer. Rather, claim 1 is directed to a specific technical improvement (e.g., efficient implementation of a reservoir-based machine learning model that maintains high accuracy with a reduced physical footprint). Therefore, Applicant respectfully submits that amended claim 1 is not directed to mental processes.
The Examiner respectfully disagrees. The claimed limitation “perform[ing] learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer” was analyzed under Step 2A Prong One as being an abstract idea. The action of performing learning of a weight is a limitation recited at a high-level generality and, under the Broadest Reasonable Interpretation (BRI), can be interpreted as a procedure of human observation, evaluation, judgement, or opinion. The claim does not provide any explicit steps or details of how the recited “learning” is being implemented that would meaningfully limit the claimed limitation to falling outside the grouping of abstract ideas. The claimed limitation can be performed mentally with the aid of pen and paper, and is therefore a mental process.
The Applicant further argues that the claim as a whole integrates the alleged exception into a practical application, as evidenced by the Specification. For example, the Specification describes the technical problem of lower calculation speed and
inefficient power consumption during execution of a large model (e.g., a reservoir-based
machine learning model) to achieve the same performance as other models where parameters of layers other than the output layer are learned. Specification at paragraphs [0002]-[0004] and [0009]-[0010]. Claim 1, as a whole and when read in light of the Specification, solves this technical problem by at least performing weighting on each of the plurality results for each of the transitioned internal states, outputting the output data based on a sum of a plurality of weighted results, and performing learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer. As a result, outputs from the intermediate layer at past time steps are used, and the number of output concatenations can be relatively increased without the need to increase the nodes in the intermediate layer. Id., at paragraphs [0082]-[0083]. Additionally, even when the number of dimensions of the intermediate layer is reduced, the number of output concatenations can be made constant by adding concatenations from past times. Id. These provide high learning performance for the reservoir-based model without the need to increase the model size, while maintaining learning performance even when reducing the model size. Id., at paragraph [0015].
The Examiner respectfully disagrees. Alleged disclosed invention as improvements to the computers of Al frameworks are not explicitly reflected in the claimed invention. Although, instant application paragraphs 0002-0004, 0009-0010, and 0082-0083 have detailed descriptions of implementing steps that provide technological improvement, these details are not explicitly present in the current claim. Specifically, data representing a certain amount of temporal change, a neural network, reducing the size of the model, auxiliaries, and concatenations from past times, are not explicitly stated in the claim. Therefore, the claimed elements or combination of elements do not reflect the steps that would to a technical improvement as alleged.
Therefore, the 35 USC 101 rejection is maintained.
35 USC 103
On page 11 of the remarks section, the Applicant argues that Werbos only teaches adjusting "a set of weights" so as to make the actual outputs Y'(t) to approximate the desired outputs Y(t). Id. However, this cannot be relied upon to allegedly teach at least the claimed "perform[ing] weighting on each of the plurality results for each of the transitioned internal states" and "perform[ing] learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer." Werbos does not describe generating the plurality results for each of the transitioned internal states from the same intermediate layer, and weighing each of the plurality
results for each of the transitioned internal states. Instead, Werbos only describe a "simple feedforward network" after describing the "basic propagation." Therefore, Werbos does not teach such subject matter.
The Examiner respectfully disagrees. Werbos alone was not cited to teach the limitation “perform[ing] weighting on each of the plurality results for each of the transitioned internal states". Rather, the combination Werbos-Lathrop does indeed teach the limitation. Werbos teaches “perform weighting on each of the plurality of results” in section 2: “We must specify the ‘topology’ (connections and equations) for a network which inputs X(t) and outputs a four-component vector Y ̂(t), an approximation to Y(t). The relation between the inputs and outputs must depend on a set of weights (parameters) W which can be adjusted. We must specify a "learning rule"-a procedure for adjusting (weighting) the weights W so as to make the actual outputs P(t) (each of the plurality results) approximate the desired outputs Y(t)”. Lathrop teaches “results for each of the transitioned internal states” in paragraph 83: “In general, reset line I 100 can provide any reset signal pattern that causes all nodes 121 connected to reset lines to be reconfigured to a desired state (transitioned internal states). For example, the reset line I 100 may be a single digital channel with value r = {0,1} that connects to inputs of all the nodes 121 in the internal network 120. For example, when r = 0, outputs from nodes 121 can be a function of their other inputs and when r = 1, node outputs from nodes 121 can all be set equal to I (results) and the network state can be reset to an all " I " state”. In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
Moreover, Werbos does indeed teach the limitation "perform[ing] learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer". On page 1551, Werbos teaches: “We must specify a "learning rule"-a procedure for adjusting the weights W (perform[ing] learning of a weight obtained by the weighting) so as to make the actual outputs P(t) approximate the desired outputs Y(t)”. On page 1552, Werbos teaches: “we calculate the derivatives of E with respect to all of the weights; this is indicated by the dotted lines in Fig. 3. If increasing a given weight would lead to more error, we adjust that weight downwards (not increasing a size)”. The intermediate layer is found in figure 5 of Werbos.
The Applicant further argues that Lanthrop cannot be relied upon to allegedly teach at least the claimed "perform[ing] weighting on each of the plurality results for each of the transitioned internal states" and "perform[ing] learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer." Lanthrop does not describe that each of the plurality of results for each of the transitioned internal states is performed within the internal network 120. Although FIG. 7 of Lanthrop illustrate a structure of the internal neural network 120, Lanthrop does not teach storing a plurality of states of nodes included in the internal network 120 to multiply each of the plurality of states by the weight coefficient, and outputting the result of the calculation. Also, Lanthrop does not teach perform learning of a weight obtained by the weighting,
thereby not increasing a size of the internal network 120. Therefore, Lanthrop does not teach such subject matter.
The Examiner respectfully disagrees. For the reasons stated above, Lanthrop alone was not cited to teach the cited limitations. Rather, the combination Werbos-Lathrop were cited to teach the cited limitations. Werbos teaches “perform weighting on each of the plurality of results” in section 2: “We must specify the ‘topology’ (connections and equations) for a network which inputs X(t) and outputs a four-component vector Y ̂(t), an approximation to Y(t). The relation between the inputs and outputs must depend on a set of weights (parameters) W which can be adjusted. We must specify a "learning rule"-a procedure for adjusting (weighting) the weights W so as to make the actual outputs P(t) (each of the plurality results) approximate the desired outputs Y(t)”. Lathrop teaches “results for each of the transitioned internal states” in paragraph 83: “In general, reset line I 100 can provide any reset signal pattern that causes all nodes 121 connected to reset lines to be reconfigured to a desired state (transitioned internal states). For example, the reset line I 100 may be a single digital channel with value r = {0,1} that connects to inputs of all the nodes 121 in the internal network 120. For example, when r = 0, outputs from nodes 121 can be a function of their other inputs and when r = 1, node outputs from nodes 121 can all be set equal to I (results) and the network state can be reset to an all " I " state”. Moreover, Werbos teaches the limitation "perform[ing] learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer". On page 1551, Werbos teaches: “We must specify a "learning rule"-a procedure for adjusting the weights W (perform[ing] learning of a weight obtained by the weighting) so as to make the actual outputs P(t) approximate the desired outputs Y(t)”. On page 1552, Werbos teaches: “we calculate the derivatives of E with respect to all of the weights; this is indicated by the dotted lines in Fig. 3. If increasing a given weight would lead to more error, we adjust that weight downwards (not increasing a size)”. The intermediate layer is found in figure 5 of Werbos. In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Additionally, although “storing a plurality of states of nodes included in the internal network… to multiply each of the plurality of states by the weight coefficient, and outputting the result of the calculation” appear to be disclosed invention, the cited steps are not explicitly recited in the amended claim language. In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993).
The Applicant further argues that Soures cannot be relied upon to allegedly teach at least the claimed "perform[ing] weighting on each of the plurality results for each of the transitioned internal states" and "perform[ing] learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer." Therefore, Applicant respectfully submits that the cited references, either individually or in combination, do not teach or suggest claim 1.
The Examiner respectfully disagrees. Soures was not cited to teach the limitations "perform[ing] weighting on each of the plurality results for each of the transitioned internal states" and "perform[ing] learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer." Rather, the combination Werbos-Lanthrop were cited to teach the aforementioned limitations.
Therefore, the 35 USC 103 rejection is maintained.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 3 recites the limitation "the interface" in line 3. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1 and 3-7 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claims 1 and 3-5 are directed towards a machine learning device. Claim 6 is directed towards an information processing method. Claim 7 is directed towards a non-transitory recording medium. Therefore, all claims are directed towards one of the four statutory categories of patent eligible subject matter.
Claim 1
Step 2A Prong 1:
Claim 1 recites:
“Perform learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer;” Performing learning of a weight obtained by the weighting is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“perform weighting on each of the plurality of results for each of the transitioned internal states;” performing weighting on each of the plurality results is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“a machine learning device; a processor configured to execute instructions;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“an input layer configured to acquire input data;” “acquire the input data;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“an intermediate layer configured to transition an internal state associated with the input data from the input layer;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“an output layer configured to output data based on values acquired by multiplying output of the intermediate layer and weight coefficients;” “output the output data based on a sum of a plurality of weighted results;” Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)).
“transition the internal state associated with the input data a plurality of times to obtain a plurality of results for each transitioned internal state;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“a machine learning device; a processor configured to execute instructions;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“an input layer configured to acquire input data;” “acquire the input data;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity. See MPEP 2106.05(g). The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
“an intermediate layer configured to transition an internal state associated with the input data from the input layer;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“an output layer configured to output data based on values acquired by multiplying output of the intermediate layer and weight coefficients;” “output the output data based on a sum of a plurality of weighted results;” Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)). This falls under Well-Understood, Routine, Conventional activity -see MPEP 2106.05(d)(II)(vi).
“transition the internal state associated with the input data a plurality of times to obtain a plurality of results for each transitioned internal state;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 3
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“wherein the processor is configured to execute the instructions to transition internal states a plurality of times during a period of time from a moment where the interface acquires the input data to a moment where the input layer acquires next input data;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“wherein the processor is configured to execute the instructions to transition internal states a plurality of times during a period of time from a moment where the interface acquires the input data to a moment where the input layer acquires next input data;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 4
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“wherein the processor is configured to execute the instructions to: transition the internal states the plurality of times during a period of time from a first moment where the input layer acquires the input data to a second moment where the input layer acquires next input data; and upon the input layer acquiring the next input data, start transitioning the next input data from a state prior to transitioning at least some of the plurality of times;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“wherein the processor is configured to execute the instructions to: transition the internal states the plurality of times during a period of time from a first moment where the input layer acquires the input data to a second moment where the input layer acquires next input data; and upon the input layer acquiring the next input data, start transitioning the next input data from a state prior to transitioning at least some of the plurality of times;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 5
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“a memory configured to store the result of the weighting;” This limitation is merely a post-solution step of storing the data—a nominal addition to the claim that does not meaningfully limit the claim. The method storing is recited at a high level of generality. Simply implementing the abstract idea in a generic method is not a practical application of the abstract idea. Therefore, storing step is an insignificant extra-solution activity. See MPEP 2106.05(g).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“a memory configured to store the result of the weighting;” These elements amount to storing… information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93; See MPEP 2106.05(d) (II)(iv). The courts have recognized the computer functions of storing as well‐understood, routine, and conventional function when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 6
Step 2A Prong 1:
Claim 6 recites:
“performing weighting on each of the plurality of results for each of the transitioned internal states;” Performing weighting on each of the plurality of results is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“performing learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer;” Performing learning of a weight obtained by the weighting is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“an information processing method by a machine learning device;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“an input layer configured to acquire input data;” “acquiring input data;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“an intermediate layer configured to transition an internal state associated with the input data from the input layer;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“an output layer configured to output data based on values acquired by multiplying output of the intermediate layer and weight coefficients;” “outputting output data based on a sum of a plurality of weighted results;” Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)).
“transition the internal state associated with the input data a plurality of times to obtain a plurality of results for each transitioned internal state;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“an information processing method by a machine learning device;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“an input layer configured to acquire input data;” “acquiring input data;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“an intermediate layer configured to transition an internal state associated with the input data from the input layer;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“an output layer configured to output data based on values acquired by multiplying output of the intermediate layer and weight coefficients;” “outputting output data based on a sum of a plurality of weighted results;” Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)).
“transition the internal state associated with the input data a plurality of times to obtain a plurality of results for each transitioned internal state;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 7
Step 2A Prong 1:
Claim 7 recites:
“performing weighting on each of the plurality of results for each of the transitioned internal states;” Performing weighting on each of the plurality of results is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“performing learning of a weight obtained by the weighting, thereby not increasing a size of the intermediate layer;” Performing learning of a weight obtained by the weighting is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“a non-transitory recording medium that stores a program having an information processing method executed by a computer;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“an input layer configured to acquire input data;” “acquire the input data;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“an intermediate layer configured to transition an internal state associated with the input data from the input layer;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“an output layer configured to output data based on values acquired by multiplying output of the intermediate layer and weight coefficients;” “output the output data based on a sum of a plurality of weighted results;” Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)).
“transition the internal state associated with the input data a plurality of times to obtain a plurality of results for each transitioned internal state;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“a non-transitory recording medium that stores a program having an information processing method executed by a computer;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“an input layer configured to acquire input data;” “acquire the input data;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity. See MPEP 2106.05(g). The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
“an intermediate layer configured to transition an internal state associated with the input data from the input layer;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“an output layer configured to output data based on values acquired by multiplying output of the intermediate layer and weight coefficients;” “output the output data based on a sum of a plurality of weighted results;” Insignificant extra-solution as the limitation amounts to necessary data outputting (MPEP 2106.05(g)(3)). This falls under Well-Understood, Routine, Conventional activity -see MPEP 2106.05(d)(II)(vi).
“transition the internal state associated with the input data a plurality of times to obtain a plurality of results for each transitioned internal state;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1 and 5-7 are rejected under 35 U.S.C. 103 as being unpatentable over Werbos (Backpropagation Through Time: What It Does and How to Do it, published 1990), hereinafter Werbos, in view of Lathrop et al. (WO 2018213399 A1), hereinafter Lathrop.
Regarding claim 1, Werbos teaches,
A machine learning device [Sect I, col 1, para 2, This paper will mainly describe a simpler version of backpropagation, which can be translated into computer code and applied directly by neural network users] comprising:
an input layer configured to acquire input data [Fig. 5, Input layer];
an intermediate layer configured to transition (Sect 2, pg. 1550, para 2, adapt the parameters of the network) associated with the input data from the input layer [Fig. 5, Hidden layer; Sect 2, pg. 1550, para 2, The goal is to adapt the parameters of the network so that it performs well for patterns from outside the training set; Sect 2, pg. 1551, col 1, para 3, We must specify the "topology" (connections and equations) for a network which inputs X(t) and outputs a four-component vector Y ̂(t)];
an output layer [Fig. 5, Output layer] configured to output data based on values acquired by multiplying output of the intermediate layer and weight coefficients [Sect 2, pg. 1551, col 1, para 3, The relation between the inputs and outputs must depend on a set of weights (parameters) W which can be adjusted; Sect 2, pg. 1552, col 1, Subroutine pseudocode, net = net + W(i, j)*x(j); Sect 2, pg. 1552, col 1, In the pseudocode, note that X and Ware technically the inputs to the subroutine, while x and Yhat are the outputs. Yhat is usually regarded as "the" output of the network];
PNG
media_image1.png
257
770
media_image1.png
Greyscale
and a processor [Sect I, col 2, para 3, control system] configured to execute instructions [Sect I, col 1, para 2, computer code] to: acquire the input data [Sect 2, pg. 1551, col 1, para 1, We assume that we already have a camera and preprocessor which can digitize the image, locate the five digits, and provide a 19 x 20 grid of ones and zeros representing the image of each digit. We want the neural network to input the 19 x 20 image, and output a classification];
transition associated with the input data a plurality of times to obtain a plurality of results for each transition [Sect 2, pg. 1551, col 1, para 2, Before adapting the parameters of the neural network, one must first obtain a training database of actual handwritten digits and correct classifications. Suppose, for example, that this database contains 2000 examples of handwritten digits. In that case, T = 2000. We may give each example a label t between 1 and 2000. For each sample t, we have a record of the input pattern and the correct classification. Each input pattern consists of 380 numbers, which may be viewed as a vector with 380 components; we may call this vector X(t)];
perform weighting on each of the plurality of results; output the output data based on a sum of a plurality of weighted results (Sect IV, pg. 1558, col 1, para 3, running sum) [Sect 2, pg. 1551, col 1, para 3 & 4, We must specify the "topology" (connections and equations) for a network which inputs X(t) and outputs a four-component vector Y ̂(t), an approximation to Y(t). The relation between the inputs and outputs must depend on a set of weights (parameters) W which can be adjusted. We must specify a "learning rule"-a procedure for adjusting the weights W so as to make the actual outputs P(t) approximate the desired outputs Y(t); Sect IV, pg. 1558, col 1, para 3, The subroutine F-ACTION is virtually identical to the old subroutine F-NET, except that we need to calculate F-W as a running sum (as we did in F-NET2)]; and
perform learning of a weight obtained by the weighting [Sect 2, pg. 1551, col 1, para 4, We must specify a "learning rule"-a procedure for adjusting the weights W so as to make the actual outputs P(t) approximate the desired outputs Y(t)], thereby not increasing a size of the intermediate layer [Sect II, pg. 1552, col 2, para 1, we calculate the derivatives of E with respect to all of the weights; this is indicated by the dotted lines in Fig. 3. If increasing a given weight would lead to more error, we adjust that weight downwards].
Werbos does not teach layer configured to transition an internal state associated with the input; transition the internal state associated with the input data a plurality of times to obtain results; results for each of the transitioned internal states.
Lathrop teaches,
layer configured to transition an internal state associated with the input; transition the internal state associated with the input data a plurality of times to obtain results [Para 121, in each clock cycle, each value u of an input signal provided by input layer 110 may represent new source data, such as a new image. Once a new image has been loaded to the input layer 110, the internal network 120 can begin to asynchronously process the image input signal, and the node states in the internal network 120 may continuously change… The delay between the input layer clock 1710 and output layer clock 1720 can be a parameter of device 100 that can be tuned to optimize the classification accuracy for a given task];
results for each of the transitioned internal states [Para 83, In general, reset line I 100 can provide any reset signal pattern that causes all nodes 121 connected to reset lines to be reconfigured to a desired state. For example, the reset line I 100 may be a single digital channel with value r = {0,1} that connects to inputs of all the nodes 121 in the internal network 120. For example, when r = 0, outputs from nodes 121 can be a function of their other inputs and when r = 1, node outputs from nodes 121 can all be set equal to I and the network state can be reset to an all " I " state].
Lathrop is analogous to the claimed invention as they both relate to reservoir computing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Werbos’ teachings to incorporate the teachings of Lathrop and provide transitioning an internal state in order to [Lathrop, para 121] give higher classification accuracy.
Regarding claim 5, Werbos-Lathrop teach the limitations of claim 1.
Werbos further teaches,
A memory (Sect I, col 2, para 3, memory) configured to store (Sect II, pg. 1559, col 1, para 2, It requires intermediate storage) the result of weighting [Sect I, col 2, para 3, This allows one to calculate the derivatives needed when optimizing an iterative analysis procedure, a neural network with memory, or a control system which maximizes performance over time; Sect II, pg. 1559, col 1, para 2, For each string, we can calculate complete derivatives and update the weights. Then we can go on to the next string. This is like pattern learning, in that the weights are updated incrementally before the entire data set is studied. It requires intermediate storage for only one string at a time.].
Regarding claim 6, Werbos teaches,
An information processing method by a machine learning device [Sect I, col 1, para 2, This paper will mainly describe a simpler version of backpropagation, which can be translated into computer code and applied directly by neural network users] comprising:
an input layer configured to acquire input data [Fig. 5, Input layer];
an intermediate layer configured to transition (adapt the parameters of the network) associated with the input data from the input layer [Fig. 5, Hidden layer; Sect 2, pg. 1550, para 2, The goal is to adapt the parameters of the network so that it performs well for patterns from outside the training set; Sect 2, pg. 1551, col 1, para 3, We must specify the "topology" (connections and equations) for a network which inputs X(t) and outputs a four-component vector Y ̂(t)];
an output layer [Fig. 5, Output layer] configured to output data based on values acquired by multiplying output of the intermediate layer and weight coefficients [Sect 2, pg. 1551, col 1, para 3, The relation between the inputs and outputs must depend on a set of weights (parameters) W which can be adjusted; Sect 2, pg. 1552, col 1, Subroutine pseudocode, net = net + W(i, j)*x(j); Sect 2, pg. 1552, col 1, In the pseudocode, note that X and Ware technically the inputs to the subroutine, while x and Yhat are the outputs. Yhat is usually regarded as "the" output of the network];
acquiring input data [Fig. 5, Input Layer];
PNG
media_image1.png
257
770
media_image1.png
Greyscale
transition associated with the input data a plurality of times to obtain a plurality of results for each transition [Sect 2, pg. 1551, col 1, para 2, Before adapting the parameters of the neural network, one must first obtain a training database of actual handwritten digits and correct classifications. Suppose, for example, that this database contains 2000 examples of handwritten digits. In that case, T = 2000. We may give each example a label t between 1 and 2000. For each sample t, we have a record of the input pattern and the correct classification. Each input pattern consists of 380 numbers, which may be viewed as a vector with 380 components; we may call this vector X(t)];
performing weighting on each of the plurality of results; outputting the output data based on a sum of a plurality of weighted results (Sect IV, pg. 1558, col 1, para 3, running sum) [Sect 2, pg. 1551, col 1, para 3 & 4, We must specify the "topology" (connections and equations) for a network which inputs X(t) and outputs a four-component vector Y ̂(t), an approximation to Y(t). The relation between the inputs and outputs must depend on a set of weights (parameters) W which can be adjusted. We must specify a "learning rule"-a procedure for adjusting the weights W so as to make the actual outputs P(t) approximate the desired outputs Y(t); Sect IV, pg. 1558, col 1, para 3, The subroutine F-ACTION is virtually identical to the old subroutine F-NET, except that we need to calculate F-W as a running sum (as we did in F-NET2)]; and
performing learning of a weight obtained by the weighting [Sect 2, pg. 1551, col 1, para 4, We must specify a "learning rule"-a procedure for adjusting the weights W so as to make the actual outputs P(t) approximate the desired outputs Y(t)], thereby not increasing a size of the intermediate layer [Sect II, pg. 1552, col 2, para 1, we calculate the derivatives of E with respect to all of the weights; this is indicated by the dotted lines in Fig. 3. If increasing a given weight would lead to more error, we adjust that weight downwards].
Werbos does not teach layer configured to transition an internal state associated with the input; transition the internal state associated with the input data a plurality of times to obtain results; results for each of the transitioned internal states.
Lathrop teaches,
layer configured to transition an internal state associated with the input; transition the internal state associated with the input data a plurality of times to obtain results [Para 121, in each clock cycle, each value u of an input signal provided by input layer 110 may represent new source data, such as a new image. Once a new image has been loaded to the input layer 110, the internal network 120 can begin to asynchronously process the image input signal, and the node states in the internal network 120 may continuously change… The delay between the input layer clock 1710 and output layer clock 1720 can be a parameter of device 1 00 that can be tuned to optimize the classification accuracy for a given task];
results for each of the transitioned internal states [Para 83, In general, reset line I 100 can provide any reset signal pattern that causes all nodes 121 connected to reset lines to be reconfigured to a desired state. For example, the reset line I 100 may be a single digital channel with value r = {0,1} that connects to inputs of all the nodes 121 in the internal network 120. For example, when r = 0, outputs from nodes 121 can be a function of their other inputs and when r = 1, node outputs from nodes 121 can all be set equal to I and the network state can be reset to an all " I " state].
Lathrop is analogous to the claimed invention as they both relate to reservoir computing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Werbos’ teachings to incorporate the teachings of Lathrop and provide transitioning an internal state in order to [Lathrop, para 121] give higher classification accuracy.
Regarding claim 7, Werbos teaches,
A non-transitory recording medium that stores a program having an information processing method executed by a computer [Sect I, col 1, para 2, This paper will mainly describe a simpler version of backpropagation, which can be translated into computer code and applied directly by neural network users; Sect I, col 2, para 3, This allows one to calculate the derivatives needed when optimizing an iterative analysis procedure, a neural network with memory, or a control system which maximizes performance over time] comprising:
an input layer configured to acquire input data [Fig. 5, Input layer];
an intermediate layer configured to transition (adapt the parameters of the network) associated with the input data from the input layer [Fig. 5, Hidden layer; Sect 2, pg. 1550, para 2, The goal is to adapt the parameters of the network so that it performs well for patterns from outside the training set; Sect 2, pg. 1551, col 1, para 3, We must specify the "topology" (connections and equations) for a network which inputs X(t) and outputs a four-component vector Y ̂(t)];
an output layer [Fig. 5, Output layer] configured to output output data based on values acquired by multiplying output of the intermediate layer and weight coefficients [Sect 2, pg. 1551, col 1, para 3, The relation between the inputs and outputs must depend on a set of weights (parameters) W which can be adjusted; Sect 2, pg. 1552, col 1, Subroutine pseudocode, net = net + W(i, j)*x(j); Sect 2, pg. 1552, col 1, In the pseudocode, note that X and Ware technically the inputs to the subroutine, while x and Yhat are the outputs. Yhat is usually regarded as "the" output of the network];
PNG
media_image1.png
257
770
media_image1.png
Greyscale
The information processing method comprising: acquiring input data [Fig. 5, Input Layer];
transition associated with the input data a plurality of times to obtain a plurality of results for each transition [Sect 2, pg. 1551, col 1, para 2, Before adapting the parameters of the neural network, one must first obtain a training database of actual handwritten digits and correct classifications. Suppose, for example, that this database contains 2000 examples of handwritten digits. In that case, T = 2000. We may give each example a label t between 1 and 2000. For each sample t, we have a record of the input pattern and the correct classification. Each input pattern consists of 380 numbers, which may be viewed as a vector with 380 components; we may call this vector X(t)];
performing weighting on each of the plurality results; outputting the output data based on a sum of a plurality of weighted results (Sect IV, pg. 1558, col 1, para 3, running sum) [Sect 2, pg. 1551, col 1, para 3 & 4, We must specify the "topology" (connections and equations) for a network which inputs X(t) and outputs a four-component vector Y ̂(t), an approximation to Y(t). The relation between the inputs and outputs must depend on a set of weights (parameters) W which can be adjusted. We must specify a "learning rule"-a procedure for adjusting the weights W so as to make the actual outputs P(t) approximate the desired outputs Y(t); Sect IV, pg. 1558, col 1, para 3, The subroutine F-ACTION is virtually identical to the old subroutine F-NET, except that we need to calculate F-W as a running sum (as we did in F-NET2)]; and
performing learning of a weight obtained by the weighting [Sect 2, pg. 1551, col 1, para 4, We must specify a "learning rule"-a procedure for adjusting the weights W so as to make the actual outputs P(t) approximate the desired outputs Y(t)] , thereby not increasing a size of the intermediate layer [Sect II, pg. 1552, col 2, para 1, we calculate the derivatives of E with respect to all of the weights; this is indicated by the dotted lines in Fig. 3. If increasing a given weight would lead to more error, we adjust that weight downwards].
Werbos does not teach layer configured to transition an internal state associated with the input; transition the internal state associated with the input data a plurality of times to obtain results; results for each of the transitioned internal states.
Lathrop teaches,
layer configured to transition an internal state associated with the input; transition the internal state associated with the input data a plurality of times to obtain results [Para 121, in each clock cycle, each value u of an input signal provided by input layer 110 may represent new source data, such as a new image. Once a new image has been loaded to the input layer 110, the internal network 120 can begin to asynchronously process the image input signal, and the node states in the internal network 120 may continuously change… The delay between the input layer clock 1710 and output layer clock 1720 can be a parameter of device 1 00 that can be tuned to optimize the classification accuracy for a given task];
results for each of the transitioned internal states [Para 83, In general, reset line I 100 can provide any reset signal pattern that causes all nodes 121 connected to reset lines to be reconfigured to a desired state. For example, the reset line I 100 may be a single digital channel with value r = {0,1} that connects to inputs of all the nodes 121 in the internal network 120. For example, when r = 0, outputs from nodes 121 can be a function of their other inputs and when r = 1, node outputs from nodes 121 can all be set equal to I and the network state can be reset to an all " I " state].
Lathrop is analogous to the claimed invention as they both relate to reservoir computing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Werbos’ teachings to incorporate the teachings of Lathrop and provide transitioning an internal state in order to [Lathrop, para 121] give higher classification accuracy.
Claim(s) 3 and 4 are rejected under 35 U.S.C. 103 as being unpatentable over Werbos in view of Lathrop, and in further view of Soures et al. (Reservoir Computing in Embedded Systems, published 2017), hereinafter Soures.
Regarding claim 3, Werbos-Lathrop teaches the limitations of claim 1 including the processor configured to execute the instructions (Werbos, Sect I, col 2, para 3 and Sect I, col 1, para 2) internal states (Lathrop, Para 121), the input data (Werbos, Sect 2, pg. 1551, col 1, para 1), and the input layer (Werbos, Fig. 5).
Werbos does not teach transition a plurality of times during a period of time from a moment where acquired input data to a moment where acquired next input data
Soures further teaches,
transition (Pg. 69, col 2, para 2, delay-differential equation (DDE)) a plurality of times (Pg. 69, col 2, para 2, At each time step t) during a period of time (Pg. 69, col 2, para 2, duration of
τ
) from a moment where acquired input data (Pg. 69, col 2, para 2, reservoir’s input is sampled) to a moment where acquired next input data (Pg. 69, col 2, para 2, At each time step t) [Pg. 69, col 2, para 2, The activation function is typically governed by a delay-differential equation (DDE), which has the form
x
˙
=
f
t
,
x
t
,
x
(
t
-
τ
)
. At each time step t, the reservoir’s input is sampled and held for a duration of
τ
, which is evenly divided into H smaller time segments with durations of
θ
. For each of the H segments, the held input is multiplied by an input weight
w
i
,
j
i
n
].
Soures is analogous to the claimed invention as they both relate to reservoir computing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Werbos’s teachings to incorporate the teachings of Soures and provide transition during a period of time until next input data is acquired [Soures, Pg. 69, col 2, para 1] in order to limit the degree of connectivity which improves efficiency by reducing the allocation of resources required during dedicated connection.
Regarding claim 4, Werbos-Lathrop teach the limitations of claim 1 including the processor configured to execute the instructions (Werbos, Sect I, col 2, para 3 and Sect I, col 1, para 2), the internal states (Lathrop, Para 121), and the input layer (Werbos, Fig. 5).
Werbos does not teach transition the plurality of times during a period of time from a first moment where the input layer acquires the input data to a second moment where the input layer acquires next input data; And upon the input layer acquiring the next input data, start transitioning the next input data from a state prior to transitioning at least some of the plurality of times.
Soures further teaches,
transition (Pg. 69, col 2, para 2, delay-differential equation (DDE)) plurality of times (Pg. 69, col 2, para 2, At each time step t) during a period of time (Pg. 69, col 2, para 2, durations of
θ
) from a first moment where layer acquires input data (Pg. 69, col 2, para 2, reservoir’s input is sampled) to a second moment where layer acquires next input data (Pg. 69, col 2, para 2, At each time step t) [Pg. 69, col 2, para 2, The activation function is typically governed by a delay-differential equation (DDE), which has the form
x
˙
=
f
t
,
x
t
,
x
(
t
-
τ
)
. At each time step t, the reservoir’s input is sampled and held for a duration of
τ
, which is evenly divided into H smaller time segments with durations of
θ
. For each of the H segments, the held input is multiplied by an input weight
w
i
,
j
i
n
… The weighted input is then added to the delayed state of the reservoir node,
x
(
t
-
τ
)
, and fed back into the reservoir node at the current time step. In this way, H components makeup the reservoir’s state corresponding to each sampled and held input. This approach is attractive because the hardware implementation usually consists of a simple circuit and a delay line without the routing overhead associated with ESNs and LSMs];
And upon layer acquiring the next input data (Pg. 69, col 2, para 2, At each time step t, the reservoir’s input is sampled), start transitioning (Pg. 69, col 2, para 2, delay-differential equation (DDE)) the next input data from a state prior (Pg. 69, col 2, para 2, weighted input is then added to the delayed state) to transitioning at least some of the plurality of times (Pg. 69, col 2, para 2, fed back into the reservoir node at the current time step) [Pg. 69, col 2, para 2, The activation function is typically governed by a delay-differential equation (DDE), which has the form
x
˙
=
f
t
,
x
t
,
x
(
t
-
τ
)
. At each time step t, the reservoir’s input is sampled and held for a duration of
τ
, which is evenly divided into H smaller time segments with durations of
θ
. For each of the H segments, the held input is multiplied by an input weight
w
i
,
j
i
n
… The weighted input is then added to the delayed state of the reservoir node,
x
(
t
-
τ
)
, and fed back into the reservoir node at the current time step. In this way, H components makeup the reservoir’s state corresponding to each sampled and held input.].
Soures is analogous to the claimed invention as they both relate to reservoir computing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Werbos’s teachings to incorporate the teachings of Soures and provide performing transitions during a period of time until next input data is acquired [Soures, Pg. 69, col 2, para 1] in order to limit the degree of connectivity which improves efficiency by reducing the allocation of resources required during dedicated connection.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SYED RAYHAN AHMED whose telephone number is (571)270-0286. The examiner can normally be reached Mon-Fri ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SYED RAYHAN AHMED/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126