Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 8/26/2026 have been fully considered but they are not persuasive.
Applicant argues that the current claims provide a “specific technical solution… [that] improves upon the prior art solutions.” Remarks 16. The prior art of record shows that that the claimed compression (tensorized neural networks) is known in the prior art, and therefore there is no improvement over the prior art solutions. The claims are just related to the technical field of tensorized neural networks.
With respect to whether the claims recite additional element that integrate the abstract idea into a practical application (step 2b) Applicant argues, that the claimed conversion process “ensures that the existing neural network… is converted into a more efficient format…” Remarks 20. The conversion is not an additional element, the conversion is part of the abstract idea. Therefore, the conversion cannot amount to significantly more than the abstract idea, because the conversion is the abstract idea.
In response to the remarks about claim 4, (Remarks 21) amendments to the material from claim 4 necessitated new art.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-3,5-11 and 18-27 rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract idea of a mathematical relationship without significantly more. The claims recite the abstract idea of:
1. (Currently Amended) …
convert, into a tensorized neural network, a predetermined machine learning routine in the form of a neural network and associated with a target machine or system or process by:
converting one or more layers of a plurality of layers of the neural network into respective one or more tensor networks; and
converting one or more first non-linearities applicable to the converted one or more layers of the neural network into one or more second non-linearities applicable to each tensor of the respective one or more tensor networks; and
add at least one gauge optimization in the tensorized neural network, the at least one gauge optimization being selected from a predetermined set of gauge optimizations and including a trainable parameter that tunes the respective gauge optimization;
…
produce at least one output about the target machine or system or process, the at least one output being inferred by the converted predetermined machine learning routine upon inputting a data set therein.
2. (Original) The apparatus of claim 1, wherein the predetermined machine learning routine is a trained machine learning routine that is converted into the tensorized neural network.
3. (Currently Amended) The apparatus of claim 1, wherein the converting of the one or more first non-linearities is performed such that the one or more second non-linearities partially or completely reproduce a behavior of the one or more first non-linearities.
5. (Original) The apparatus of claim 1, further comprising at least one memory adapted to store the neural network and the tensorized neural network, wherein the tensorized neural network occupies less space in the at least one memory than the neural network.
6. (Original) The apparatus of claim 1, wherein the converted one or more layers of the plurality of layers of the neural network comprises all layers of the plurality of layers of the neural network.
…
18. (Currently Amended) A method, comprising:
converting, into a tensorized neural network, a predetermined machine learning routine in the form of a neural network and associated with a target machine or system or process by:
converting one or more layers of a plurality of layers of the neural network into respective one or more tensor networks; and
converting one or more non-linearities applicable to the converted one or more layers of the neural network into one or more non-linearities applicable to each tensor of the respective one or more tensor networks;
adding at least one gauge optimization in the tensorized neural network, the at least one gauge optimization being selected from a predetermined set of gauge optimizations and
including a trainable parameter that tunes the respective gauge optimization;
training the predetermined machine learning routine in the form of the tensorized neural network with a training data set such that the training adjusts the trainable parameter of the at least one gauge optimization; and
producing at least one output about the target machine or system or process, the at least one output being inferred by the converted predetermined machine learning routine upon inputting a data set therein.
… [Claims 19-22 are similar to Claims 1-11 and likewise rejected]
23. (New) The method of claim 18, wherein the converting of the one or more first non-linearities is performed such that the one or more second non-linearities partially or completely reproduce a behavior of the one or more first non-linearities.
24. (New) …
converting, into a tensorized neural network, a predetermined machine learning routine in the form of a neural network and associated with a target machine or system or process by:
converting one or more layers of a plurality of layers of the neural network into respective one or more tensor networks;
converting one or more non-linearities applicable to the converted one or more layers of the neural network into one or more non-linearities applicable to each tensor of the respective one or more tensor networks;
adding at least one gauge optimization in the tensorized neural network, the at least one gauge optimization being selected from a predetermined set of gauge optimizations and including a trainable parameter that tunes the respective gauge optimization;
training the predetermined machine learning routine in the form of the tensorized neural network with a training data set such that the training adjusts the trainable parameter of the at least one gauge optimization; and
producing at least one output about the target machine or system or process, the at least one output being inferred by the converted predetermined machine learning routine upon inputting a data set therein.
25. (New) The non-transitory computer-readable storage medium of claim 24, wherein the converting of the one or more non-linearities applicable to the converted one or more layers of the neural network is performed such that the one or more non-linearities applicable to each tensor of the respective one or more tensor networks partially or completely reproduce a behavior of the one or more non-linearities applicable to the converted one or more layers of the neural network.
26. (New) The non-transitory computer-readable storage medium of claim 24, wherein the instructions cause the at least one computing device to further obtain at least part of the data set from at least one or more sensors or one or more computing devices or a combination thereof.
27. (New) The non-transitory computer-readable storage medium of claim 24, wherein the instructions cause the at least one computing device to further provide, at least based on the at least one output, at least one instruction for actuation of one or more actuators or controllers or a combination thereof of the target machine or system or process.
This judicial exception is not integrated into a practical application because the additional elements of:
various industrial applications of neural networks (e.g. claims 8-11, 19-22 and 26 and 27), is merely linking to a field of use; and
Training the neural network is mere instructions to use computers as a tool to perform an existing process.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements of memory and classical processors are claims to generic computer parts.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 5-8, 10-11, 18-21 and 23-26 are rejected under 35 U.S.C. 103 as being unpatentable over Nonlinear tensor train format for deep neural network compression by Wang et al, US20220108218A1 to Wall et al and US20220026879A1 to Kale.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Nonlinear tensor train format for deep neural network compression by Wang et al, US20220108218A1 to Wall et al, US20220026879A1 to Kale and Tensor-based framework for training flexible neural networks by Zniyed et al.
Claims 9, 22 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Nonlinear tensor train format for deep neural network compression by Wang et al, US20220108218A1 to Wall et al, US20220026879A1 to Kale and US 20210065010 A1 to Pfeil.
Wang teaches claims 1, 18 and 24. An apparatus comprising at least one classical processor configured to:
convert, into a tensorized neural network, a predetermined machine learning routine (Wang abs “DNN”) in the form of a neural network and associated with a target machine or system or process by: (The system is the system with the DNN in it.)
converting one or more layers of a plurality of layers of the neural network into respective one or more tensor networks; and (Wang sec. 3.1 “compressing a typical FC layer in DNNs y =f(xW) in tensorizing way…” compressing in a tensorizing way is converting a layer to a tensor network.)
converting one or more first non-linearities applicable to the converted one or more layers of the neural network into one or more second non-linearities applicable to each tensor of the respective one or more tensor networks; (Wang sec. 3.1 To obtain the corresponding mapping way from Eq. (9), all the Gk should be separated as new independent weights based on Eq. (7). Naturally, inserting extra nonlinear activation functions after contracting every Gk could reach this requirement, then we have… which we term it nonlinear TT or NTT [non-linear tensor train].” Wang abs. “a novel nonlinear tensor train (NTT) format, which contains extra nonlinear activation functions embedded in sequenced contractions and convolutions on the top of the normal TT decomposition and the proposed TT format connected by convolutions, to compensate the accuracy loss that normal TT cannot give.”)
train the predetermined machine learning routine in the form of the tensorized neural network with a training data set (Wang sec. 3.1 To obtain the corresponding mapping way from Eq. (9), all the Gk should be separated as new independent weights based on Eq. (7).
Wang doesn’t teach gauge optimization.
However, Wall teaches how to add at least one gauge optimization in the tensorized neural network, the at least one gauge optimization being selected from a predetermined set of gauge optimizations and including a trainable parameter that tunes the respective gauge optimization; (Wall para 57 “the optimal choice of gauge may require a global optimization across all tensors. … The heuristic guiding the scheme can be to ensure that operations are as ‘diagonal’ as possible…” Wall para 59 “ While the location of the diagonality center can again be used as an optimization parameter, the diagonality center may be set to an isometry that is initially an identity matrix.”)
train the predetermined machine learning routine in the form of the tensorized neural network with a training data set such that the training adjusts the trainable parameter of the at least one gauge optimization; and (Wall para 34 “optimization may be performed where a TN model may be learned from a collection of quantum training data.” Wall para 49 “optimized by a DMRG-style procedure using gradient descent, where the gradient may be taken with respect to the tensors of the MPS.”)
Wang, Wall and the claims are all directed to machine learning. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to gauge optimize because “risk of errors can be substantially reduced.” Wall para 27.
Wang doesn’t make predictions.
However, Kale teaches how to produce at least one output about the target machine or system or process, the at least one output being inferred by the converted predetermined machine learning routine upon inputting a data set therein. (Kale para 34 “The system further includes a computing device (e.g., an SSD that includes the memory and a processing device) to predict a maintenance service for one or more of the machines based on an output from an artificial neural network (ANN).”)
Kale, Wang and the claims all use a neural network. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to use an ANN to do machine control/maintenance because “the prediction allows the service to be scheduled at a convenient time, for which it is more technically and/or cost efficient to implement the service.” Kale para 29.
Wang teaches claim 2. The apparatus of claim 1, wherein the predetermined machine learning routine is a (Wang uses pre-trained weights W, Wang abs, “tensorizing neural weights into higher-order tensors for better decomposition, and directly mapping efficient tensor structure to neural architecture with nonlinear activation functions…” Wang is not just tensorizing random weights, or weights set to some arbitrary starting value, that means the weights are trained weights. Wang sec. 4.2.2 also says that the DNN is trained on MNIST and UCF11 before the weights are compressed.)
Wang is not explicit about the trained NN.
However, Zniyed teaches tensorizing a trained machine learning routine. (Zniyed abs “This technique fuses the first and zeroth order information of the NN, where the first-order information is contained in a Jacobian tensor, following a constrained canonical polyadic decomposition (CPD). The proposed algorithm can handle different decomposition bases. The goal of this method is to compress large pretrained NN models, by replacing subnetworks, i.e., one or multiple layers of the original network, by a new flexible layer. The approach is applied to a pretrained convolutional neural network (CNN) used for character classification.” (emphasis added))
Zniyed, Wang and the claims all tensorize NNs. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to tensorize an already trained network because the goal of tensorizing is to “compress large pretrained NN models…” Zniyed abs.
Wang teaches claims 3. The apparatus of claim 1, wherein the converting of the one or more first non-linearities is performed such that the one or more second non-linearities partially or completely reproduce a behavior of the one or more first non-linearities. (Wang abs. “we propose a novel nonlinear tensor train (NTT) format, which contains extra nonlinear activation functions embedded in sequenced contractions and convolutions on the top of the normal TT decomposition and the proposed TT format connected by convolutions, to compensate the accuracy loss that normal TT cannot give…. our NTT format can almost maintain the accuracy at least on MNIST, UCF11 and CIFAR-10 datasets…”)
Wang teaches claim 5. The apparatus of claim 1, further comprising at least one memory adapted to store the neural network and the tensorized neural network, wherein the tensorized neural network occupies less space in the at least one memory than the neural network. (Wang sec. 6 “a novel compression method termed as NTT format for both weight matrices and convolutional kernels. Our NTT can significantly ameliorate the accuracy loss existing widely in traditional TT DNNs…”)
Wang teaches claim 6. The apparatus of claim 1, wherein the converted one or more layers of the plurality of layers of the neural network comprises all layers of the plurality of layers of the neural network. (Wang sec. 3.1 tensorizes all “d layers” in the DNN.)
Wang teaches claim 7. The apparatus of claim 1, further configured to train the predetermined machine learning routine in the form of the tensorized neural network with a training data set. (Wang sec. 4.2.2 “For MNIST dataset, we follow …to make each image as a sequence of vectors, and design two hidden layers with 625 and 1296 neurons, respectively. Concretely, 625 is … During training, the initial learningrateis0.0001…” MNIST is the training dataset. Wang sec. 4.3.1 “Except training the above 3 CNNs in original, TT, and NTT formats, NTT with residual connections …, have also been examined.” NTT is the tensorizes format.)
Kale teaches claim 8. The apparatus of claim 1, further configured to obtain at least part of the data set from at least one or more sensors or one or more computing devices or a combination thereof. (Kale abs. “An artificial neural network (ANN) is configured to receive the sensor data stream…”)
Kale teaches claim 9. The apparatus of claim 1, further configured to provide, at least based on the at least one output, at least one instruction for (Kale abs. “An artificial neural network (ANN) is configured to receive the sensor data stream and predict a maintenance service for the machine based on the sensor data stream.”)
Kale doesn’t teach an actuator.
However, Pfeil teaches actuation of one or more actuators or controllers… (Pfeil para 45 “Control unit 13 controls an actuator as a function of the output variable of deep neural network 12…”)
Kale, Pfeil and the claims are NN applied to industry. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to use Kale’s NN to control an actuator so as to implement the maintenance service prescribed by Kale’s NN.
Kale teaches claims 10 and 19. The apparatus of claim 1, wherein the target machine or system or process comprises one of: a computing device or system, a factory line or a machine thereof, a factory, a production process, means of transportation or an automatic control unit thereof, an automatic transportation controlling process, an electric grid or network, an energy power plant, an electric power station, an electric power generation process, and an electrical energy allocation process. (Kale para 25 “The machine may be, for example, a manufacturing machine in an assembly line of several machines (e.g., in a semiconductor wafer fabrication facility), an autonomous vehicle, or systems that provide electrical, chemical, mechanical, data, and/or communication services for a physical structure (e.g., a building or house).”)
Kale teaches claims 11 and 20. The apparatus of claim 1, wherein the at least one output comprises one of: prediction of a failure of a machine, determination of a predictive maintenance of a machine, production amount of energy, production amount of a substance or an object, and actuation of a control unit of means of transportation. (Kale para 25 “The machine may be, for example, a manufacturing machine in an assembly line of several machines (e.g., in a semiconductor wafer fabrication facility), an autonomous vehicle, or systems that provide electrical, chemical, mechanical, data, and/or communication services for a physical structure (e.g., a building or house).” Kale abs. “An artificial neural network (ANN) is configured to receive the sensor data stream and predict a maintenance service for the machine based on the sensor data stream.”)
Kale teaches claim 21. (New) The method of claim 18, further comprising obtaining at least part of the data set from at least one or more sensors or one or more computing devices or a combination thereof. (Kale abs. “An artificial neural network (ANN) is configured to receive the sensor data stream…”)
Kale teaches claim 22. (New) The method of claim 18, further comprising providing, at least based on the at least one output, at least one instruction for (Kale abs. “An artificial neural network (ANN) is configured to receive the sensor data stream and predict a maintenance service for the machine based on the sensor data stream.”)
Kale doesn’t teach an actuator.
However, Pfeil teaches actuation of one or more actuators or controllers… (Pfeil para 45 “Control unit 13 controls an actuator as a function of the output variable of deep neural network 12…”)
Kale, Pfeil and the claims are NN applied to industry. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to use Kale’s NN to control an actuator so as to implement the maintenance service prescribed by Kale’s NN.
Wang teaches claim 23. (New) The method of claim 18, wherein the converting of the one or more first non-linearities is performed such that the one or more second non-linearities partially or completely reproduce a behavior of the one or more first non-linearities. (Wang abs. “we propose a novel nonlinear tensor train (NTT) format, which contains extra nonlinear activation functions embedded in sequenced contractions and convolutions on the top of the normal TT decomposition and the proposed TT format connected by convolutions, to compensate the accuracy loss that normal TT cannot give…. our NTT format can almost maintain the accuracy at least on MNIST, UCF11 and CIFAR-10 datasets…”)
Wang teaches claim 24. (New) A non-transitory computer-readable storage medium storing a computer program including instructions which, when the program is executed by at least one computing device with at least one classical processor, (Wang abs “DNN”) cause the at least one computing device to carry out:
converting, into a tensorized neural network, a predetermined machine learning routine in the form of a neural network and associated with a target machine or system or process by: (Wang sec. 3.1 “compressing a typical FC layer in DNNs y =f(xW) in tensorizing way…” compressing in a tensorizing way is converting a layer to a tensor network.)
converting one or more layers of a plurality of layers of the neural network into respective one or more tensor networks; (Wang sec. 3.1 “compressing a typical FC layer in DNNs y =f(xW) in tensorizing way…” compressing in a tensorizing way is converting a layer to a tensor network.)
converting one or more non-linearities applicable to the converted one or more layers of the neural network into one or more non-linearities applicable to each tensor of the respective one or more tensor networks; (Wang sec. 3.1 To obtain the corresponding mapping way from Eq. (9), all the Gk should be separated as new independent weights based on Eq. (7). Naturally, inserting extra nonlinear activation functions after contracting every Gk could reach this requirement, then we have… which we term it nonlinear TT or NTT [non-linear tensor train].” Wang abs. “a novel nonlinear tensor train (NTT) format, which contains extra nonlinear activation functions embedded in sequenced contractions and convolutions on the top of the normal TT decomposition and the proposed TT format connected by convolutions, to compensate the accuracy loss that normal TT cannot give.”)
training the predetermined machine learning routine in the form of the tensorized neural network with a training data set (Wang sec. 3.1 To obtain the corresponding mapping way from Eq. (9), all the Gk should be separated as new independent weights based on Eq. (7).)
Wang doesn’t teach gauge optimization.
However, Wall teaches how to adding at least one gauge optimization in the tensorized neural network, the at least one gauge optimization being selected from a predetermined set of gauge optimizations and including a trainable parameter that tunes the respective gauge optimization; (Wall para 57 “the optimal choice of gauge may require a global optimization across all tensors. … The heuristic guiding the scheme can be to ensure that operations are as ‘diagonal’ as possible…” Wall para 59 “ While the location of the diagonality center can again be used as an optimization parameter, the diagonality center may be set to an isometry that is initially an identity matrix.”)
training the predetermined machine learning routine in the form of the tensorized neural network with a training data set such that the training adjusts the trainable parameter of the at least one gauge optimization; and (Wall para 34 “optimization may be performed where a TN model may be learned from a collection of quantum training data.” Wall para 49 “optimized by a DMRG-style procedure using gradient descent, where the gradient may be taken with respect to the tensors of the MPS.”)
Wang, Wall and the claims are all directed to machine learning. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to gauge optimize because “risk of errors can be substantially reduced.” Wall para 27.
Wang doesn’t make predictions.
However, Kale teaches how to producing at least one output about the target machine or system or process, the at least one output being inferred by the converted predetermined machine learning routine upon inputting a data set therein. (Kale para 34 “The system further includes a computing device (e.g., an SSD that includes the memory and a processing device) to predict a maintenance service for one or more of the machines based on an output from an artificial neural network (ANN).”)
Kale, Wang and the claims all use a neural network. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to use an ANN to do machine control/maintenance because “the prediction allows the service to be scheduled at a convenient time, for which it is more technically and/or cost efficient to implement the service.” Kale para 29.
Wang teaches claim 25. (New) The non-transitory computer-readable storage medium of claim 24, wherein the converting of the one or more non-linearities applicable to the converted one or more layers of the neural network is performed such that the one or more non-linearities applicable to each tensor of the respective one or more tensor networks partially or completely reproduce a behavior of the one or more non-linearities applicable to the converted one or more layers of the neural network. (Wang abs. “we propose a novel nonlinear tensor train (NTT) format, which contains extra nonlinear activation functions embedded in sequenced contractions and convolutions on the top of the normal TT decomposition and the proposed TT format connected by convolutions, to compensate the accuracy loss that normal TT cannot give…. our NTT format can almost maintain the accuracy at least on MNIST, UCF11 and CIFAR-10 datasets…”)
Kale teaches claim 26. (New) The non-transitory computer-readable storage medium of claim 24, wherein the instructions cause the at least one computing device to further obtain at least part of the data set from at least one or more sensors or one or more computing devices or a combination thereof. (Kale abs. “An artificial neural network (ANN) is configured to receive the sensor data stream…”)
Kale teaches claim 27. (New) The non-transitory computer-readable storage medium of claim 24, wherein the instructions cause the at least one computing device to further provide, at least based on the at least one output, at least one instruction for (Kale abs. “An artificial neural network (ANN) is configured to receive the sensor data stream and predict a maintenance service for the machine based on the sensor data stream.”)
Kale doesn’t teach an actuator.
However, Pfeil teaches actuation of one or more actuators or controllers… (Pfeil para 45 “Control unit 13 controls an actuator as a function of the output variable of deep neural network 12…”)
Kale, Pfeil and the claims are NN applied to industry. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to use Kale’s NN to control an actuator so as to implement the maintenance service prescribed by Kale’s NN.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Austin Hicks whose telephone number is (571)270-3377. The examiner can normally be reached Monday - Thursday 8-4 PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AUSTIN HICKS/Primary Examiner, Art Unit 2142