DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1, 5-7, 9-12, 15, and 17-18 are presented for examination.
Response to Amendment
The prior objections to the claims, specification, and drawings have been obviated by the amendments. The prior double patenting rejections have also been obviated by the amendments. Thus, these objections/rejections are withdrawn.
Claim Rejections - 35 USC § 101
Claims 1, 5-7, 9-12, 15, and 17-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance (“2019 PEG”).
Claim 1
Step 1: The claim recites a computing device, and is therefore directed to the statutory category of machines.
Step 2A Prong 1: The claim recites:
“identifying and extracting paths of an electric circuit between a plurality of designated components that represent the electric circuit”: This limitation encompasses mentally identifying and extracting paths of a graphical representation of an electric circuit.
“converting at least one extracted path of the extracted paths to a path embedding comprising a vector of a fixed length”: This limitation encompasses mentally converting at least one of the extracted paths to a path embedding comprising a vector of a fixed length.
“predicting…an efficiency rating and a voltage output of the plurality of designated components based on an input of circuit parameters and the path embedding of the electric circuit”: This limitation encompasses mentally predicting an efficiency rating and a voltage output of the plurality of designated components based on circuit parameters and the path embedding of the electric circuit.
“concatenating and linearly transforming outputs of the attention mechanism functions…”: This limitation encompasses mentally concatenating and linearly transforming outputs.
“map…[ping] the electric circuit to a scalar value”: This limitation encompasses mentally mapping the electric circuit to a scalar value.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “A computing device comprising: a processor; a storage device coupled to the processor, wherein the storage device stores instructions to cause the processor to perform acts to provide circuit performance modeling.” However, this limitation amounts to mere instructions to apply a judicial exception using a generic computer (MPEP § 2106.05(f)). The claim further recites “training, by a training model, a circuit representation-learning model to perform the circuit performance modeling,” however, this limitation merely generally links the above-mentioned abstract ideas to the technological environment of model training (MPEP § 2106.05(h)). The claim further recites “wherein the circuit representation-learning model comprises a transformer model including a stack of multi-head attention modules,” that the predicting is performed “by the circuit representation-learning model,” and “wherein the predicting comprises: operating, by the transformer model, attention mechanism functions of the multi-head attention modules in parallel,” however, these limitations amount to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)). The claim further recites that the concatenating and linearly transforming outputs is “for input to a multi-layer perceptron network” and that the multi-layer perceptron network performs the mapping, however this limitation amounts to mere instruct to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)), as it merely amounts to applying the abstract idea using a generic MLP network.
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above. As an ordered whole, the claim is directed to a mentally performable process of identifying paths of a circuit, converting at least one of the paths to a vector of a fixed length, and predicting an efficiency rating and a voltage output of the circuit. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Claim 5
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“embedding the circuit parameters of the electric circuit…”: This limitation encompasses mentally embedding the circuit parameters.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “as an input to the transformer model.” However, this limitation amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The “as an input to the transformer model” limitation amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)) as stated above.
Claim 6
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites the same judicial exceptions as claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “wherein the converting of the at least one extracted path to the path embedding is performed by a bidirectional Long Short-Term Memory (Bi-LSTM) network.” However, this limitation amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)). The claim also recites “wherein the instructions cause the processor to perform an additional act comprising outputting, by the Bi-LSTM network, the vector of the fixed length for each path embedding.” However, this limitation amounts to the insignificant extra-solution activity of mere data gathering and outputting (MPEP § 2106.05(g)).
Step 2B: The claim does not contain significantly more than the judicial exception. The Bi-LSTM network limitation amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)) as stated above. The outputting the vector of a fixed length for each path embedding limitation, in addition to being insignificant extra-solution activity, is also directed to the well-understood, routine, and conventional activity of storing and retrieving information in memory (MPEP § 2106.05(d)(II) Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93).
Claim 7
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“representing the electric circuit as a device embedding…”: This limitation encompasses mentally representing the electric circuit as a device embedding by mentally converting the representation of the circuit into a vector embedding.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites that the device embedding is “input to the Bi-LSTM network.” However, this limitation amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The “input to the Bi-LSTM network” limitation amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)) as stated above.
Claim 9
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites the same judicial exceptions as claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “wherein the training model comprises a stochastic gradient descent-based model.” However, this limitation merely further limits the training model that performs the training limitation of claim 1, which is still merely generally linking the above-mentioned abstract ideas to the technological environment of model training (MPEP § 2106.05(h)).
Step 2B: The claim does not contain significantly more than the judicial exception. The training limitation amounts to generally linking the judicial exception to a particular technological environment (MPEP § 2106.05(h)) as stated above.
Claim 10
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites the same judicial exceptions as claim 9.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “processing, by the stochastic gradient descent-based model, parameters in the path embedding, the transformer model, and the multi-layer perceptron network.” However, this limitation merely generally links the above-mentioned abstract ideas to the technological environment of model training (MPEP § 2106.05(h)).
Step 2B: The claim does not contain significantly more than the judicial exception. The processing parameters limitation amounts to generally linking the above-mentioned abstract ideas to the technological environment of model training (MPEP § 2106.05(h)) as stated above.
Claim 11
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites the same judicial exceptions as claim 10 above.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “wherein the multi-layer perceptron network has an input size that is the same as an output size of the transformer model.” However, this limitation merely further limits the multi-layer perceptron network that performs the mental process of mapping the circuit to a scalar value, and still amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The multi-layer perceptron network limitation amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)) as stated above.
Claim 12
Step 1: The claim recites a computer-implemented method of circuit performance modeling, and therefore is directed to the statutory category of processes.
Step 2A Prong 1: The claim recites:
“identifying and extracting paths of an electric circuit between a plurality of designated components that represent the electric circuit”: This limitation encompasses mentally identifying and extracting paths of a graphical representation of an electric circuit.
“converting one or more of the extracted paths to respective path embeddings including a corresponding vector of a fixed length”: This limitation encompasses mentally converting one or more of the extracted paths to respective path embeddings including a corresponding vector of a fixed length.
“predicting…an efficiency rating and a voltage output of the electric circuit based on an input of circuit parameters and the respective path embeddings of the electric circuit”: This limitation encompasses mentally predicting an efficiency rating and a voltage output of the plurality of designated components based on circuit parameters and the path embedding of the electric circuit.
“concatenating and linearly transforming outputs of the attention mechanism functions…”: This limitation encompasses mentally concatenating and linearly transforming outputs.
“map…[ping] the electric circuit to a scalar value”: This limitation encompasses mentally mapping the electric circuit to a scalar value.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “training, by a training model, a circuit representation-learning model to perform the circuit performance modeling,” however, this limitation merely generally links the above-mentioned abstract ideas to the technological environment of model training (MPEP § 2106.05(h)). The claim further recites “wherein the circuit representation-learning model comprises a transformer model including a stack of multi-head attention modules,” that the predicting is performed “by the circuit representation-learning model,” and “wherein the predicting comprises: operating, by the transformer model, attention mechanism functions of the multi-head attention modules in parallel,” however, these limitations amount to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)). The claim further recites that the concatenating and linearly transforming outputs is “for input to a multi-layer perceptron network” and that the multi-layer perceptron network performs the mapping, however this limitation amounts to mere instruct to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)), as it merely amounts to applying the abstract idea using a generic MLP network.
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above. As an ordered whole, the claim is directed to a mentally performable process of identifying paths of a circuit, converting at least one of the paths to a vector of a fixed length, and predicting an efficiency rating and a voltage output of the circuit. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Claim 15
Step 1: A process, as above.
Step 2A Prong 1: The claim recites:
“embedding the circuit parameters of the electric circuit…”: This limitation encompasses mentally embedding the circuit parameters.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “as an input to the transformer model. However, this limitation amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The “as an input to the transformer model” limitation amounts to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)) as stated above.
Claim 17
Step 1: A process, as above.
Step 2A Prong 1: The claim recites the same judicial exceptions as claim 12.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “training the training model to process parameters for the respective path embeddings.” However, this limitation merely generally links the above-mentioned abstract ideas to the technological environment of model training (MPEP § 2106.05(h)).
Step 2B: The claim does not contain significantly more than the judicial exception. The “training the training model to process parameters for the respective path embeddings” limitation amounts to generally linking the above-mentioned abstract ideas to the technological environment of model training (MPEP § 2106.05(h)) as stated above.
Claim 18
Step 1: The claim recites a computer program product, and therefore is directed to the statutory category of articles of manufacture.
Step 2A Prong 1: The claim recites:
“identifying and extracting paths of an electric circuit between a plurality of designated components that represent the electric circuit”: This limitation encompasses mentally identifying and extracting paths of a graphical representation of an electric circuit.
“converting one or more of the extracted paths to respective path embeddings including a corresponding vector of a fixed length”: This limitation encompasses mentally converting one or more of the extracted paths to respective path embeddings including a corresponding vector of a fixed length.
“predicting…an efficiency rating and a voltage output of the electric circuit based on an input of circuit parameters and the respective path embeddings of the electric circuit”: This limitation encompasses mentally predicting an efficiency rating and a voltage output of the plurality of designated components based on circuit parameters and the path embedding of the electric circuit.
“concatenating and linearly transforming outputs of the attention mechanism functions…”: This limitation encompasses mentally concatenating and linearly transforming outputs.
“map…[ping] the electric circuit to a scalar value”: This limitation encompasses mentally mapping the electric circuit to a scalar value.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “training, by a training model, a circuit representation-learning model to perform the circuit performance modeling,” however, this limitation merely generally links the above-mentioned abstract ideas to the technological environment of model training (MPEP § 2106.05(h)). The claim further recites “wherein the circuit representation-learning model comprises a transformer model including a stack of multi-head attention modules,” that the predicting is performed “by the circuit representation-learning model,” and “wherein the predicting comprises: operating, by the transformer model, attention mechanism functions of the multi-head attention modules in parallel,” however, these limitations amount to mere instructions to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)). The claim further recites that the concatenating and linearly transforming outputs is “for input to a multi-layer perceptron network” and that the multi-layer perceptron network performs the mapping, however this limitation amounts to mere instruct to apply a judicial exception using a generic computer programmed with a generic class of computer algorithms (MPEP § 2106.05(f)), as it merely amounts to applying the abstract idea using a generic MLP network.
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above. As an ordered whole, the claim is directed to a mentally performable process of identifying paths of a circuit, converting at least one of the paths to a vector of a fixed length, and predicting an efficiency rating and a voltage output of the circuit. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
Claims 1, 5, 12, 15, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Cao et al. (US20240273272) (“Cao2”) in view of Hakhamaneshi et al. (“Pretraining Graph Neural Networks for Few-Shot Analog Circuit Modeling and Design”) (“Hakhamaneshi”), further in view of Sarkar et al. (US20220207351), and further in view of Nath et al. (“TransSizer: A Novel Transformer-Based Fast Gate Sizer”).
Regarding claim 1, Cao2 discloses
“…identifying and extracting paths of an electric circuit between a plurality of designated components that represent the electric circuit (Cao2, [0038]: “In S2, static timing analysis is performed on the circuit after placement in S1, and the timing and physical information of all the stages of cell in the path is extracted from the static timing analysis report and the layout information to form the feature sequences of the path” and [0037]: “For example, 80% of all paths extracted from six of the nine ISCAS and OpenCores circuits are randomly selected as training set data for training the model, and the remaining 20% of the paths extracted from the six circuits are used for verifying the prediction performance of the model on the known circuits, and all paths extracted from the remaining three circuits are used for verifying the prediction performance of the model on an unknown circuit”);
converting at least one extracted path of the extracted paths to a path embedding comprising a vector of a fixed length (Cao2, [0042]: “S33: Feature sequence padding is performed to ensure that the feature sequences have the same length. A maximum feature sequence length of the training set data is set as max_len, and feature sequences with a length less than max_len are filled with “0” at the end until the length of the feature sequences is max_len” and [0044]: “S41: an input feature sequence with a dimension of (samples, max_len) is converted into a tensor with a dimension of (samples,max_len,dim.sub.k), wherein samples is the number of samples, max_len is a maximum path length, dim.sub.k is a designated word vector dimension of a k.sup.th feature in an embedding layer, k=1, 2, . . . , n, and n is the number of input feature sequences; in the positional encoding process, trigonometric functions shown by formula (2) and formula (3) are used to assist the network in understanding the positional relationship of features of all the stages of cells in the path…for each feature sequence, the tensor output after positional encoding is added with a tensor output after input embedding, such that n new tensors with a dimension of (samples, max_len,dim.sub.k) are obtained” and [0045]: “S42: the n new tensors obtained in S41 are merged to obtain a tensor with a dimension of (samples,max_len,dim)”; Examiner notes that the final merged tensor in S42 resulting from embedding/encoding the feature sequences of a path corresponds to “a path embedding comprising a vector of a fixed length);
training, by a training model, a circuit representation-learning model to perform the circuit performance modeling, wherein the circuit representation-learning model comprises a transformer model including a… multi-head attention module[[s]] (Cao2, [0043]: “In S4, the transformer network comprises input embedding and positional encoding, a multi-head self-attention mechanism, a fully connected feedforward network, and adding and normalization” and [0036]: “S4: a post-routing path delay prediction model is established” and [0037]: “S5: the model established in S4 is trained and verified… During training, an Adam optimizer is used, the learning rate is 0.001, the number of training batches is 1080, and the loss function is the root-mean-square error (RMSE)”; Examiner notes that the post-routing path delay prediction model corresponds to “a circuit representation- learning model” and the Adam optimizer corresponds to “a training model”); and
predicting, by the circuit representation-learning model… [a post-routing path delay] based on an input of… the path embedding of the electric circuit (Cao, [0045]: “S42: the n new tensors obtained in S41 are merged to obtain a tensor with a dimension of (samples,max_len,dim), which is used as an input X of the multi-head self-attention mechanism” and [0050]: “…finally, the pre-routing path delay and the pre-routing and post-routing path delay residual are added to obtain the final predicted post-routing path delay”), wherein the predicting comprises:
operating, by the transformer model, attention mechanism functions of the multi-head attention module[[s]] in parallel (Cao2, [0045]: “…h is the number of heads of the self-attention mechanism, and Q, K and V represent query, key and value. An attention function is performed on the h groups of matrices Q.sub.i, K.sub.i and V.sub.i parallelly, wherein a calculation formula of the dot product-based attention mechanism is formula (4)”); and
concatenating and linearly transforming outputs of the attention mechanism functions for input to a… [feed forward] network... (Cao2, [0046]: “Calculation results head.sub.i of the h-head attention mechanism are merged, and linear transform is performed by means of a trainable matrix W.sup.O to obtain an output Multillead(X) of the multi-head self-attention mechanism” and [0048]: “S44: the output, normalized in S43, of the multi-head self-attention mechanism is input to the fully connected feedforward neural network”).
Cao2 does not appear to explicitly disclose the further limitations of the claim.
However, Hakhmaneshi discloses “predicting, by the circuit representation-learning model…a voltage output of the plurality of designated components based on an input of circuit parameters” (Hakhmaneshi, III. C. From Node Embedding to Graph Property Prediction: “To perform a graph property prediction task, we need to combine the node embeddings into a single graph embedding…Fig. 3 demonstrates this architecture for fine tuning on a graph property prediction task” and Fig. 8. “(b) Shows the test Acc@200 (e.g., 0.5% prediction error) of predicting the output voltage for our fine-tuned method (FT-PT) versus training from scratch with no knowledge transfer” and Hakmaneshi, III. A. New Graph Representation: “Each node is associated with a feature vector which describes the type of the node as well as the device parameters that the node belongs to (e.g., the value of a resistor)”; Examiner notes that node features correspond to “circuit parameters” which are used as input to predict output voltage (see Fig. 3.)) and “…transforming outputs of the attention mechanism functions for input to a multi-layer perceptron network that maps the electric circuit to a scalar value” (Hakhmaneshi, Fig. 3: “The graph embedding is then fed to an MLP to predict the output. The node to graph embedding is three cross-attention layers between a learned embedding and the node features”; Examiner notes that output voltage is a scalar value).
Hakhmaneshi and the instant application both relate to circuit representation-learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Cao2 such that the predicted value is a voltage output of the plurality of designated components, the prediction is additionally based on an input of circuit parameters, and such that the concatenated and linearly transformed outputs are for input to “a multi-layer perceptron network that maps the electric circuit to a scalar value,” and one would have been motivated to do so, as doing so would allow for determining the optimal circuit design for achieving a desired output voltage, with improved sample efficiency (see Hakhmaneshi, IV. F. Improving Sample Efficiency of Model-Based Optimization by Using Pretrained Models).
Neither Cao2 nor Hakhmaneshi appear to explicitly disclose the further limitations of the claim.
However, Sarkar discloses “A computing device, comprising: a processor, a storage device coupled to the processor, wherein the storage device stores instructions to cause the processor to perform acts to provide circuit performance modeling” (Sarkar, [0031]: “The semiconductor design system 100 includes one or more processors 121, which may be formed in a substrate configured to execute one or more machine executable instructions or pieces of software, firmware, or a combination thereof. The processors 121 can be semiconductor-based—that is, the processors can include semiconductor material that can perform digital logic. The semiconductor design system 100 can also include one or more memory devices 123. The memory devices 123 may include any type of storage device that stores information in a format that can be read and/or executed by the processor(s) 121. The memory devices 123 may store executable instructions that when executed by the processor(s) 121 are configured to perform the functions discussed herein”) and “predicting, by the circuit representation model, an efficiency rating of the plurality of designated components…based on an input of circuit parameters” (Sarkar, [0027]: “The neural network 114 may include or define one or more predictive models 124, where each predictive model 124 corresponds to a different characteristic or performance metric (e.g., efficiency, breakdown voltage, threshold voltage, etc.). For example, a predictive model 124 relating to efficiency may predict the efficiency of the semiconductor system based on a given set of inputs” and [0028]: “The semiconductor design system 100 includes an optimizer 126 configured to operate in conjunction with the predictive model(s) 124 of the neural network 114 to generate the design model 136 in accordance with input parameters 101… For example, the optimizer 126 (in conjunction with the predictive model(s) 124) may determine the process parameters 138, circuit parameters 140 and/or device parameters 142 such that characteristics (e.g., efficiency, breakdown voltage, threshold voltage, etc.) of the predictive model(s) 124 achieve a threshold result (e.g., maximized, minimized) while meeting constraints 128 and/or goals 130 of the optimizer 126” and [0023]: “The circuit parameters 140 may include parameters for the structure (e.g., connections, wiring) of a circuit and/or parameters for circuit elements as values for resistors, capacitors, and inductors, and parameters related to the size of active semiconductor devices, etc”).
Sarkar and the instant application both relate to machine learning for circuit design and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi with Sarkar such that the method is performed on a computing device, and to include predicting an efficiency rating, and one would have been motivated to do so. Doing so would allow for determining the optimal circuit parameters for achieving a desired efficiency threshold (see Sarkar, [0028]).
Neither Cao2, Hakhmaneshi, nor Sarkar appear to explicitly disclose the further limitations of the claim.
However, Nath discloses “a stack of multi-head attention modules” (Nath, 2.2 Standard Transformer: “For our specific application, we propose to use the basic transformer model architecture[17], which comprises of an encoder and a decoder, each of which is a stack of 𝑁 identical blocks. The encoder block consists of two sub-layers, a multi-head self-attention layer and a position-wise feed-forward network (FFN)”).
Nath and the instant application both relate to machine learning for circuit design and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi/Sarkar with the teachings of Nath such that the transformer model includes a stack of multi-head attention modules, and one would have been motivated to do so, as doing so would increase the ability of the model to highlight the important features of the input sequence (see Nath, 3.1 Encoder design, paragraph 1).
Regarding claim 5, the rejection of claim 1 is incorporated. Hakhmaneshi further discloses “embedding the circuit parameters of the electric circuit as an input to… [a cross attention network]” (Hakhmaneshi, C. From Node Embedding to Graph Property Prediction: “Once we have the contextualized node embeddings at the output of the GNN, we flatten them as a set of feature vectors and pass them through a cross attention network”).
Hakhmaneshi and the instant application both relate to machine learning for circuit design and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Sarkar/Nath with the teachings of Hakhmaneshi to include embedding the circuit parameters of the electric circuit as an input to the transformer model, and one would have been motivated to do so. Doing so would allow for effectively capturing the structural dependencies of a circuit (see Hakhmaneshi, II. Related Work, GNNs in Chip Design).
Regarding claim 12, Cao2 discloses
“A computer-implemented method of circuit performance modeling, the computer-implemented method comprising:
identifying and extracting paths between a plurality of designated components that represent an electric circuit (Cao2, [0038]: “In S2, static timing analysis is performed on the circuit after placement in S1, and the timing and physical information of all the stages of cell in the path is extracted from the static timing analysis report and the layout information to form the feature sequences of the path” and [0037]: “For example, 80% of all paths extracted from six of the nine ISCAS and OpenCores circuits are randomly selected as training set data for training the model, and the remaining 20% of the paths extracted from the six circuits are used for verifying the prediction performance of the model on the known circuits, and all paths extracted from the remaining three circuits are used for verifying the prediction performance of the model on an unknown circuit”);
converting one or more extracted paths to respective path embeddings including a vector of a fixed length (Cao2, [0042]: “S33: Feature sequence padding is performed to ensure that the feature sequences have the same length. A maximum feature sequence length of the training set data is set as max_len, and feature sequences with a length less than max_len are filled with “0” at the end until the length of the feature sequences is max_len” and [0044]: “S41: an input feature sequence with a dimension of (samples, max_len) is converted into a tensor with a dimension of (samples,max_len,dim.sub.k), wherein samples is the number of samples, max_len is a maximum path length, dim.sub.k is a designated word vector dimension of a k.sup.th feature in an embedding layer, k=1, 2, . . . , n, and n is the number of input feature sequences; in the positional encoding process, trigonometric functions shown by formula (2) and formula (3) are used to assist the network in understanding the positional relationship of features of all the stages of cells in the path…for each feature sequence, the tensor output after positional encoding is added with a tensor output after input embedding, such that n new tensors with a dimension of (samples, max_len,dim.sub.k) are obtained” and [0045]: “S42: the n new tensors obtained in S41 are merged to obtain a tensor with a dimension of (samples,max_len,dim)”; Examiner notes that the final merged tensor in S42 resulting from embedding/encoding the feature sequences of a path corresponds to “a path embedding comprising a vector of a fixed length);
training, by a training model, a circuit representation-learning model to perform the circuit performance modeling, wherein the circuit representation-learning model comprises a transformer model including a… multi-head attention module[[s]] (Cao2, [0043]: “In S4, the transformer network comprises input embedding and positional encoding, a multi-head self-attention mechanism, a fully connected feedforward network, and adding and normalization” and [0036]: “S4: a post-routing path delay prediction model is established” and [0037]: “S5: the model established in S4 is trained and verified… During training, an Adam optimizer is used, the learning rate is 0.001, the number of training batches is 1080, and the loss function is the root-mean-square error (RMSE)”; Examiner notes that the post-routing path delay prediction model corresponds to “a circuit representation- learning model” and the Adam optimizer corresponds to “a training model”); and
predicting, by the circuit representation-learning model… [a post-routing path delay] based on an input of… the respective path embeddings of the electric circuit (Cao, [0045]: “S42: the n new tensors obtained in S41 are merged to obtain a tensor with a dimension of (samples,max_len,dim), which is used as an input X of the multi-head self-attention mechanism” and [0050]: “…finally, the pre-routing path delay and the pre-routing and post-routing path delay residual are added to obtain the final predicted post-routing path delay”), wherein the predicting comprises:
operating, by the transformer model, attention mechanism functions of the multi-head attention module[[s]] in parallel (Cao2, [0045]: “…h is the number of heads of the self-attention mechanism, and Q, K and V represent query, key and value. An attention function is performed on the h groups of matrices Q.sub.i, K.sub.i and V.sub.i parallelly, wherein a calculation formula of the dot product-based attention mechanism is formula (4)”); and
concatenating and linearly transforming outputs of the attention mechanism functions for input to a… [feed forward] network... (Cao2, [0046]: “Calculation results head.sub.i of the h-head attention mechanism are merged, and linear transform is performed by means of a trainable matrix W.sup.O to obtain an output Multillead(X) of the multi-head self-attention mechanism” and [0048]: “S44: the output, normalized in S43, of the multi-head self-attention mechanism is input to the fully connected feedforward neural network”).
Cao2 does not appear to explicitly disclose the further limitations of the claim.
However, Hakhmaneshi discloses “predicting, by the circuit representation-learning model…a voltage output of the electric circuit based on an input of circuit parameters” (Hakhmaneshi, III. C. From Node Embedding to Graph Property Prediction: “To perform a graph property prediction task, we need to combine the node embeddings into a single graph embedding…Fig. 3 demonstrates this architecture for fine tuning on a graph property prediction task” and Fig. 8. “(b) Shows the test Acc@200 (e.g., 0.5% prediction error) of predicting the output voltage for our fine-tuned method (FT-PT) versus training from scratch with no knowledge transfer” and Hakmaneshi, III. A. New Graph Representation: “Each node is associated with a feature vector which describes the type of the node as well as the device parameters that the node belongs to (e.g., the value of a resistor)”; Examiner notes that node features correspond to “circuit parameters” which are used as input to predict output voltage (see Fig. 3.)) and “…transforming outputs of the attention mechanism functions for input to a multi-layer perceptron network that maps the electric circuit to a scalar value” (Hakhmaneshi, Fig. 3: “The graph embedding is then fed to an MLP to predict the output. The node to graph embedding is three cross-attention layers between a learned embedding and the node features”; Examiner notes that output voltage is a scalar value).
Hakhmaneshi and the instant application both relate to circuit representation-learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Cao2 such that the predicted value is a voltage output of the plurality of designated components, the prediction is additionally based on an input of circuit parameters, and such that the concatenated and linearly transformed outputs are for input to “a multi-layer perceptron network that maps the electric circuit to a scalar value,” and one would have been motivated to do so, as doing so would allow for determining the optimal circuit design for achieving a desired output voltage, with improved sample efficiency (see Hakhmaneshi, IV. F. Improving Sample Efficiency of Model-Based Optimization by Using Pretrained Models).
Neither Cao2 nor Hakhmaneshi appear to explicitly disclose the further limitations of the claim.
However, Sarkar discloses “predicting, by the circuit representation model, an efficiency rating of the electric circuit…based on an input of circuit parameters” (Sarkar, [0027]: “The neural network 114 may include or define one or more predictive models 124, where each predictive model 124 corresponds to a different characteristic or performance metric (e.g., efficiency, breakdown voltage, threshold voltage, etc.). For example, a predictive model 124 relating to efficiency may predict the efficiency of the semiconductor system based on a given set of inputs” and [0028]: “The semiconductor design system 100 includes an optimizer 126 configured to operate in conjunction with the predictive model(s) 124 of the neural network 114 to generate the design model 136 in accordance with input parameters 101… For example, the optimizer 126 (in conjunction with the predictive model(s) 124) may determine the process parameters 138, circuit parameters 140 and/or device parameters 142 such that characteristics (e.g., efficiency, breakdown voltage, threshold voltage, etc.) of the predictive model(s) 124 achieve a threshold result (e.g., maximized, minimized) while meeting constraints 128 and/or goals 130 of the optimizer 126” and [0023]: “The circuit parameters 140 may include parameters for the structure (e.g., connections, wiring) of a circuit and/or parameters for circuit elements as values for resistors, capacitors, and inductors, and parameters related to the size of active semiconductor devices, etc”).
Sarkar and the instant application both relate to machine learning for circuit design and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi with Sarkar to include predicting an efficiency rating, and one would have been motivated to do so. Doing so would allow for determining the optimal circuit parameters for achieving a desired efficiency threshold (see Sarkar, [0028]).
Neither Cao2, Hakhmaneshi, nor Sarkar appear to explicitly disclose the further limitations of the claim.
However, Nath discloses “a stack of multi-head attention modules” (Nath, 2.2 Standard Transformer: “For our specific application, we propose to use the basic transformer model architecture[17], which comprises of an encoder and a decoder, each of which is a stack of 𝑁 identical blocks. The encoder block consists of two sub-layers, a multi-head self-attention layer and a position-wise feed-forward network (FFN)”).
Nath and the instant application both relate to machine learning for circuit design and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi/Sarkar with the teachings of Nath such that the transformer model includes a stack of multi-head attention modules, and one would have been motivated to do so, as doing so would increase the ability of the model to highlight the important features of the input sequence (see Nath, 3.1 Encoder design, paragraph 1).
Regarding claim 15, the rejection of claim 12 is incorporated. Hakhmaneshi further discloses “embedding the circuit parameters of the electric circuit as an input to… [a cross attention network]” (Hakhmaneshi, C. From Node Embedding to Graph Property Prediction: “Once we have the contextualized node embeddings at the output of the GNN, we flatten them as a set of feature vectors and pass them through a cross attention network”).
Hakhmaneshi and the instant application both relate to machine learning for circuit design and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Sarkar/Nath with the teachings of Hakhmaneshi to include embedding the circuit parameters of the electric circuit as an input to the transformer model, and one would have been motivated to do so. Doing so would allow for effectively capturing the structural dependencies of a circuit (see Hakhmaneshi, II. Related Work, GNNs in Chip Design).
Regarding claim 18, Cao2 discloses
“…identifying and extracting paths between a plurality of designated components that represent an electric circuit (Cao2, [0038]: “In S2, static timing analysis is performed on the circuit after placement in S1, and the timing and physical information of all the stages of cell in the path is extracted from the static timing analysis report and the layout information to form the feature sequences of the path” and [0037]: “For example, 80% of all paths extracted from six of the nine ISCAS and OpenCores circuits are randomly selected as training set data for training the model, and the remaining 20% of the paths extracted from the six circuits are used for verifying the prediction performance of the model on the known circuits, and all paths extracted from the remaining three circuits are used for verifying the prediction performance of the model on an unknown circuit”);
converting one or more of the extracted paths to respective path embeddings including a vector of a fixed length (Cao2, [0042]: “S33: Feature sequence padding is performed to ensure that the feature sequences have the same length. A maximum feature sequence length of the training set data is set as max_len, and feature sequences with a length less than max_len are filled with “0” at the end until the length of the feature sequences is max_len” and [0044]: “S41: an input feature sequence with a dimension of (samples, max_len) is converted into a tensor with a dimension of (samples,max_len,dim.sub.k), wherein samples is the number of samples, max_len is a maximum path length, dim.sub.k is a designated word vector dimension of a k.sup.th feature in an embedding layer, k=1, 2, . . . , n, and n is the number of input feature sequences; in the positional encoding process, trigonometric functions shown by formula (2) and formula (3) are used to assist the network in understanding the positional relationship of features of all the stages of cells in the path…for each feature sequence, the tensor output after positional encoding is added with a tensor output after input embedding, such that n new tensors with a dimension of (samples, max_len,dim.sub.k) are obtained” and [0045]: “S42: the n new tensors obtained in S41 are merged to obtain a tensor with a dimension of (samples,max_len,dim)”; Examiner notes that the final merged tensor in S42 resulting from embedding/encoding the feature sequences of a path corresponds to “a path embedding comprising a vector of a fixed length);
training, by a training model, a circuit representation-learning model to perform the circuit performance modeling, wherein the circuit representation-learning model comprises a transformer model including a… multi-head attention module[[s]] (Cao2, [0043]: “In S4, the transformer network comprises input embedding and positional encoding, a multi-head self-attention mechanism, a fully connected feedforward network, and adding and normalization” and [0036]: “S4: a post-routing path delay prediction model is established” and [0037]: “S5: the model established in S4 is trained and verified… During training, an Adam optimizer is used, the learning rate is 0.001, the number of training batches is 1080, and the loss function is the root-mean-square error (RMSE)”; Examiner notes that the post-routing path delay prediction model corresponds to “a circuit representation- learning model” and the Adam optimizer corresponds to “a training model”); and
predicting, by the circuit representation-learning model… [a post-routing path delay] based on an input of… the respective path embeddings of the electric circuit (Cao, [0045]: “S42: the n new tensors obtained in S41 are merged to obtain a tensor with a dimension of (samples,max_len,dim), which is used as an input X of the multi-head self-attention mechanism” and [0050]: “…finally, the pre-routing path delay and the pre-routing and post-routing path delay residual are added to obtain the final predicted post-routing path delay”), wherein the predicting comprises:
operating, by the transformer model, attention mechanism functions of the multi-head attention module[[s]] in parallel (Cao2, [0045]: “…h is the number of heads of the self-attention mechanism, and Q, K and V represent query, key and value. An attention function is performed on the h groups of matrices Q.sub.i, K.sub.i and V.sub.i parallelly, wherein a calculation formula of the dot product-based attention mechanism is formula (4)”); and
concatenating and linearly transforming outputs of the attention mechanism functions for input to a… [feed forward] network...” (Cao2, [0046]: “Calculation results head.sub.i of the h-head attention mechanism are merged, and linear transform is performed by means of a trainable matrix W.sup.O to obtain an output Multillead(X) of the multi-head self-attention mechanism” and [0048]: “S44: the output, normalized in S43, of the multi-head self-attention mechanism is input to the fully connected feedforward neural network”).
Cao2 does not appear to explicitly disclose the further limitations of the claim.
However, Hakhmaneshi discloses “predicting, by the circuit representation-learning model…a voltage output of the electric circuit based on an input of circuit parameters” (Hakhmaneshi, III. C. From Node Embedding to Graph Property Prediction: “To perform a graph property prediction task, we need to combine the node embeddings into a single graph embedding…Fig. 3 demonstrates this architecture for fine tuning on a graph property prediction task” and Fig. 8. “(b) Shows the test Acc@200 (e.g., 0.5% prediction error) of predicting the output voltage for our fine-tuned method (FT-PT) versus training from scratch with no knowledge transfer” and Hakmaneshi, III. A. New Graph Representation: “Each node is associated with a feature vector which describes the type of the node as well as the device parameters that the node belongs to (e.g., the value of a resistor)”; Examiner notes that node features correspond to “circuit parameters” which are used as input to predict output voltage (see Fig. 3.)) and “…transforming outputs of the attention mechanism functions for input to a multi-layer perceptron network that maps the electric circuit to a scalar value” (Hakhmaneshi, Fig. 3: “The graph embedding is then fed to an MLP to predict the output. The node to graph embedding is three cross-attention layers between a learned embedding and the node features”; Examiner notes that output voltage is a scalar value).
Hakhmaneshi and the instant application both relate to circuit representation-learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Cao2 such that the predicted value is a voltage output of the plurality of designated components, the prediction is additionally based on an input of circuit parameters, and such that the concatenated and linearly transformed outputs are for input to “a multi-layer perceptron network that maps the electric circuit to a scalar value,” and one would have been motivated to do so, as doing so would allow for determining the optimal circuit design for achieving a desired output voltage, with improved sample efficiency (see Hakhmaneshi, IV. F. Improving Sample Efficiency of Model-Based Optimization by Using Pretrained Models).
Neither Cao2 nor Hakhmaneshi appear to explicitly disclose the further limitations of the claim.
However, Sarkar discloses “A computer program product, comprising: one or more computer-readable storage media; and program instructions stored on at least one of the one or more computer-readable storage media to perform operations” (Sarkar, [0005]: “According to an aspect, a non-transitory computer-readable medium storing executable instructions that when executed by at least one processor is configured to cause the at least one processor to…”) and “predicting, by the circuit representation model, an efficiency rating of the electric circuit…based on an input of circuit parameters” (Sarkar, [0027]: “The neural network 114 may include or define one or more predictive models 124, where each predictive model 124 corresponds to a different characteristic or performance metric (e.g., efficiency, breakdown voltage, threshold voltage, etc.). For example, a predictive model 124 relating to efficiency may predict the efficiency of the semiconductor system based on a given set of inputs” and [0028]: “The semiconductor design system 100 includes an optimizer 126 configured to operate in conjunction with the predictive model(s) 124 of the neural network 114 to generate the design model 136 in accordance with input parameters 101… For example, the optimizer 126 (in conjunction with the predictive model(s) 124) may determine the process parameters 138, circuit parameters 140 and/or device parameters 142 such that characteristics (e.g., efficiency, breakdown voltage, threshold voltage, etc.) of the predictive model(s) 124 achieve a threshold result (e.g., maximized, minimized) while meeting constraints 128 and/or goals 130 of the optimizer 126” and [0023]: “The circuit parameters 140 may include parameters for the structure (e.g., connections, wiring) of a circuit and/or parameters for circuit elements as values for resistors, capacitors, and inductors, and parameters related to the size of active semiconductor devices, etc”).
Sarkar and the instant application both relate to machine learning for circuit design and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi with Sarkar to include a computer program product and to include predicting an efficiency rating, and one would have been motivated to do so. Doing so would allow for determining the optimal circuit parameters for achieving a desired efficiency threshold (see Sarkar, [0028]).
Neither Cao2, Hakhmaneshi, nor Sarkar appear to explicitly disclose the further limitations of the claim.
However, Nath discloses “a stack of multi-head attention modules” (Nath, 2.2 Standard Transformer: “For our specific application, we propose to use the basic transformer model architecture[17], which comprises of an encoder and a decoder, each of which is a stack of 𝑁 identical blocks. The encoder block consists of two sub-layers, a multi-head self-attention layer and a position-wise feed-forward network (FFN)”).
Nath and the instant application both relate to machine learning for circuit design and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi/Sarkar with the teachings of Nath such that the transformer model includes a stack of multi-head attention modules, and one would have been motivated to do so, as doing so would increase the ability of the model to highlight the important features of the input sequence (see Nath, 3.1 Encoder design, paragraph 1).
Claims 6 and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Cao2 in view of Hakhmaneshi, Sarkar, and Nath, and further in view of Cao et al. (US20230195986) ("Caol").
Regarding claim 6, the rejection of claim 1 is incorporated. Neither Cao2, Hakhmaneshi, Sarkar, nor Nath appear to explicitly disclose the further limitations of the claim.
However, Cao1 discloses “wherein the converting of the at least one extracted path to the path embedding is performed by a bidirectional Long Short-Term Memory (Bi-LSTM) network, and wherein the instructions cause the processor to perform an additional act comprising outputting, by the Bi-LSTM network, the vector of the fixed length for each path embedding” (Cao1, [0018]: “In step S3, the topology information of the path includes the cell type, the cell size, and the corresponding load capacitance sequence, for two category-type sequence features of the cell type and the cell size, the problem of inconsistent sequence lengths of the two sequence features is first resolved through padding, then padded sequences are inputted into an embedding layer, a vector representation of an element is obtained through network learning, load capacitances are binned and filled by using a padding operation to a uniform length, and vector expressions are learned by using the embedding layer, next, vector splicing is performed on the foregoing expressions obtained after the learning using the embedding layer, and finally spliced vectors are inputted into the bi-directional long short-term memory neural network (BLSTM) to perform training” and [0021]: “ finally, a pooling layer is connected after the bi-directional long short-term memory neural network (BLSTM), dimension reduction is performed on the second dimension outputted by the sequence, a dimension of an outputted vector is N.sub.s×h.sub.o”).
Caol and the instant application both relate to predicting circuit performance using neural
networks and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the
effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi/Sarkar/Nath with the teachings of Cao1 such that the converting of the at least one extracted path to the path embedding is performed by a bidirectional Long Short-Term Memory (Bi-LSTM) network, and wherein the instructions cause the processor to perform an additional act comprising outputting, by the Bi-LSTM network, the vector of the fixed length for each path embedding, and one would have been motivated to do so, as doing so would achieve prediction with higher precision through more effective feature engineering processing in a case of low simulation overheads (see Caol, Abstract).
Regarding claim 7, the rejection of claim 6 is incorporated. Neither Cao2, Hakhmaneshi, Sarkar, nor Nath appear to explicitly disclose the further limitations of the claim.
However, Caol discloses "representing the electric circuit as a device embedding input to the Bi-LSTM network" (Caol, [0018]: "In step S3, the topology information of the path includes the cell type, the cell size, and the corresponding load capacitance sequence, for two category-type sequence features of the cell type and the cell size, the problem of inconsistent sequence lengths of the two sequence features is first resolved through padding, then padded sequences are inputted into an embedding layer, a vector representation of an element is obtained through network learning, load capacitances are binned and filled by using a padding operation to a uniform length, and vector expressions are learned by using the embedding layer, next, vector splicing is performed on the foregoing expressions obtained after the learning using the embedding layer, and finally spliced vectors are inputted into the bi-directional long short-term memory neural network (BLSTM) to perform training"; Examiner notes that the vector expressions learned by using the embedding layer correspond to device embeddings because they are embeddings that represent information of cells on a path in a circuit).
Cao1 and the instant application both relate to predicting circuit performance using neural
networks and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the
effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi/Sarkar/Nath with the teachings of Cao1 to include representing the electric circuit as a device embedding input to the Bi-LSTM network, as disclosed by Caol, and one would have been motivated to do so for the purpose of achieving prediction with higher precision through more effective feature engineering processing in a case of low simulation overheads (see Caol, Abstract).
Claims 9-10 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Cao2 in view of Hakhmaneshi, Sarkar, and Nath, and further in view of Zhang et al. (“Student Network Learning via Evolutionary Knowledge Distillation”) ("Zhang").
Regarding claim 9, the rejection of claim 1 is incorporated. Neither Cao2, Hakhmaneshi, Sarkar, nor Nath appear to explicitly disclose the further limitations of the claim.
However, Zhang discloses "wherein the training model comprises a stochastic gradient
descent-based model" (Zhang, III, Algorithm 1, line 8: "compute the total loss of teacher with Eq.
(11) and line 9: "Compute gradient to model parameters Wt and update with the SGD [stochastic
gradient descent] optimizer").
Zhang and the instant application both relate to neural networks and are analogous. It would have
been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi/Sarkar/Nath to have the training be performed by a stochastic gradient descent-based training model, as disclosed by Zhang, and one would have been motivated to do so, as doing so would achieve more robust performance and improved generalization ability of the student model (see Zhang, II.A).
Regarding claim 10, the rejection of claim 9 is incorporated. Cao2 further discloses
"processing, by… [the training model] parameters in the path embedding, the transformer model" (Cao2, [0037]: "S5: the model established in S4 is trained and verified…During training, an Adam optimizer is used, the learning rate is 0.001, the number of training batches is 1080, and the loss function is the root-mean-square error (RMSE)"), but does not appear to explicitly disclose the further limitations of the claim.
However, Hakhmaneshi discloses “processing by… [the training model] parameters in…the multi-layer perceptron network” (Hakmaneshi, IV. E. Domain Knowledge Transfer by Reusing the Learned GNN Features for Predicting Graph Level Tasks: “In this section, we use the architecture outlined in Section III-C to reuse the learned node features of the pre trained GNN by using a cross-attention pooling layer to construct a graph feature from aggregating node features and use that for the prediction of graph-level metrics. To learn the model, we update all parameters (including the GNN back bone parameters) using Adam optimizer on the downstream training dataset”).
Hakhmaneshi and the instant application both relate to machine learning for circuit design and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cao2/Sarkar/Nath/Zhang with the teachings of Hakhmaneshi to include processing parameters in the multi-layer perceptron network, and one would have been motivated to do so, as doing so would improve the accuracy of the circuit property prediction task (see Hakhmaneshi, Fig. 8).
Neither Cao2, Hakhmaneshi, Sarkar, nor Nath appear to explicitly disclose the further limitations of the claim.
However, Zhang discloses "processing, by the stochastic gradient descent-based model,
parameters...” (Zhang, III, Algorithm 1, line 8: "compute the total loss of teacher with Eq. (11) and
line 9: "Compute gradient to model parameters Wt and update with the SGD [stochastic gradient
descent] optimizer").
Zhang and the instant application both relate to neural networks and are analogous. It would have
been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention,
to have modified the combination of Cao2/Hakhmaneshi/Sarkar/Nath with the teachings of Zhang such that the parameters are processed by the stochastic gradient descent model, and one would have been motivated to do so for the purpose of achieving more robust performance and improved generalization ability of the student model (see Zhang, II.A).
Regarding claim 17, the rejection of claim 12 is incorporated. Cao2 further discloses “…the training model…process[es] parameters for the respective path embeddings” (Cao2, [0037]: “S5: the model established in S4 is trained and verified… During training, an Adam optimizer is used, the learning rate is 0.001, the number of training batches is 1080, and the loss function is the root-mean-square error (RMSE)”).
Neither Cao2, Hakmaneshi, Sakar, nor Nath appear to explicitly disclose “training the training model.”
However, Zhang discloses "training the training model to process parameters" (Zhang, III,
Algorithm 1, line 8: "compute the total loss of teacher with Eq. (11) and line 9: "Compute gradient to model parameters Wt and update with the SGD [stochastic gradient descent] optimizer"; Examiner notes that the teacher corresponds to a “training model” as it supervises the learning of a student model, see Zhang, III: “In our evolutionary knowledge distillation (EKD) approach, the teacher and student network are trained almost synchronously, and an evolutionary teacher can provide supervision information for the learning of student”).
Zhang and the instant application both relate to neural networks and are analogous. It would have
been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention,
to have modified the combination of Cao2/Hakmaneshi/Sakar/Nath with the teachings of Zhang to include "training the training model to process parameters for the respective path embeddings" and one would have been motivated to do so for the purpose of achieving more robust performance and improved generalization ability of the student model (see Zhang, II.A).
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Cao2 in view of Hakhmaneshi, Sarkar, Nath, Zhang, and further in view of An et al. (“LPViT: A Transformer Based Model for PCB Image Classification and Defect Detection”) (“An”).
Regarding claim 11, the rejection of claim 10 is incorporated. Cao2, Hakhmaneshi, Sarkar, Nath, and Zhang do not appear to explicitly disclose the further limitations of the claim.
However, An discloses "the multi-layer perceptron network has an input size that is the same as
an output size of the transformer model" (An, III.B: "3) TRANSFORMER The transformer module is used to extract features, which are then fed into the final classifier to produce the required
classification results... 4) CLASSIFIER After enough good features have been retrieved, all that is
required for reliable classification results is a simple network topology, and a shallow network like
MLP(multilayer perceptron) can achieve incredibly high metrics" and Figure 3. ViT architecture:
Transformer Encoder to MLP Head); Examiner notes that the input size of the multi-layer
perceptron network must be the same as the output size of the transformer model given that the output
features of the transformer module are input to the multi-layer perceptron network classifier).
An and the instant application both relate to circuit-performance prediction using a transformer
network and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the
effective filing date of the claimed invention, to have modified the combination of Cao2/Hakhmaneshi/Sarkar/Nath/Zhang with the teachings of An such that the multi-layer perceptron network has an input size that is the same as an output size of the transformer model, and one would have been motivated to do, as doing so would achieve increase performance metrics (see An, III.B.4).
Response to Arguments
Applicant's arguments filed 6/9/2026 regarding the rejections under 35 U.S.C. 101 have been fully considered but they are not persuasive.
Applicant argues regarding Step 2A Prong One that the claimed features are “inextricably tied to a machine and a computer technology” and thus “cannot be considered as an abstract idea such as mental processes.” However, Applicant provides no evidence as to why the limitations analyzed as being mental processes in the 101 rejections are not mentally performable beyond the assertion that they are tied to a machine, thus Examiner does not find this argument persuasive.
Applicant further argues regarding Step 2A Prong Two that the alleged abstract idea is integrated into a practical implementation. Applicant asserts that the use of a transformer model “provides accurate predictions about circuit performance with less computation time than conventional software simulators” and operating multiple attention mechanisms in parallel leads to “faster and more accurate circuit representation and prediction.” Examiner submits that merely using a generic transformer model to perform the abstract idea amounts to mere instructions to apply an exception on a generic computer programmed with a generic class of computer algorithms (2106.05(f)) and therefore does not integrate the judicial exception into a practical application. Applicant further argues that embedding all paths into fixed-length vectors “results in a more stable and scalable computational framework.” Examiner submits that this amounts to an assertion that the judicial exception itself provides the purported improvement, and the judicial exception alone cannot provide the improvement (see MPEP 2106.05(a)).
Applicant’s arguments regarding the rejections under 35 U.S.C. 103 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GWYNEVERE A DETERDING whose telephone number is (571)272-7657. The examiner can normally be reached Mon-Fri. 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/G.A.D./Examiner, Art Unit 2125
/KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125