DETAILED ACTION
This action is responsive to the application filed on 05/25/2026. Claims 1,2,4-12, and 14-20 are pending and have been examined. This action is Final.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C.
120, 121, 365(c), or 386(c) is acknowledged.
Response to Arguments
Argument 1: The applicant argues that amended independent claims 1 and 11 are not anticipated by Kuo because the claims now require a specific paired relationship for each neural network operation: the operation itself is mapped to a categorical feature vector, while that same operation's configuration setting is mapped to a corresponding numerical feature vector. According to the applicant, Kuo merely combines unrelated categorical and numerical property attributes, such as occupancy type and insurance coverage, and therefore does not disclose operation and configuration-setting pairs. The applicant also argues that Kuo applies positional encoding and transformer processing only to the categorical embeddings before the numerical variables are concatenated, rather than applying positional encoding to a feature-embedded sequence containing both categorical and numerical feature vectors as required. Because claims 2 and 12 depend from claims 1 and 11, the applicant requests withdrawal of the 35 U.S.C. 102 rejection of claims 1, 2, 11, and 12. For the 35 U.S.C. 103 rejections, the applicant contends that Gao does not cure these deficiencies for claims 4 through 7 and 14 through 17, and Vaswani does not cure them for claims 8 through 10 and 18 through 20. The applicant therefore requests withdrawal of all pending rejections and allowance of the claims. Claims 3 and 13 were canceled.
Response to Argument 1: The applicant’s arguments have been fully considered but are not persuasive because they address the former anticipation rejection based principally on Kuo rather than the present obviousness rejection based on the combined teachings of Chau, Deng, Zhou, and Lu in light of the amendments. The present rejection does not rely on Kuo to disclose the operation-specific paired relationship or the application of positional encoding to the claimed combined feature representation. Instead, Chau teaches training a predictor to predict execution performance of neural-network models on a specified hardware platform; Deng expressly teaches representing each individual operation using a tuple comprising the operation type, TY, and that same operation’s numerical configuration settings, KW, KH, and CH, retrieving corresponding vectors from type, kernel-size, and channel-ratio lookup tables, and concatenating those vectors to form the representation of that particular layer, thereby teaching the claimed categorical vector and corresponding numerical vectors for the same operation; Zhou teaches training the predictor using performance values obtained from an accelerator simulator; and Lu teaches applying positional information to an embedded operation representation before processing the resulting representation through multiple Transformer encoder layers having multi-head self-attention and a regressor. Thus, when Lu’s positional encoding and Transformer processing are applied to Deng’s feature-embedded operation representation, the positional encoding is applied to the representation containing both the categorical operation vector and that operation’s corresponding numerical configuration vectors, rather than only to unrelated categorical data. Kuo is relied upon only for the additional concatenation limitation of dependent claims 2 and 12, while Gao and Vaswani are relied upon only for the additional limitations of their respective dependent claims, not to remedy the amended independent-claim limitations. Accordingly, the assertion that Kuo, Gao, or Vaswani individually fails to cure the independent-claim deficiencies does not overcome the rejection based on the combined teachings and articulated reasons to combine, and the rejections of claims 1, 2, 4-12, and 14-20 under 35 U.S.C. 103 are maintained.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this
Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not
identically disclosed as set forth in section 102, if the differences between the claimed invention and the
prior art are such that the claimed invention as a whole would have been obvious before the effective filing
date of the claimed invention to a person having ordinary skill in the art to which the claimed invention
pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are
summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 4, 11, and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over WO2022071685A1, by Chau. et. al. (referred herein as Chau) in view of NPL reference “Peephole: Predicting Network Performance Before Training” in view of NPL reference “Rethinking Co-design of Neural Architectures and Hardware Accelerators”, by Zhou et. al. (referred herein as Zhou) further in view of NPL reference “TNASP: ATransformer-based NAS Predictor with a Self-evolution Framework”, by Lu et. al. (referred herein as Lu).
Regarding claim 1, Chau teaches:
A method of training a prediction engine for predicting performance of a neural network model executed on a hardware platform, comprising: ([Chau, [0047]] “a computer implemented method uses a trained predictor for predicting performance of a neural network on a hardware arrangement” AND [Chau, 0145] “The predictor may be trained with measurements of the performance of neural network models on hardware arrangements”, wherein the examiner interprets “trained predictor” to be the same as “prediction engine” because they are both directed to a trained machine-learning model that generates predicted performance information, and interprets “hardware arrangement” to be the same as “hardware platform” because they are both directed to the physical computing environment on which the neural network model is implemented and executed).
receiving, by the prediction engine during training, a plurality of training neural networks compiled for the hardware platform, ([Chau, [0141] “each of the predictors needs to be trained and Figure 10 shows a method of training a machine learning predictor, according to an embodiment. Operation S1000 may include generating a plurality of inputs.” AND [Chau, [0146] “the predictor may be trained using a randomly sampled set of 900 models from a database, such as the NAS-Bench-201 dataset” AND [Chau, 0109] “it is necessary to input the hardware arrangement on which each of the neural network models are to be implemented”, wherein the examiner interprets “generating a plurality of inputs” using a “randomly sampled set of 900 models” to be the same as receiving, during training, a plurality of training neural networks because they are both directed to supplying multiple neural-network models as training inputs to a performance predictor, and interprets neural-network models “to be implemented” on an identified hardware arrangement to be the same as neural networks compiled for the hardware platform because they are both directed to neural-network models prepared or configured for execution on specified target hardware).
each training neural network including a plurality of layers and each layer defined by a set of operations ([Chau, 0106] “the nodes N1 to N8 each represent an operation of a layer within the model and the computational flow is represented by an edge”, wherein the examiner interprets nodes respectively representing operations of layers within a model to be the same as each training neural network including a plurality of layers defined by a set of operations because they are both directed to representing a neural-network model as multiple layers performing respective computational operations).
generating, by the predicting engine, a performance metric of executing the given training neural network on the hardware platform; and ([Chau, 0110] “Operation S410 may include predicting the performance of each input neural network model on the fixed hardware design…the performance of each model on the hardware may also be output. Typically, overall latency may be predicted together with the latency associated with each operation.”, wherein the examiner interprets predicting the performance and latency of an input neural-network model “on the fixed hardware design” to be the same as generating a performance metric of executing the given training neural network on the hardware platform because they are both directed to producing a quantitative performance value for a neural-network model executed on specified hardware).
Chau does not teach:
configuration settings corresponding to the set of operations, wherein the prediction engine is a transformer-based neural network; generating a feature embedded sequence for a given training network in the plurality of neural networks, the feature embedded sequence including both a sequence of categorical feature vectors and a sequence of numerical feature vectors, wherein for each operation of the given training neural network, the operation is mapped to a respective one of the categorical feature vectors and the operation's configuration setting is mapped to a respective one of the numerical feature vectors such that each categorical feature vector has a corresponding numerical feature vector; updating a categorical mapping and a numerical mapping of the prediction engine based on a difference between the performance metric and a simulated performance metric obtained from the given training neural network; wherein generating the performance metric for each training neural network further comprises: applying positional encoding to the feature embedded sequence including both the sequence of categorical feature vectors and the sequence of numerical feature vectors, followed by applying a series of attention functions to an output of the positional encoding, to generate an encoded sequence; and reducing dimensions of the encoded sequence to output the performance metric of executing the given training neural network on the hardware platform.
Deng teaches:
configuration settings corresponding to the set of operations, ([Deng, page 3, sec. 3.1] “For each layer, we encode it with a type id (TY), a kernel width (KW), a kernel height (KH), and a channel number (CH)”, wherein the examiner interprets “a kernel width (KW), a kernel height (KH), and a channel number (CH)” to be the same as configuration settings corresponding to the set of operations because they are both directed to numerical settings that configure the operation performed by the respective neural-network layer).
generating a feature embedded sequence for a given training network in the plurality of neural networks, ([Deng, page 3, Figure 2] “Given a network architecture, it first encodes each layer into a vector through integer coding and layer embedding. Subsequently, it applies a recurrent network with LSTM units to integrate the information of individual layers following the network topology into a structural feature” AND [Deng, page 4, sec. 3.2] “Specifically, we adopt the Long-Short Term Memory (LSTM), an effective variant of RNN, for integrating the information along a sequence of layers.”, wherein the examiner interprets encoding each layer of a network architecture into a vector and integrating the information “along a sequence of layers” to be the same as generating a feature embedded sequence for a given training network because they are both directed to generating an ordered sequence of embedded layer representations for an individual neural-network architecture).
the feature embedded sequence including both a sequence of categorical feature vectors and a sequence of numerical feature vectors, wherein for each operation of the given training neural network, the operation is mapped to a respective one of the categorical feature vectors and the operation's configuration setting is mapped to a respective one of the numerical feature vectors such that each categorical feature vector has a corresponding numerical feature vector; ([Deng, page 4, sec. 3.1] “Overall, we can represent a common operation by a tuple of four integers in the form of (TY, KW, KH, CH), where TY is an integer id that indicates the type of the computation, KW and KH are respectively the width and height of the kernel, while CH represents the ratio of output-input channels” and “this module is associated with three lookup tables, respectively for layer types, kernel sizes, and channel ratios. Note that the kernel size table is used to encode both KW and KH. Given a tuple of integers, we can convert its element into a real vector by retrieving from the corresponding lookup table. Then by concatenating all the embedded vectors derived respectively from individual integers, we can form a vector representation of the layer”, wherein the examiner interprets the real vector retrieved for TY, which “indicates the type of the computation,” to be the same as a categorical feature vector because they are both directed to representing the category of an operation as a vector. Further, the examiner interprets the real vectors retrieved for KW, KH, and CH to be the same as numerical feature vectors because they are both directed to vector representations derived from numerical configuration settings of the operation. Lastly, the examiner interprets concatenating the type, kernel-width, kernel-height, and channel vectors to form the representation of each layer to be the same as each categorical feature vector having a corresponding numerical feature vector because they are both directed to associating the categorical operation representation with the numerical configuration representations belonging to the same operation).
updating a categorical mapping and a numerical mapping of the prediction engine ([Deng, page 4, sec. 3.1] “this module is associated with three lookup tables, respectively for layer types, kernel sizes, and channel ratios “AND [Deng, page 3, Figure 2] “the embeddings, the LSTM, and the MLP, are jointly learned in an end-to-end manner” AND [Deng, page 4, sec 3.1] “At each step, it takes an input xt, decides the value of all the gates, yields an output ut, and updates both the hidden state ht and the cell memory”, wherein the examiner interprets the lookup table for “layer types” to be the same as a categorical mapping because they are both directed to mapping categorical operation types to vector representations; interprets the lookup tables for “kernel sizes, and channel ratios” to be the same as a numerical mapping because they are both directed to mapping numerical operation configurations to vector representations; and interprets the embeddings being “jointly learned in an end-to-end manner” to be the same as updating the categorical mapping and the numerical mapping because they are both directed to modifying the learned mapping parameters during training).
Chau and Deng do not teach:
based on a difference between the performance metric and a simulated performance metric obtained from the given training neural network; wherein the prediction engine is a transformer-based neural network; wherein generating the performance metric for each training neural network further comprises: applying positional encoding to the feature embedded sequence including both the sequence of categorical feature vectors and the sequence of numerical feature vectors, followed by applying a series of attention functions to an output of the positional encoding, to generate an encoded sequence; and reducing dimensions of the encoded sequence to output the performance metric of executing the given training neural network on the hardware platform.
Zhou teaches:
based on a difference between the performance metric and a simulated performance metric obtained from the given training neural network; ([Zhou, page 6, sec. 3.5.2] “we train a cost model with random generated samples using an in-house accelerator simulator” AND [Zhou, page 5, sec 3.3] “The target device is an industry-standard, highly parameterized edge accelerator which allows us to create various configurations in a large design space with tradeoffs between performance, power, area, and cost…Unlike the NAS search space, the HAS search space contains many invalid points, which makes training a cost model or joint search with the in-house simulator more challenging”, wherein the examiner interprets “in house accelerator simulator…which allows us to create various configurations in a large design space with tradeoffs between performance, power, area, and cost” to be the same as basing difference between performance/simulated performance based on a trained NN, because they are both directed to determining prediction error between predicted hardware performance and simulator-generated hardware performance).
Chau, Deng, and Zhou do not teach:
wherein the prediction engine is a transformer-based neural network; wherein generating the performance metric for each training neural network further comprises: applying positional encoding to the feature embedded sequence including both the sequence of categorical feature vectors and the sequence of numerical feature vectors, followed by applying a series of attention functions to an output of the positional encoding, to generate an encoded sequence; and reducing dimensions of the encoded sequence to output the performance metric of executing the given training neural network on the hardware platform.
Lu teaches:
wherein the prediction engine is a transformer-based neural network; ([Lu, page 4, Fig. 1] “Our Transformer-based NAS predictor mainly consists of an encoder and a regressor. We first encode the information of operations and connections into continuous representation, followed by 3 Transformer encoder layers, and the regressor uses the output feature of Transformer encoder layers to derive the final prediction”, wherein the examiner interprets “Transformer-based NAS predictor” to be the same as a prediction engine that is a transformer-based neural network because they are both directed to a neural-network performance predictor that uses Transformer encoder layers as its prediction architecture).
wherein generating the performance metric for each training neural network further comprises: applying positional encoding to the feature embedded sequence including both the sequence of categorical feature vectors and the sequence of numerical feature vectors, as to applying positional encoding to the feature embedded sequence recited above ([Lu, page 4, sec. 3.2] “we first get the operation feature [eq] by transforming the operation vector κ with an embedding matrix E ∈ RF×M…Then we use a linear layer to map the Laplacian matrix (L) to a continuous feature vector e2, whose dimension is same as the embedding vector e1,”, wherein the examiner interprets the embedded operation feature e1 to be the same as the feature embedded sequence because they are both directed to a continuous embedded representation of the operations of the neural-network architecture, and interprets adding positional feature e2 to embedded operation feature e1 before Transformer processing to be the same as applying positional encoding to the feature embedded sequence because they are both directed to combining positional information with an embedded architecture representation before attention processing).
followed by applying a series of attention functions to an output of the positional encoding, to generate an encoded sequence; and ([Lu, page 4, Fig. 1] “We first encode the information of operations and connections into continuous representation, followed by 3 Transformer encoder layers” AND [Lu, page 4, sec. 3.2] “Finally, we get the the continuous representation from the Transformer encoder with multi-head self-attention module”, wherein the examiner interprets “3 Transformer encoder layers” having a “multi-head self-attention module” to be the same as a series of attention functions because they are both directed to multiple successive Transformer attention-processing layers, and interprets the resulting “continuous representation” to be the same as an encoded sequence because they are both directed to an encoded representation generated by applying Transformer attention to the combined operational and positional representations).
reducing dimensions of the encoded sequence to output the performance metric of executing the given training neural network on the hardware platform. ([Lu, page 4, Fig. 1] “the regressor uses the output feature of Transformer encoder layers to derive the final prediction” AND [Lu, page 5, sec 3.2] “we only choose a simple regressor, specifically 2 Multi-Layer Perceptions (MLP), to estimate the final accuracy”, wherein the examiner interprets using an MLP regressor to convert “the output feature of Transformer encoder layers” into “the final prediction” to be the same as reducing dimensions of the encoded sequence because they are both directed to converting a multidimensional Transformer-encoded representation into a lower-dimensional final performance prediction).
Chau, Deng, Zhou, Lu, and the instant application are analogous art because they are all directed to predicting neural-network performance from encoded representations of neural-network architectures and their operational characteristics.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the operation-based feature representation disclosed by Chau to include the layer-type, kernel-width, kernel-height, and channel representations disclosed by Deng. One would have been motivated to do so to provide the performance predictor with a uniform representation containing both the identity and numerical configuration of each neural-network operation, as suggested by Deng ([Deng, page 4, sec. 3.1] “Then by concatenating all the embedded vectors derived respectively from individual integers, we can form a vector representation of the layer”).
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the simulator-derived performance labels and mean-squared-error training objective disclosed by Zhou. One would have been motivated to do so to train the prediction engine to reproduce platform-specific execution performance without requiring a new cycle-accurate simulator query each time a neural-network configuration is evaluated as suggested by Zhou ([Zhou, page 6, sec. 3.5.2] “we train a cost model with random generated samples using an in-house accelerator simulator” AND [Zhou, page 5, sec 3.3] “The target device is an industry-standard, highly parameterized edge accelerator which allows us to create various configurations in a large design space with tradeoffs between performance, power, area, and cost…Unlike the NAS search space, the HAS search space contains many invalid points, which makes training a cost model or joint search with the in-house simulator more challenging”)
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the Transformer encoder, positional encoding, and multi-head self-attention disclosed by Lu. One would have been motivated to do so to efficiently transforming neural-network operations and connections into continuous representations, incorporating graph-topology positional information, processing the resulting representation through multiple Transformer encoder layers, and using a regressor to derive the final performance prediction, as suggested by Lu ([Lu, page 4, Fig. 1] “Our Transformer-based NAS predictor mainly consists of an encoder and a regressor. We first encode the information of operations and connections into continuous representation, followed by 3 Transformer encoder layers, and the regressor uses the output feature of Transformer encoder layers to derive the final prediction” AND [Lu, page 4, sec. 3.2] “we first get the operation feature [eq] by transforming the operation vector κ with an embedding matrix E ∈ RF×M…Then we use a linear layer to map the the Laplacian matrix (L) to a continuous feature vector e2, whose dimension is same as the embedding vector e1,”). Claim 11 is analogous to claim 1, aside from claim type and minute differences, thus the same mapping applies as above.
Regarding claim 4, Chau, Deng, Zhou, and Lu teaches The method of claim 1 (see rejection of claim 1).
Deng further teaches wherein the set of operations are categorized into a set of operation groups, the operation groups including one or more of convolutions, pooling, and activation functions. ([Deng, page 3, sec. 3.1] “Table 1: The coding table of the Unified Layer Code, where each row corresponds to a layer type” and “We notice that the operations commonly used in a CNN, including convolution, pooling, and nonlinear activation, can all be considered as applying a kernel to the input feature map. To produce an output value, the kernel takes a local part of the feature map as input, applies a linear or nonlinear transform, and then yields an output. In particular, an element-wise activation function can be considered as a nonlinear kernel of size 1 x 1”, wherein the examiner interprets “each row corresponds to a layer type” to be the same as the set of operations are categorized into a set of operation groups because they are both directed to assigning neural-network operations to respective identified categories based on the type of operation performed, and interprets “nonlinear activation” and “element-wise activation function” to be the same as activation functions because they are both directed to functions that apply a nonlinear transformation to an input feature map).
Chau, Deng, Zhou, Lu, and the instant application are analogous art because they are all directed to categorizing and representing neural-network operations, including convolution, pooling, and activation operations, for use in predicting neural-network performance.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Chau, Deng, Zhou, and Lu to include the “operations commonly used in a CNN, including convolution, pooling, and nonlinear activation” disclosed by Deng. One would be motivated to do so to efficiently standardize and distinguish the different types of neural-network operations within the embedded architecture representation for improved performance prediction, as suggested by Deng ([Deng, page 3, sec. 3.1] “we propose Unified Layer Code (ULC), a uniform scheme to encode various layers into numerical vectors”). Claim 14 is analogous to claim 4, aside from claim type and minute differences, thus the same mapping applies as above.
Claim(s) 2 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chau in view of Deng in view of Zhou in view Lu further in view of NPL reference “Embeddings and Attention in Predictive Modeling”, by Kuo et. al. (referred herein as Kuo).
Regarding claim 2, Chau, Deng, Zhou, and Lu teaches The method of claim 1, (see rejection of claim 1).
Chau, Deng, Zhou, and Lu do not teach wherein performing feature embedding further comprises: concatenating the sequence of the categorical feature vectors for all layers of the neural network model and the sequence of the numerical feature vectors to generate the feature embedded sequence; ([Kuo, page 9, sec. 4] “The categorical inputs go through one-dimensional embedding layers ... The embeddings are then concatenated with the numeric predictors, which have been normalized in data pre-processing, before being passed through a feedforward layer (with 8 hidden units and ReLU activation) to obtain a scalar value between 0 and 1, as constrained by a sigmoid output activation,” wherein the examiner interprets categorical “embeddings” to be the same as “categorical feature vectors” as they both involve mapping categorical data into vectors for use in downstream processing in the context of machine learning; the examiner also interprets “categorical inputs go through one-dimensional embedding layers...The embeddings are then concatenated with the numeric predictors” to be the same as “concatenating the sequence of the categorical feature vectors for all layers of the neural network model and the sequence of the numerical feature vectors to generate the feature embedded sequence”. This is further illustrated in the figure below from Kuo where “Concat” is interpreted to be the same as concatenation.).
Chau, Deng, Zhou, Lu, Kuo, and the instant application are analogous art because they are all directed to generating feature representations containing categorical and numerical information for downstream neural-network prediction and performance analysis.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Chau, Deng, Zhou, and Lu to include the process in which “The embeddings are then concatenated with the numeric predictors” disclosed by Kuo. One would be motivated to do so to effectively generate a unified feature-embedded sequence containing both categorical feature vectors and numerical feature vectors for downstream prediction processing, as suggested by Kuo ([Kuo, page 9, sec. 4] “before being passed through a feedforward layer (with 8 hidden units and ReLU activation) to obtain a scalar value between 0 and 1”). Claim 12 is analogous to claim 2, aside from claim type and minute differences, thus the same mapping applies as above.
Claim(s) 5-7, and 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chau in view of Deng in view of Zhou in view Lu in view of NPL reference “Resource-Guided Configuration Space Reduction for Deep Learning Models”, by Gao et. al. (referred herein as Gao).
Regarding claim 5, Chau, Deng, Zhou, and Lu teach The method of claim 1 (see rejection of claim 1).
Chau, Deng, Zhou, and Lu do not teach further comprising: training the feature embedding to map each operation to a categorical feature vector that has a trainable vector value and a predetermined embedding size.
Gao teaches further comprising: training the feature embedding to map each operation to a categorical feature vector that has a trainable vector value and a predetermined embedding size. ([Gao, pages 1-2] “Another useful constraint is that the size of a model’s weights cannot exceed a certain upper bound. ... The inputs and outputs of such a computation graph and its nodes are tensors (multi-dimensional arrays of numerical values). The shape of a tensor is the element number in each dimension plus the element data type. Each node represents the invocation of a mathematical operation called an operator (e.g., elementwise matrix addition). An edge delivers an output tensor and specifies the execution dependency,” and [Gao, page 4] “execution of a DL model can be represented as iterative forward and backward propagation on its computation graph,” wherein the examiner interprets “Each node represents the invocation of a mathematical operation called an operator” to be the same as map each operation to a categorical feature vector, “iterative forward and backward propagation on such a computation graph” to be the same as training the feature embedding, and “size of a model’s weights cannot exceed a certain upper bound” to be the same as a predetermined embedding size).
Chau, Deng, Zhou, Lu, Gao, and the instant application are analogous art because they are all directed to training feature embeddings to improve the representation of operations in a neural network model for performance prediction.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Chau, Deng, Zhou, and Lu to include the “execution of a DL model can be represented as iterative forward and backward propagation on its computation graph” disclosed by Gao. One would be motivated to do so to effectively improve the learning of trainable feature embeddings, as suggested by Gao ([Gao, page 4] “Each node represents the invocation of a mathematical operation called an operator”). Claim 15 is analogous to claim 5, aside from claim type and minute differences, thus the same mapping applies as above.
Regarding claim 6, Chau, Deng, Zhou, and Lu teach The method of claim 1 (see rejection of claim 1).
Chau, Deng, Zhou, and Lu do not teach wherein one or more of the numerical feature vectors indicate height, width, and number of channels in a corresponding convolution operation.
Gao teaches wherein one or more of the numerical feature vectors indicate height, width, and number of channels in a corresponding convolution operation; ([Gao, page 179, sec. 4] “The following symbols are used to denote the hyperparameters and tensor shapes. Sf is the size of input data type (e.g., 4 bytes for FLOAT32 data). N represents batch size. Hk and Wk are kernel (filter) height and width,” AND [Gao, page 177, sec. 3] “Table I lists some commonly used hyperparameters with their domains,” wherein the examiner interprets “Hk and Wk are kernel (filter) height and width” to be the same as numerical feature vectors indicate height and width, and “Table I lists some commonly used hyperparameters with their domains” to be the same as numerical feature vectors indicate number of channels in a corresponding convolution operation).
Chau, Deng, Zhou, Lu, Gao, and the instant application are analogous art because they are all directed to representing convolutional operations using numerical feature vectors to describe their properties for downstream performance prediction.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Chau, Deng, Zhou, and Lu to include the kernels “Hk and Wk are kernel (filter) height and width” disclosed by Gao. One would be motivated to do so to effectively enhance the representation of convolutional operations for improved model performance analysis, as suggested by Gao ([Gao, page 179, sec. 4] “The following symbols are used to denote the hyperparameters and tensor shapes.”). Claim 16 is analogous to claim 6, aside from claim type and minute differences, thus the same mapping applies as above.
Regarding claim 7, Chau, Deng, Zhou, and Lu teach The method of claim 1 (see rejection of claim 1).
Chau, Deng, Zhou, and Lu do not teach wherein the performance metric includes one or more of: latency, execution cycles, and power consumption.
Gao teaches wherein the performance metric includes one or more of: latency, execution cycles, and power consumption; ([Gao, page 179, sec. 4] “In this paper, we consider four representative computational constraints with respect to the model, namely weight size, number of floating-point operations, inference time, and GPU memory consumption,” wherein the examiner interprets “inference time” to be the same as latency because they are both directed to the amount of time required to execute the neural network model, and interprets “number of floating-point operations” to be the same as execution cycles because they are both directed to quantitative measures of the computational operations required to execute the neural network model).
Chau, Deng, Zhou, Lu, Gao, and the instant application are analogous art because they are all directed to predicting performance metrics of neural network models based on computational constraints and operational characteristics.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Chau, Deng, Zhou, and Lu to include the “inference time” disclosed by Gao. One would be motivated to do so to effectively predict latency as part of the performance metrics, as suggested by Gao ([Gao, page 179, sec. 4] “In this paper, we consider four representative computational constraints with respect to the model, namely weight size, number of floating-point operations, inference time, and GPU memory consumption”). Claim 17 is analogous to claim 7, aside from claim type and minute differences, thus the same mapping applies as above.
Claim(s) 8-10 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chau in view of Deng in view of Zhou in view Lu in view of NPL reference “Attention Is All You Need”, by Vaswani et. al. (referred herein as Vaswani).
Regarding claim 8, Chau, Deng, Zhou, and Lu teach The method of claim 1 (see rejection of claim 1).
Chau, Deng, Zhou, and Lu do not teach wherein reducing the dimensions of the encoded sequence further comprises: reducing the dimensions of the encoded sequence using a series of fully-connected layers.
Vaswani teaches wherein reducing the dimensions of the encoded sequence further comprises: reducing the dimensions of the encoded sequence using a series of fully-connected layers; ([Vaswani, page 2, sec. 3] “The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder, shown in the left and right halves of Figure 1, respectively..... In this work we employ h = 8 parallel attention layers, or heads. For each of these we use dk = dv = dmodel/h = 64. Due to the reduced dimension of each head, the total computational cost is similar to that of single-head attention with full dimensionality,” wherein the examiner interprets “point-wise, fully connected layers” to be the same as using a series of fully-connected layers because they are both directed to applying fully connected neural-network layers to encoded representations, and interprets “due to the reduced dimension of each head” to be the same as reducing the dimensions of the encoded sequence because they are both directed to processing the encoded representation in a reduced-dimensional space).
Chau, Deng, Zhou, Lu, Vaswani, and the instant application are analogous art because they are all directed to reducing the dimensions of encoded sequences to optimize neural network operations.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Chau, Deng, Zhou, and Lu to include the “point-wise, fully connected layers for both the encoder and decoder” disclosed by Vaswani. One would be motivated to do so to efficiently reduce the computational complexity of encoding operations, as suggested by Vaswani ([Vaswani, page 2, sec. 3] “Due to the reduced dimension of each head, the total computational cost is similar to that of single-head attention with full dimensionality”). Claim 18 is analogous to claim 8, aside from claim type and minute differences, thus the same mapping applies as above.
Regarding claim 9, Chau, Deng, Zhou, and Lu teach The method of claim 1 (see rejection of claim 1).
Chau, Deng, Zhou, and Lu do not teach wherein the series of attention functions include a series of multi-head attention functions that identify correlations among vectors.
Vaswani teaches wherein the series of attention functions include a series of multi-head attention functions that identify correlations among vectors in the sequence. ([Vaswani, page 2-3, sec. 3.1] “The encoder is composed of a stack of N = 6 identical layers. Each layer has two sub-layers. The first is a multi-head self-attention mechanism, and the second is a simple, position wise fully connected feed-forward network,” wherein the examiner interprets “multi-head self-attention mechanism” to be the same as “a series of multi-head attention functions…The encoder is composed of a stack of N = 6 identical layers” to be the same as “the series of attention functions”).
Chau, Deng, Zhou, Lu, Vaswani, and the instant application are analogous art because they are all directed to the use of attention mechanisms in neural networks to identify relationships among inputs in a sequence.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 Chau, Deng, Zhou, and Lu to include the “multi-head self-attention mechanism” disclosed by Vaswani. One would be motivated to do so to efficiently enhance the ability to identify correlations among vectors in a sequence, as suggested by Vaswani ([Vaswani, page 2-3, sec. 3.1] “The encoder is composed of a stack of N = 6 identical layers ... fully-connected feed-forward network”). Claim 19 is analogous to claim 9, aside from claim type and minute differences, thus the same mapping applies as above.
Regarding claim 10, Chau, Deng, Zhou, and Lu teach The method of claim 1 (see rejection of claim 1).
Chau, Deng, Zhou, and Lu do not teach further comprising: adding input and output of each attention function to generate a sequence of sums; and normalizing the sequence of sums to output to a feed-forward network.
Vaswani teaches further comprising: adding input and output of each attention function to generate a sequence of sums; and normalizing the sequence of sums to output to a feed-forward network. ([Vaswani, page 7, sec. 5.4] “We apply dropout [27] to the output of each sub-layer, before it is added to the sub-layer input and normalized,” AND [Vaswani, pages 2-3, sec. 3.1] “The first is a multi-head self-attention mechanism, and the second is a simple, position wise fully connected feed-forward network. We employ a residual connection [i.e. adding input and output] [10] around each of the two sub-layers, followed by layer normalization [1]. That is, the output of each sub-layer is LayerNorm(x + Sublayer(x)), where Sublayer(x) is the function implemented by the sub-layer itself,” wherein the examiner interprets “the output of each sub-layer, before it is added to the sub-layer input” and “residual connection” to be the same as adding input and output of each attention function to generate a sequence of sums because they are both directed to adding the input of a neural-network sub-layer to the output generated by that sub-layer, and interprets “normalized” and “position wise fully connected feed-forward network” together with “LayerNorm(x + Sublayer(x))” to be the same as normalizing the sequence of sums to output to a feed-forward network because they are both directed to normalizing the summed input and sub-layer output in connection with processing by a feed-forward network).
Chau, Deng, Zhou, Lu, Vaswani, and the instant application are analogous art because they are all directed to improving the processing and transformation of input data in neural networks through attention mechanisms and normalization.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Chau, Deng, Zhou, and Lu to include the process of determining the “residual connection [i.e. adding input and output]” disclosed by Vaswani. One would be motivated to do so to efficiently combine input and output for enhanced model stability and performance, as suggested by Vaswani ([Vaswani, page 7, sec. 5.4] “We apply dropout [27] to the output of each sub-layer, before it is added to the sub-layer input and normalized”). Claim 20 is analogous to claim 10, aside from claim type and minute differences, thus the same mapping applies as above.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DEVAN KAPOOR/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126