Detailed Action
This action is written in response to the application filed 15 March 2024. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections for Lack of Utility
Claims 1-20 are also rejected under 35 U.S.C. 101 because the claimed invention lacks a specific and substantial utility.
The claimed invention lacks a specific utility.
A "specific utility" is specific to the subject matter claimed and can "provide a well-defined and particular benefit to the public." In re Fisher, 421 F.3d 1365, 1371, 76 USPQ2d 1225, 1230 (Fed. Cir. 2005). This contrasts with a general utility that would be applicable to the broad class of the invention. Office personnel should distinguish between situations where an applicant has disclosed a specific use for or application of the invention and situations where the applicant merely indicates that the invention may prove useful without identifying with specificity why it is considered useful.” (MPEP 2107.01(I)(A))
Claim 1 recites “A neural network model optimization method”. There are no meaningful limitations in the claim regarding what the neural network is for—ie what it does—nor limitations as to why or how the neural network is optimized. The same is true for independent claims 9/17.
The applicant’s specification speaks to the wide applicability of neural networks (NNs):
[0003] “A neural network model can complete tasks such as target detection, target classification, machine translation, and speech recognition, and therefore is widely used in various fields such as security protection, transportation, and industrial production.”
However, despite the wide applicability of NNs, no particular real-world problem is specified in the claims.
The claimed invention lacks a substantial utility.
"[A]n application must show that an invention is useful to the public as disclosed in its current form, not that it may prove useful at some future date after further research. Simply put, to satisfy the ‘substantial’ utility requirement, an asserted use must show that the claimed invention has a significant and presently available benefit to the public." In re Fisher, 421 F.3d 1365, 1374, 76 USPQ2d 1225, 1232. See also MPEP 2107.01.
Each of independent claims 1/9/17 is directed to a system for optimizing an unspecified aspect of the recited NN in an unspecified way for an unspecified purpose. Because the recited optimization is unclear, the Examiner finds that the claimed invention does not have any significant benefit to the public that was available at the time of filing based on the specification.
Claim Rejections - 35 USC § 112(b) - Indefiniteness
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b), as being indefinite for failing to particularly point out and distinctly claim the subject matter which applicant regards as the invention.
Claim 1 recites:
“an input of at least one feature transformation module … is obtained based on an output feature of at least one non-adjacent previous network layer of the optimized attention layer.”
However, every pair of nodes (or layers) which is directly connected (ie by synapse connections); thus, it is not clear what “non-adjacent” means in the context of the claim.
Because it is not clear which of the above interpretations is applicable, the term is ambiguous, and consequently a person of ordinary skill would not be able to understand the scope of the claim with reasonable certainty. Therefore the claim is indefinite.
This rejection applies equally to independent claims 9/17, as well as all pending dependent claims, which inherit this deficiency.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 6, 9-10, 14 and 17-18 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Vaswani.
Vaswani, Ashish, et al. "Attention is all you need." Advances in neural information processing systems 30 (2017).
Regarding claims 1, 9 and 17, Vaswani discloses a neural network model optimization method, (and a related device and non-transitory computer-readable medium) comprising:
performing optimization processing on a first neural network model to obtain a second neural network model, wherein the second neural network model comprises an optimized attention layer and at least two previous network layers of the optimized attention layer, and the at least two previous network layers are connected in series; and
PNG
media_image1.png
442
210
media_image1.png
Greyscale
PNG
media_image2.png
266
142
media_image2.png
Greyscale
Applicant’s fig. 2.
Vaswani fig. 2 (left).
P. 7, sec. 5.3, “We used the Adam optimizer [17] with β1 = 0:9, β2 = 0:98 and ϵ = 10-9. We varied the learning rate over the course of training”.
PNG
media_image3.png
512
356
media_image3.png
Greyscale
P. 3, fig. 1, illustrating attention layers with multiple preceding layers.
the optimized attention layer comprises an optimized query feature transformation module, an optimized key feature transformation module, and an optimized value feature transformation module, wherein
P. 4, fig. 2 (reproduced below).
PNG
media_image4.png
342
640
media_image4.png
Greyscale
‘query feature transformation module’ :: Q (linear layer)
‘key feature transformation module’ :: K (linear layer)
‘value feature transformation module’ :: V (linear layer)
an input of the optimized query feature transformation module is obtained based on an output feature of at least one previous network layer of the optimized attention layer;
See fig. 1 (reproduced above): Several other layers precede each attention layer.
an input of the optimized key feature transformation module is obtained based on an output feature of at least one previous network layer of the optimized attention layer;
Id.
an input of the optimized value feature transformation module is obtained based on an output feature of at least one previous network layer of the optimized attention layer; and
Id.
an input of at least one feature transformation module in the optimized query feature transformation module, the optimized key feature transformation module, and the optimized value feature transformation module is obtained based on an output feature of at least one non-adjacent previous network layer of the optimized attention layer.
See fig. 1 (reproduced above), illustrating skipping connections (ie connections which skip one ore more layers).
Regarding independent claims 9 and 17, the computer hardware components recited therein (ie a processor, memory, and a computer-readable medium) are inherent throughout Vaswani.
Regarding claims 2, 10 and 18, Vaswani discloses the further limitation wherein an input of a target feature transformation module is the output feature of the at least one previous network layer of the optimized attention layer, and the target feature transformation module is he optimized query feature transformation module, the optimized key feature transformation module, or the optimized value feature transformation module.
Fig. 2 (reproduced supra), “Scaled dot-product attention”.
Regarding claims 6 and 14, Vaswani discloses the further limitation wherein an input of a target feature transformation module is an input feature obtained by performing weighted summation on output features of the at least two previous network layers of the optimized attention layer and weights of the previous network layers; and
P. 4, fig. 2, “scaled dot-product attention”.
the target feature transformation module is the optimized query feature transformation module, the optimized key feature transformation module, or the optimized value feature transformation module.
P. 4, eqn. 1:
PNG
media_image5.png
62
352
media_image5.png
Greyscale
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
The following are the references relied upon in the rejections below:
Vaswani (Vaswani, Ashish, et al. "Attention is all you need." Advances in neural information processing systems 30 (2017).)
Wang (Wang, Yujing, et al. "Textnas: A neural architecture search space tailored for text representation." Proceedings of the AAAI conference on artificial intelligence. Vol. 34. No. 05. 2020.)
Claims 3-5, 11-13 and 10-20 are rejected under 35 U.S.C. 103 as being unpatentable over Vaswani and Wang.
Regarding claims 3, 11 and 19, Vaswani discloses the further limitation wherein the first neural network model comprises an attention layer and at least two previous network layers of the attention layer that are connected in series, and the at least two previous network layers are connected in series; and
See fig. 1 (reproduced supra).
the performing of optimization processing on a first neural network model to obtain a second neural network model comprises:
Wang discloses the following further limitations which Vaswani does not disclose:
determining a search space of the first neural network model, wherein elements in the search space comprise previous network layers that can be connected to first query feature transformation module, first key feature transformation module, and first value feature transformation module at the attention layer; and
P. 9243, “neural architecture search”.
See also figs. 1 and 2.
determining the optimized attention layer according to a search space-based search algorithm;
Id.
for determining, based on a search condition, a first previous network layer connected to the optimized query feature transformation module, a second previous network layer connected to the optimized key feature transformation module, and a third previous network layer connected to the optimized value feature transformation module; and
See figs. 1 and 2, illustrating a search space comprising four layers.
the first previous network layer, the second previous network layer, and /or the third previous network layer is a non-adjacent previous network layer of the optimized attention layer.
Id. Some of the illustrated layers are non-adjacent.
At the time of filing, it would have been obvious to a skilled machine learning engineer to apply neural architecture search (as taught by Wang) to the Vaswani system because this could improve network performance in the context of the task at hand.
Regarding claims 4, 12 and 20, Wang discloses the further limitation 3 wherein the search space-based search algorithm comprises an evolutionary algorithm, a reinforcement learning algorithm, or a network structure search algorithm.
P. 9243, “neural architecture search”.
Regarding claims 5 and 13, Wang discloses the further limitation wherein the elements in the search space further comprise:
an optional activation function of the first neural network model, an optional normalization operation of the first neural network model, an operation type of an optional feature map of the first neural network model, a quantity of optional parallel branches of the first neural network model, a quantity of modules in an optional search unit, and /or an optional connection manner between previous network layers other than the attention layer.
[The Examiner notes that this is a Markush group.]
The Examiner interprets “optional connection manner” according to its broadest reasonable interpretation as encompassing the different combinations and permutations of connected layers illustrated in figs. 1 and 2 (reproduced supra).
Additional Relevant Prior Art
The following references were identified by the Examiner as being relevant to the disclosed invention, but are not relied upon in any rejection:
Shazeer discloses an attention neural network including techniques for multi-head attention. (US 2021/1279576 A1)
Allowable Subject Matter
Dependent claims 7-8 and 15-16 are allowable over the prior art.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Vincent Gonzales whose telephone number is (571) 270-3837. The examiner can normally be reached on Monday-Friday 7 a.m. to 4 p.m. MT. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang, can be reached at (571) 270-7092.
Information regarding the status of an application may be obtained from the USPTO Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov.
/Vincent Gonzales/Primary Examiner, Art Unit 2124