Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: 501, 502, and 503, denoting steps for determining a data generation model corresponding to each label type of the to-be-optimized model (¶[0079], [0082], [0085]). Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either "Replacement Sheet" or "New Sheet" pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Objections
Claim 3 is objected to because of the following informalities: in the limitation "wherein a distribution of the first training data in the data set and second training data for obtaining the to-be-optimized model meet a similarity condition," "meet" should read --meets-- to agree with its singular subject "a distribution."
Claim 6 is objected to because of the following informalities: the phrase "before generating the at least one piece" should be amended to recite --before generating the at least one piece of fourth training data-- for consistency with the antecedent basis established in claim 5 and with the parallel recitation in claim 15.
Claim 9 is objected to because of the following informalities: the phrase "a performance indicator value of a neural network model" is inconsistent with the parallel recitation "a performance indicator value indicated by a neural network model" in claim 18.
Claim 12 is objected to because of the following informalities: in the limitation "wherein a distribution of the first training data and second training data for obtaining the to-be-optimized model meet a similarity condition," "meet" should read --meets-- to agree with its singular subject "a distribution."
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Claim 1 recites the limitation “performing neural architecture search processing in a search space based on the model file to obtain a neural network architecture that meets the optimization requirement.” Claims 10 and 19 recite the same limitation in device and computer program product form, respectively. The limitation is recited entirely in terms of an input (the model file) and a desired result (an architecture meeting the optimization requirement). It recites no manner of using the model file in the search and no mechanism by which the resulting architecture is caused to meet either the recited performance requirement or the recited hardware requirement.
The specification does not describe this limitation in the breadth in which it is claimed. The only passages that track the claim language merely restate the recited function. The specification at ¶[0064] states that “after receiving the model file and the optimization requirement, the search apparatus obtains the search space, and performs neural architecture search processing in the search space based on the model file, to obtain the neural network architecture that meets the optimization requirement.” The specification at ¶¶[0007], [0043], and [0146] repeats the same sentence in substance. None of these passages discloses how the search is performed, how the model file is used, or how the optimization requirement is enforced.
The specification describes exactly one manner of using the model file to obtain the architecture. At ¶¶[0071]–[0087], the model file is used to derive at least one data generation model; that model generates a synthetic data set whose distribution and the distribution of the original training data of the to-be-optimized model meet a similarity condition; and the architecture search is then performed on the synthetic data set. In that disclosed embodiment, the model file is not an input to the architecture search at all. It is an input to the data generation stage (¶¶[0073]–[0076]), and the search operates on the generated data set (¶¶[0085]–[0086], [0111]–[0138]). Similarly, the only disclosed mechanism by which the resulting architecture is caused to meet the hardware requirement is the removal from the search space of architectures that do not meet it, described at ¶¶[0112]–[0114].
Claims 1, 10, and 19 thus claim a genus — all manners of performing neural architecture search based on the model file so as to yield an architecture meeting the optimization requirement — for which the specification supplies neither a representative number of species nor any structural feature common to the genus. The single disclosed species is claimed separately in dependent claims 3 and 12 and in dependent claims 9 and 18. A claim that is defined by a desired result rather than by the means of achieving it is not supported merely because one means of achieving the result is disclosed. See Abbvie Deutschland GmbH & Co. v. Janssen Biotech, Inc., 759 F.3d 1285, 1300–01 (Fed. Cir. 2014); Ariad Pharmaceuticals, Inc. v. Eli Lilly & Co., 598 F.3d 1336, 1349–51 (Fed. Cir. 2010) (en banc); MPEP § 2163.
Further, because claims 1, 10, and 19 require a computer to perform a specialized function, the specification must disclose the algorithm by which that function is performed. The recitation in claim 10 of a memory and a processor “configured to execute the instructions,” and the recitation in claim 19 of instructions “stored on a computer-readable medium,” do not supply the missing algorithm. See MPEP § 2161.01(I); Vasudevan Software, Inc. v. MicroStrategy, Inc., 782 F.3d 671, 682–83 (Fed. Cir. 2015).
To the extent that claims 1, 10, and 19 substantially track the “first aspect” described at ¶[0007], original claims constitute their own written description only insofar as the disclosure as a whole conveys possession of the full claimed scope. Here it does not. See In re Koller, 613 F.2d 819, 823 (CCPA 1980); MPEP § 2163.
Claims 2, 11, and 20 are rejected for the same reason by virtue of their dependency. Claim 2 and its counterparts add only that the hardware requirement comprises at least one of a hardware specification or an occupied memory size of an optimized model during deployment — a limitation that is itself supported at ¶¶[0009] and [0062], but that does not narrow the deficient search limitation.
Claim 5 recites generating at least one piece of fourth training data “wherein each of the at least one piece of the fourth training data comprises input data and a label predicted value,” and further recites that “a loss value corresponding to each of the at least one piece of the fourth training data is based on the inference result and a calibration predicted value of each of the at least one piece of the fourth training data.”
Claim 14 recites the same two limitations in device form. Claim 5 therefore recites two distinct values associated with the same piece of fourth training data: a label predicted value that is a component of the training data, and a calibration predicted value against which the loss value is computed.
The specification describes only one such value. At ¶[0089] the specification states that the generated training data “includes input data and a calibration predicted value.” At ¶[0091] it states that “[t]he calibration predicted value indicates that a value of a probability that the input data belongs to the target label type is 1, and the calibration predicted value may also be referred to as a label.” At ¶¶[0097]–[0098] the loss value is determined as the difference between the inference result and that same “calibration predicted value.” The summary at ¶[0014] is to the same effect. The specification at ¶[0108], describing FIG. 7, applies the phrase “label predicted value” to the very value against which the loss is computed, confirming that the specification uses “label predicted value” and “calibration predicted value” interchangeably to denote a single value.
Because distinct claim terms are presumed to have distinct meanings, claims 5 and 14 encompass an embodiment in which the fourth training data carries a label predicted value while the loss value is computed from a separate calibration predicted value. The specification nowhere describes such an embodiment, and nowhere describes the origin of a calibration predicted value that is not the value carried by the generated training data. The specification therefore fails to reasonably convey to one skilled in the relevant art that the inventors had possession of the claimed subject matter at the time the application was filed. See MPEP § 2163.
Claims 6, 7, 15, and 16 are rejected for the same reason by virtue of their dependency. The limitations those claims add — obtaining a target neural network model and randomly initializing its weight parameter (¶¶[0016], [0092]–[0093]), and specifying the type of inference data as a text type or an image type (¶¶[0017], [0094]) — are themselves supported, but do not remedy the deficiency inherited from claims 5 and 14.
The remaining dependent claims not touched upon are rejected as being dependent to a rejected base claim.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b), as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Claim 2 recites "an occupied memory size of an optimized model during deployment of the optimized model". Claim 1, from which claim 2 depends, recites only a "to-be-optimized model" (the model being optimized) and a "neural network architecture" (the search result). Claim 2 introduces a new entity, "an optimized model," without identifying whether it refers to (i) the "neural network architecture" obtained through the search of claim 1, (ii) the "to-be-optimized model" after some unrecited optimization step, or (iii) a still further entity such as a model obtained by later training or fine-tuning the returned architecture (specification at ¶[0066], describing that after receiving the neural network architecture the user "may perform full training or model fine-tuning ... to obtain a neural network model suitable for the local service"). Because a PHOSITA cannot determine which structure "an optimized model" denotes, the metes and bounds of the hardware requirement limitation cannot be ascertained with reasonable certainty. For purposes of examination, "an optimized model" in claim 2 is interpreted under BRI to encompass the neural network architecture obtained by the search of claim 1, since specification at [0062] uses "optimized model" in the same sentence structure as the model resulting from the search process, and no other spec passage links the hardware requirement to a separately trained or fine-tuned model.
Claim 11 recites "an occupied memory size of an optimized model during deployment of the optimized model" and depends from claim 10, which (like claim 1) recites only a "to-be-optimized model" and a "neural network architecture." For the same reasons given for claim 2, claim 11 is indefinite because "an optimized model" is introduced without antecedent linkage to any previously recited claim element. For purposes of examination, "an optimized model" in claim 11 is interpreted under BRI to encompass the neural network architecture obtained by the search of claim 10, for the same reasons given with respect to claim 2.
Claim 20 recites "an occupied memory size of an optimized model during deployment of the optimized model" and depends from claim 19, which recites only a "to-be-optimized model" and a "neural network architecture." For the same reasons given for claims 2 and 11, claim 20 is indefinite.
For purposes of examination, "an optimized model" in claim 20 is interpreted under BRI to encompass the neural network architecture obtained by the search of claim 19, for the same reasons given with respect to claim 2.
Claim 3 recites that "a distribution of the first training data in the data set and second training data for obtaining the to-be-optimized model meet a similarity condition." This limitation is grammatically and semantically ambiguous. The singular subject "a distribution" does not agree with the plural verb "meet," and it is unclear whether the claim requires (i) the distribution of the first training data and the distribution of the second training data to satisfy a similarity condition relative to each other, or (ii) a single "distribution" (an aggregate statistical property) to be directly compared against "second training data" (raw, undistributed data points), which is not a coherent comparison. A PHOSITA cannot determine with reasonable certainty what comparison the claim requires. For purposes of examination, this limitation is interpreted under BRI to require that a distribution of the first training data and a distribution of the second training data satisfy a similarity condition (e.g., a matching or close variance/average value as between the two distributions), consistent with specification at ¶¶[0010]-[0011] and [0074]-[0076], which describe generated training data whose distribution is "the same as or close to at least one of a variance and an average value" of the original training data.
Claim 12 recites the analogous limitation "a distribution of the first training data and second training data for obtaining the to-be-optimized model meet a similarity condition." For the same reasons given with respect to claim 3, this limitation is grammatically ambiguous (singular "a distribution" with plural verb "meet") and it is unclear whether the claim compares two distributions or a single distribution against raw data. For purposes of examination, this limitation is interpreted under BRI to require that a distribution of the first training data and a distribution of the second training data satisfy a similarity condition, for the same reasons given with respect to claim 3.
Claim 4 recites "generating third training data corresponding to each label type" and "the first data generation model corresponds to each label type of the to-be-optimized model." Neither claim 4 nor claim 3 (from which it depends) previously introduces a plurality of "label types" or any "label type" of the to-be-optimized model. Because "each label type" presupposes an antecedent plurality of label types that is never recited, there is insufficient antecedent basis for this limitation, and the scope of "each label type" (how many, and which, label types are required) cannot be ascertained. For purposes of examination, "each label type" is interpreted under BRI to refer to each of a plurality of classification categories that the to-be-optimized model is configured to output, consistent with specification at ¶¶[0079]-[0090], which describe the to-be-optimized model as having multiple "label types" (e.g., "eight label types" for an eight-class classifier) even though the claim itself never introduces this plurality.
Claim 13 recites "generate third training data corresponding to each label type" and "the first data generation model corresponds to each label type of the to-be-optimized model." For the same reasons given with respect to claim 4, "each label type" lacks antecedent basis because no plurality of label types is previously recited in claim 13 or claim 12 (from which it depends). For purposes of examination, "each label type" is interpreted under BRI in the same manner as for claim 4, for the same reasons.
Claim 5 depends from claim 4 and recites "for a target label type of the to-be-optimized model." Because claim 4's antecedent-basis deficiency regarding "each label type" is not cured, and because claim 5 presupposes that "a target label type" is one of the (never-established) plurality of label types, claim 5 inherits the indefiniteness of claim 4. For purposes of examination, "a target label type" in claim 5 is interpreted under BRI as one of the plurality of label types identified for claim 4, above.
Claim 6 recites "wherein before generating the at least one piece, the method further comprises:" Claim 5 (from which claim 6 depends) introduces "at least one piece of fourth training data," but claim 6 refers back only to "the at least one piece" without repeating "of fourth training data." By contrast, the analogous device claim 15 correctly recites "before generating the at least one piece of the fourth training data." As literally worded, "the at least one piece" in claim 6 lacks a complete antecedent basis and its scope (piece of what) is not clear from the claim language alone. In addition, claim 6 inherits the antecedent-basis deficiency of claims 4-5 with respect to "each label type"/"a target label type." For purposes of examination, "the at least one piece" in claim 6 is interpreted under BRI to refer to "the at least one piece of fourth training data" recited in claim 5, consistent with the correctly-worded parallel language of device claim 15.
Claim 7 depends from claim 6 and inherits the antecedent-basis deficiencies of claims 4, 5, and 6 discussed above, none of which are cured by the additional limitations of claim 7. For purposes of examination, claim 7 is interpreted under BRI consistent with the interpretations given for claims 4-6, above.
Claim 14 depends from claim 13 and recites "for a target label type of the to-be-optimized model," inheriting the antecedent-basis deficiency of claim 13 for the same reasons given with respect to claim 5. For purposes of examination, "a target label type" in claim 14 is interpreted under BRI in the same manner as claim 5, above.
Claim 15 depends from claim 14 and inherits the antecedent-basis deficiencies of claims 13-14 discussed above. Claim 15 itself is correctly worded ("the at least one piece of the fourth training data") and introduces no new indefiniteness. For purposes of examination, claim 15 is interpreted under BRI consistent with the interpretations given for claims 13-14, above.
Claim 16 depends from claim 15 and inherits the antecedent-basis deficiencies of claims 13-15 discussed above. For purposes of examination, claim 16 is interpreted under BRI consistent with the interpretations given for claims 13-15, above.
Claim 5 recites that fourth training data "comprises input data and a label predicted value," but later recites that a loss value "is based on the inference result and a calibration predicted value of each of the at least one piece of the fourth training data." The claim does not establish whether "label predicted value" and "calibration predicted value" refer to the same attribute of the fourth training data or to two distinct attributes. The specification uses only the single term "calibration predicted value" for this attribute (specification at ¶¶[0089], [0091]) and does not use or define a separate "label predicted value." The inconsistent terminology renders it unclear what data element(s) claim 5 actually requires. For purposes of examination, "a label predicted value" and "a calibration predicted value" in claim 5 are interpreted under BRI as referring to the same attribute — i.e., a value indicating that the input data belongs to the target label type with probability 1 — consistent with specification at ¶[0091] ("The calibration predicted value indicates that a value of a probability that the input data belongs to the target label type is 1, and the calibration predicted value may also be referred to as a label"), which is the only definition of a predicted-value attribute of training data disclosed in the specification.
Claim 14 recites the analogous limitations using both "a label predicted value" and "a calibration predicted value" for what appears to be the same attribute of the fourth training data. For the same reasons given with respect to claim 5, this inconsistent terminology renders claim 14 indefinite. For purposes of examination, these terms in claim 14 are interpreted under BRI in the same manner as claim 5, above.
Claim 6 depends from claim 5 and inherits the inconsistent-terminology deficiency of claim 5 discussed above. For purposes of examination, claim 6 is interpreted under BRI consistent with the interpretation given for claim 5, above.
Claim 7 depends from claim 6 and inherits the inconsistent-terminology deficiency of claim 5 discussed above. For purposes of examination, claim 7 is interpreted under BRI consistent with the interpretation given for claim 5, above.
Claim 15 depends from claim 14 and inherits the inconsistent-terminology deficiency of claim 14 discussed above. For purposes of examination, claim 15 is interpreted under BRI consistent with the interpretation given for claim 14, above.
Claim 16 depends from claim 15 and inherits the inconsistent-terminology deficiency of claim 14 discussed above. For purposes of examination, claim 16 is interpreted under BRI consistent with the interpretation given for claim 14, above.
Claim 9 recites that each of the pieces of third training data "comprises a performance indicator value of a neural network model and the performance requirement." As written, this limitation is ambiguous as to whether the recited training data comprises (i) a performance indicator value, and, separately, the performance requirement itself (an abstract target/criterion, e.g., a target latency or accuracy threshold) as a literal component of the training data, or (ii) a performance indicator value that corresponds to or is measured against the performance requirement. Requiring training data to literally "comprise" the performance requirement produces an unclear result, since a requirement is not the kind of thing ordinarily stored as a data value alongside a performance indicator value. For purposes of examination, this limitation is interpreted under BRI to require that each piece of training data comprises a performance indicator value of a neural network model together with an indication of the performance requirement metric to which that indicator value corresponds (e.g., a latency value paired with an indication that it is a latency measurement), consistent with specification at ¶¶[0121], [0125]-[0127], which describe generating training data comprising "a performance indicator value indicated by a neural network model and the performance requirement" for purposes of training a corresponding evaluation model for that performance indicator.
Claim 18 recites the analogous limitation that each piece of third training data "comprises a performance indicator value indicated by a neural network model and the performance requirement." For the same reasons given with respect to claim 9, it is ambiguous whether the training data literally comprises the performance requirement itself or a value corresponding to it. For purposes of examination, this limitation is interpreted under BRI in the same manner as claim 9, above.
The remaining dependent claims not touched upon are rejected as being dependent to a rejected base claim.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 2, 10, 11, 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Data-Free Neural Architecture Search via Recursive Label Calibration to Liu et al. (hereinafter Liu) in view of US Pat. Pub. No. 2022/0019869A1 to Li et al. (hereinafter Li).
Per claim 1, Liu discloses A method (Liu: p. 1, Abstract…Liu presents a procedure that performs neural architecture search using0 nothing but a model that was previously trained on the unavailable data, which constitutes the recited method under BRI, "This paper aims to explore the feasibility of neural architecture search (NAS) given only a pre-trained model without using any original training data"), comprising:
performing neural architecture search processing in a search space based on the model file to obtain a neural network architecture that meets the optimization requirement (Liu: p. 2, Section 1…Liu's search operates on the pre-trained model alone, the model being the only artifact shared, so the file embodying that model is the basis on which the search proceeds, which constitutes performing the search based on the model file, "We assume that we only have a pre-trained model, and the original dataset is not accessible during the neural architecture search process"; p. 4, Section 3…Liu searches a defined search space S for the architecture that optimizes the search objective, and the architecture so returned is the one that satisfies that objective, which constitutes obtaining an architecture in a search space that meets the optimization requirement, "we aim to use synthesized data…to search for the high-performing architecture A in the search space of S"); and…
Liu does not expressly disclose, but Li does teach:
receiving an optimization request comprising a model file and an optimization requirement of a to-be-optimized model, wherein the optimization requirement comprises a performance requirement and a hardware requirement, and wherein the performance requirement comprises at least one of an inference latency, a recall rate, or an accuracy; (Li: ¶[0033]…Li's architecture search system takes in over a published network interface whatever artifact the remote user uploads to initiate a search, and in the combination with Liu that uploaded artifact is Liu's pre-trained model file rather than Li's training data, Li supplying the request-intake path and Liu supplying the model file the request carries, which together constitute receiving an optimization request comprising a model file, "the system 100 can receive training data as an upload from a remote user of the system over a data communication network, e.g., using an application programming interface (API) made available by the system 100"; ¶[0010]…Li conditions the same search on the particular hardware the found architecture must run on and on objectives expressed in accuracy and latency terms, which constitutes the hardware requirement and the at least one of an inference latency or an accuracy that the optimization requirement carries, "the systems and techniques described herein may use a search space augmented with operations that are specific to the target set of hardware resources and multi-objective performance metrics that take both accuracy and latency into account when selecting neural network architectures"); …
sending the neural network architecture (Li: ¶[0050]…Li transmits the architecture the search produced back to the same user who submitted the upload that initiated it, which constitutes sending the neural network architecture, "the neural network search system 100 can output the architecture data 150 to the user that submitted the training data").
Liu and Li are analogous art because they are from the same field of endeavor, specifically the automated search for neural network architectures. They address the same problem of producing an architecture that suits the conditions under which the resulting model will actually be trained and deployed.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to carry out Liu's data-free architecture search inside Li's user-facing search service, so that what the user uploads is the pre-trained model rather than the training data and the request carrying that upload also states the accuracy or latency objective and the hardware the architecture must run on.
The suggestion/motivation for doing so would have been provided by Liu itself, which identifies the exchange of the model in place of the data as an affirmative advantage of its approach, "Also, models are usually smaller in size than large-scale datasets, which makes them easier to exchange and store" (Liu: p. 2, Section 1), and which frames the data-free setting as one arising precisely where an owner will share a model but not the data behind it, "a practical task for application scenarios in which privacy or logistical concerns restrict sharing of the original training data but permit sharing of a model trained by such data for NAS" (Liu: p. 2, Section 1). Li supplies the corresponding intake and return path for exactly such an upload (Li: ¶[0033], ¶[0050]) and already conditions its search on the requester's hardware and on accuracy and latency objectives (Li: ¶[0010]), so a PHOSITA reading the two together would have had every reason to let Liu's shareable model be the artifact Li's service receives. Furthermore, this is the combination of prior art elements according to known methods to yield predictable results, the rationale of MPEP § 2143(A).
Per claim 2, Liu combined with Li discloses claim 1. Liu does not expressly disclose, but Li does teach the hardware requirement comprises at least one of a hardware specification or an occupied memory size of an optimized model during deployment of the optimized model (Li: ¶[0009]…Li names the target hardware by the specific class of accelerator device the architecture is to be deployed on, which constitutes a hardware specification within the hardware requirement, "such a target set of hardware resources may correspond to one or more datacenter accelerators including one or more tensor processing units (TPUs), one or more graphics processing units (GPUs), or a combination thereof"). The rationale to combine Li with Liu is the same as the parent claim.
Per claim 10, Liu discloses:
…
perform neural architecture search processing in a search space based on the model file to obtain a neural network architecture that meets the optimization requirement (Liu: p. 2, Section 1…as in claim 1, Liu's search draws on the pre-trained model as its sole input, "We assume that we only have a pre-trained model, and the original dataset is not accessible during the neural architecture search process"); and…
Liu does not expressly disclose, but Li does teach:
A device (Li: ¶[0005]…Li implements its architecture search as programs running on one or more computers, which constitutes the recited device under BRI, "This specification describes how a system implemented as computer programs on one or more computers in one or more locations"), comprising:
a memory configured to store instructions (Li: ¶[0083]…Li identifies the memory device holding the instructions as an essential element of the computer that carries out its method, which constitutes a memory configured to store instructions, "The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data"); and
a processor coupled to the memory and configured to execute the instructions to cause the device to (Li: ¶[0083]…the same paragraph recites the central processing unit that receives and executes those stored instructions and describes that unit and the memory as joined, which constitutes a processor coupled to the memory and configured to execute the instructions, "The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry"):
receive an optimization request comprising a model file and an optimization requirement of a to-be-optimized model, wherein the optimization requirement comprises a performance requirement and a hardware requirement, and wherein the performance requirement comprises at least one of an inference latency, a recall rate, or an accuracy; (Li: ¶[0033]…as in claim 1, the request arrives as the remote user's upload over the network interface Li exposes and, in the combination with Liu, carries Liu's pre-trained model file, "the system 100 can receive training data as an upload from a remote user of the system over a data communication network, e.g., using an application programming interface (API) made available by the system 100"; ¶[0042]…Li measures the elapsed time the candidate takes to produce an output on the requester's hardware, which constitutes the inference latency the performance requirement may comprise, "the target hardware deployment engine 130 may measure or determine (i) a latency of generating an output using the candidate neural network when deployed on the target set of hardware resources"); …
send the neural network architecture (Li: ¶[0050]…as in claim 1, the resulting architecture is returned to the requesting user, "the neural network search system 100 can output the architecture data 150 to the user that submitted the training data").
Liu and Li are analogous art because they are from the same field of endeavor, specifically the automated search for neural network architectures. They address the same problem of producing an architecture that suits the conditions under which the resulting model will actually be trained and deployed.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to host Liu's data-free architecture search on the programmed computer Li describes, whose processor executes the stored instructions that receive the user's upload, run the search and return the architecture, as claim 10 recites.
The suggestion/motivation for doing so would have been provided by Li itself, which states that its architecture search is realized as computer programs executing on one or more computers, "This specification describes how a system implemented as computer programs on one or more computers in one or more locations" (Li: ¶[0005]), and identifies the processor and memory as the essential elements of such a computer, "The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data" (Li: ¶[0083]). Liu's method is itself a computation and must run on some such machine, so a PHOSITA had a plain reason to embody it in the hardware Li specifies. Furthermore, this is the use of a known technique to improve a similar device in the same way, the rationale of MPEP § 2143(C).
Per claim 11, Liu combined with Li discloses claim 10. Liu does not expressly disclose, but Li does teach wherein the hardware requirement comprises at least one of a hardware specification or an occupied memory size of an optimized model during deployment of the optimized model. (Li: ¶[0009]…as in claim 2, Li identifies the target hardware by the class of accelerator device on which the architecture is to run, "such a target set of hardware resources may correspond to one or more datacenter accelerators including one or more tensor processing units (TPUs), one or more graphics processing units (GPUs), or a combination thereof"). The rationale to combine Li with Liu is the same as the parent claim.
Per claim 19, Liu discloses:
…
perform neural architecture search processing in a search space based on the model file to obtain a neural network architecture that meets the optimization requirement (Liu: p. 2, Section 1…as in claim 1, "We assume that we only have a pre-trained model, and the original dataset is not accessible during the neural architecture search process"); and …
Liu does not expressly disclose, but Li does teach:
A computer program product comprising computer-executable instructions that are stored on a computer-readable medium and that, when executed by a processor, cause a device to (Li: ¶[0077]…Li stores the very instructions that carry out its architecture search on a tangible non-transitory medium for execution by a data processing apparatus, which constitutes the recited computer program product and its stored computer-executable instructions, "Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus"; ¶[0076]…Li describes those instructions as causing the apparatus to carry out the recited operations when executed, "the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions"):
receive an optimization request comprising a model file and an optimization requirement of a to-be-optimized model, wherein the optimization requirement comprises a performance requirement and a hardware requirement, and wherein the performance requirement comprises at least one of an inference latency, a recall rate, or an accuracy (Li: ¶[0033]…as in claim 1, and in the combination with Liu the uploaded artifact is Liu's pre-trained model file, "the system 100 can receive training data as an upload from a remote user of the system over a data communication network, e.g., using an application programming interface (API) made available by the system 100"; ¶[0010]…as in claim 1, "multi-objective performance metrics that take both accuracy and latency into account when selecting neural network architectures"); …
send the neural network architecture (Li: ¶[0050]…as in claim 1, "the neural network search system 100 can output the architecture data 150 to the user that submitted the training data").
Liu and Li are analogous art because they are from the same field of endeavor, specifically the automated search for neural network architectures. They address the same problem of producing an architecture that suits the conditions under which the resulting model will actually be trained and deployed.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to distribute Liu's data-free architecture search as instructions held on the computer-readable medium Li describes, which when executed cause the device to receive the request, run the search and send back the architecture, as claim 19 recites.
The suggestion/motivation for doing so would have been provided by Li itself, which states that the operations of its architecture search are embodied as program instructions on a tangible storage medium, "one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus" (Li: ¶[0077]). Storing Liu's procedure in that same form is no more than the ordinary means of delivering it to the machine that must run it. Furthermore, this is the combination of prior art elements according to known methods to yield predictable results, the rationale of MPEP § 2143(A).
Per claim 20, Liu combined with Li discloses claim 19. Liu does not expressly disclose, but Li does teach the hardware requirement comprises at least one of a hardware specification or an occupied memory size of an optimized model during deployment of the optimized model (Li: ¶[0009]…as in claim 2, "such a target set of hardware resources may correspond to one or more datacenter accelerators including one or more tensor processing units (TPUs), one or more graphics processing units (GPUs), or a combination thereof"). The rationale to combine Li with Liu is the same as the parent claim.
Claims 3-5, 8, 12-14 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Li, as applied in the rejection of claims 1 and 10 above, and further in view of Data-Free Learning of Student Networks to Chen et al. (hereinafter Chen).
Per claim 3, Liu combined with Li discloses claim 1. Liu further teaches wherein performing the neural architecture search processing comprises:
…
further performing the neural architecture search processing in the search space based on the data set to obtain the neural network architecture (Liu: p. 5, Section 3…Liu feeds the synthesized data set, rather than any original data, into the architecture search that yields the returned architecture, which constitutes further performing the search in the search space based on the data set, "Such that we can conduct data-free NAS by integrating the synthesized data and their corresponding soft labels with existing NAS algorithms").
Liu combined with Li does not expressly disclose, but Chen does teach: …
generating a data set using at least one data generation model that corresponds to the to-be-optimized model and that is based on the model file, wherein the data set comprises first training data, and wherein a distribution of the first training data in the data set and second training data for obtaining the to-be-optimized model meet a similarity condition (Chen: p. 1, Abstract…Chen builds a generator network whose sole training signal comes from the already-trained network being compressed, so the generator is derived from and tied to that network, which constitutes a data generation model that corresponds to the to-be-optimized model and is based on its model file, "we propose a novel framework for training efficient deep neural networks by exploiting generative adversarial networks (GANs). To be specific, the pre-trained teacher networks are regarded as a fixed discriminator and the generator is utilized for derivating training samples which can obtain the maximum response on the discriminator"; p. 6, Section 4.1…Chen reports as an achieved result that the data its generator emits reproduces the distribution of the network's original training data better than substitute data does, which constitutes the recited similarity condition between the generated first training data and the second training data used to obtain the to-be-optimized model, "the accuracy of student network using the proposed algorithm is superior to these using other data (normal distribution, USPS dataset and reconstructed dataset using “meta data”), which suggest that our method could imitate the distribution of training dataset better"); and …
Liu, Li and Chen are analogous art because all three references are from the same field of endeavor, specifically the automated design and compression of neural networks in settings where the original training data is withheld. Liu and Chen address the same problem of recovering usable training data from a network whose own data cannot be shared.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to produce the data set on which Liu runs its architecture search with Chen's generator network, trained against the uploaded model's own outputs, rather than by Liu's direct optimization of the pixels.
The suggestion/motivation for doing so would have been provided by Liu itself, which surveys the generator-based route as an established alternative for synthesizing data from a pre-trained model and cites Chen by name for it, "Recently, Chen et al. [5] and Xu et al. [28] proposed to use a generator to synthesize images from a pre-trained model and simultaneously train the student network" (Liu: p. 4, Section 2). Chen in turn states that its generator was devised for the same withheld-data predicament Liu confronts, "training data for the given deep network are often unavailable due to some practice problems (e.g. privacy, legal issue, and transmission)" (Chen: p. 1, Abstract), and both references synthesize data by driving it against the fixed pre-trained network's own responses, so the two techniques are interchangeable inputs to the same downstream search. Furthermore, this is the simple substitution of one known element for another to obtain predictable results, the rationale of MPEP § 2143(B).
Per claim 4, Liu combined with Li and Chen discloses claim 3. Liu combined with Li does not expressly disclose, but Chen does teach: wherein generating the data set comprises generating third training data corresponding to each label type using a first data generation model, wherein the first data generation model corresponds to each label type of the to-be-optimized model and is based on the model file, and wherein the third training data forms the data set (Chen: p. 4, Section 3.2…Chen trains one generator and drives it with a loss that holds the generated data balanced across every class the network recognizes, so that single generator answers to each label type rather than to only one, and the data it emits is the whole set used downstream, which constitutes a first data generation model corresponding to each label type whose generated third training data forms the data set, "We employ the information entropy loss to measure the class balance of generated images"; p. 1, Abstract…the generator is trained solely against the already-trained network standing in as a fixed discriminator, which is what makes the generator based on the model file, "the pre-trained teacher networks are regarded as a fixed discriminator and the generator is utilized for derivating training samples which can obtain the maximum response on the discriminator"). The rationale to combine Chen with Liu and Li is the same as the parent claim.
Per claim 5, Liu combined with Li and Chen discloses claim 4. Liu combined with Li does not expressly disclose, but Chen does teach: generating, for a target label type of the to-be-optimized model using an initial data generation model corresponding to the target label type, at least one piece of fourth training data corresponding to the target label type, wherein each of the at least one piece of the fourth training data comprises input data and a label predicted value (Chen: p. 4, Algorithm 1…Chen draws a batch of latent vectors and runs the generator on them to emit the training samples, which constitutes generating pieces of training data using the generator, "Randomly generate a batch of vector: {zᵢ}ⁿᵢ₌₁; Generate the training samples: x ← G(z)"; p. 4, Section 3.2…each emitted sample is paired with the label the network itself predicts for it, which constitutes each piece comprising input data and a label predicted value, "The predicted labels {t₁, t₂, ⋯ , tₙ} are then calculated by tᵢ = arg max(yᵢᵀ)ⱼ");
inputting the input data into the to-be-optimized model to obtain an inference result corresponding to each of the at least one piece of the fourth training data, wherein a loss value corresponding to each of the at least one piece of the fourth training data is based on the inference result and a calibration predicted value of each of the at least one piece of the fourth training data (Chen: p. 4, Section 3.2…Chen passes each generated sample through the network being compressed to collect that network's output, which constitutes obtaining an inference result for each piece, "Inputting these images into the teacher network, we can obtain the outputs {y₁ᵀ, y₂ᵀ, ⋯ , yₙᵀ} with yᵢᵀ = NT(xᵢ)"; p. 4, Section 3.2…the loss is then formed by comparing that collected output against the pseudo label taken as ground truth for the sample, which constitutes a loss value based on the inference result and a calibration predicted value, "By taking {t₁, t₂, ⋯ , tₙ} as pseudo ground-truth labels, we formulate the one-hot loss function"); and
updating a first weight parameter of the initial data generation model based on the loss value to obtain an updated initial data generation model, wherein a second data generation model corresponding to the target label type is based on the updated initial data generation model. (Chen: p. 4, Algorithm 1…Chen carries that loss back into the generator and revises the generator's own weights with it, and the generator that emerges from the repeated revision is the one thereafter used to produce data, which constitutes updating a first weight parameter to obtain an updated model on which the second data generation model is based, "Calculate the loss function L_Total (Fcn.7): Update weights in G using back-propagation").
The rationale to combine Chen with Liu and Li is the same as the parent claim.
Per claim 8, Liu combined with Li and Chen discloses claim 3. Liu combined with Li does not expressly disclose, but Chen does teach the at least one data generation model is a generative adversarial network (GAN) model (Chen: p. 1, Abstract…Chen's data generation model is the generator half of a generative adversarial network, the network being compressed standing in as the fixed discriminator, which constitutes the data generation model being a GAN model, "we propose a novel framework for training efficient deep neural networks by exploiting generative adversarial networks (GANs)") . The rationale to combine Chen with Liu and Li is the same as the parent claim.
Claims 12-14 and 17 are substantially similar in scope and spirit as claims 3-5 and 8. Therefore the rejections of claims 3-5 and 8 are applied accordingly.
Claims 6, 7, 15 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Li and Chen, as applied in the rejection of claims 5 and 14 above, and further in view of Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks to Radford et al. (hereinafter Radford).
Per claim 6, Liu combined with Li and Chen discloses claim 5. Liu combined with Li does not expressly disclose, but Chen does teach wherein before generating the at least one piece, the method further comprises: obtaining a target neural network model for generating data (Chen: p. 4, Algorithm 1…Chen takes a generator network and initializes it as the first step of the procedure, before any sample is generated, which constitutes obtaining a target neural network model for generating data prior to the generating step, "Initialize the generator G, the student network NS with fewer memory usage and computational complexity"; p. 6, Section 4.1…Chen identifies that network as a standard convolutional generator taken from the deep convolutional generative adversarial network design of Radford, "We use a deep convolutional generator following [19] and add a batch normalization at the end of the generator to smooth the sample values"); and …
Liu combined with Li combined with Chen does not expressly disclose, but Radford does teach: performing random initialization on a second weight parameter of the target neural network model to obtain the initial data generation model (Radford: p. 3, Section 4…Radford specifies that the weights of the deep convolutional generative adversarial network it defines - the very generator design Chen adopts by reference - are set at the outset by drawing each from a zero-centred normal distribution, which constitutes performing random initialization on a weight parameter of the target neural network model to obtain the initial data generation model, "All weights were initialized from a zero-centered Normal distribution with standard deviation 0.02").
Liu, Li, Chen and Radford are analogous art because all four references reside within the field of endeavor of neural network design and training, and Radford is reasonably pertinent to the same problem Chen addresses in that it defines the convolutional generator architecture and training procedure Chen builds upon.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to initialize the weights of Chen's generator in the manner Radford prescribes, by drawing each weight from a zero-centred normal distribution, so that the generator Chen begins from is a randomly initialized instance of the target network.
The suggestion/motivation for doing so would have been provided by Chen itself, which does not merely resemble Radford's design but expressly adopts it, "We use a deep convolutional generator following [19] and add a batch normalization at the end of the generator to smooth the sample values" (Chen: p. 6, Section 4.1), Chen's reference [19] being Radford, "Alec Radford, Luke Metz, and Soumith Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks" (Chen: p. 9, References). A PHOSITA implementing Chen's generator would therefore have turned to Radford for the initialization Chen leaves unstated, and would there have found it prescribed, "All weights were initialized from a zero-centered Normal distribution with standard deviation 0.02" (Radford: p. 3, Section 4). Furthermore, this is the use of a known technique to improve a similar method in the same way, the rationale of MPEP § 2143(C).
Per claim 7, Liu combined with Li, Chen and Radford discloses claim 6. Liu combined with Chen and Radford does not expressly disclose, but Li does teach the optimization requirement further comprises a type of inference data inferred by the to-be-optimized model, wherein the type is a text type or an image type, and wherein the data are the same type as the type of the inference data (Li: ¶[0059]…the request Li's system receives states the particular machine learning task the architecture is to perform, so the request itself carries the character of the data the resulting network will infer upon, which constitutes the optimization requirement further comprising a type of inference data, "The system receives training data for performing a particular machine learning task"; ¶[0054]…Li names that task as an image processing task and states that the network input for it is image data, which constitutes the type being an image type and the data being of the same type as the inference data, "the particular machine learning task that the neural network architecture 200 is configured to perform may be an image processing task. In this example, the network input 202 may correspond to data representing one or more images"). The rationale to combine Li with Liu is the same as claim 1.
Claims 15 and 16 are substantially similar in scope and spirit as claims 6 and 7. Therefore the rejections of claims 6 and 7 are applied accordingly.
Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Li and Chen, as applied in the rejection of claims 3 and 12 above, and further in view of Once-for-All: Train One Network and Specialize it for Efficient Deployment to Cai et al. (hereinafter Cai).
Per claim 9, Liu combined with Li and Chen discloses claim 3. Li further teaches optimizing the search space based on the hardware requirement to obtain an optimized search space (Li: ¶[0010]…Li tailors the set of candidate operations the search may draw from to the particular hardware the architecture must run on, so the space the search actually explores is one already shaped by the hardware requirement, which constitutes optimizing the search space based on the hardware requirement to obtain an optimized search space, "the systems and techniques described herein may use a search space augmented with operations that are specific to the target set of hardware resources"; ¶[0035]…Li repeats that the operations populating its candidate search space are chosen for the target hardware, "the set or list of operations reflected in the candidate architecture search space 111 may include operations that are specific to the target set of hardware resources on which the candidate neural network architectures are intended to run");
Liu further teaches training a super net in the optimized search space based on the data set to obtain a trained super net (Liu: p. 9…Liu's Single Path One-Shot instantiation trains a SuperNet over the search space on the synthesized data set before any architecture is selected, which constitutes training a super net in the search space based on the data set to obtain a trained super net, "After the supernet is trained, the weights can be used for evolutionary search in the second separate step"); …
Liu combined with Li combined with Chen does not expressly disclose, but Cai does teach:
generating pieces of third training data using the trained super net, wherein each of the pieces of the third training data comprises a performance indicator value of a neural network model and the performance requirement (Cai: p. 6, Section 3.4…Cai draws sub-networks from the trained super net and measures each one's accuracy, yielding a body of paired records in which a neural network model is coupled to its measured performance indicator value, which constitutes generating pieces of training data each comprising a performance indicator value of a neural network model, "we randomly sample 16K sub-networks with different architectures and input image sizes, then measure their accuracy on 10K validation images sampled from the original training set"; p. 6, Section 3.4…the indicator recorded in those records is the very accuracy-and-latency objective the search is required to satisfy, which constitutes each piece further comprising the performance requirement, "The goal is to search for a neural network that satisfies the efficiency (e.g., latency, energy) constraints on the target hardware while optimizing the accuracy");
training an evaluation model based on the pieces of the third training data to obtain a trained evaluation model (Cai: p. 6, Section 3.4…Cai fits a predictor on exactly those paired records so that the predictor thereafter estimates a candidate's performance from its architecture alone, which constitutes training an evaluation model on the pieces of training data to obtain a trained evaluation model, "These [architecture, accuracy] pairs are used to train an accuracy predictor to predict the accuracy of a model given its architecture and input image size"; p. 14, Appendix A…Cai identifies the trained evaluation model as a neural network in its own right, "We use a three-layer feedforward neural network that has 400 hidden units in each layer as the accuracy predictor"); and
obtaining the neural network architecture through a search based on the trained evaluation model, the pieces of the third training data, and an evolution algorithm. (Cai: p. 6, Section 3.4…Cai runs an evolutionary search whose fitness signal is supplied by the trained predictor rather than by measurement, and returns the architecture that search selects, which constitutes obtaining the neural network architecture through a search based on the trained evaluation model and an evolution algorithm, "Given the target hardware and latency constraint, we conduct an evolutionary search (Real et al., 2019) based on the neural-network-twins to get a specialized sub-network").
Liu, Li, Chen and Cai are analogous art because all four references reside within the field of endeavor of automated neural network architecture design. Cai and Li are reasonably pertinent to the same problem Liu and Chen address in that each is directed to selecting an architecture that satisfies stated accuracy and hardware constraints.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to carry out the selection step of Liu's super-net-based search in the manner Cai teaches, by sampling sub-networks from the trained super net, fitting a predictor on the sampled architecture-and-performance records, and running the evolutionary search against that predictor, as claim 9 recites.
The suggestion/motivation for doing so would have been provided by Cai itself, which states that predicting rather than measuring each candidate's performance is what removes the repeated cost of the search, "It eliminates the repeated search cost by substituting the measured accuracy/latency with predicted accuracy/latency (twins)" (Cai: p. 6, Section 3.4), and which directs that search at the same accuracy-under-hardware-constraint objective Li imposes, "Given the target hardware and latency constraint, we conduct an evolutionary search" (Cai: p. 6, Section 3.4). Liu already trains a super net and already selects from it by evolutionary search (Liu: p. 9), so Cai's predictor is added to a pipeline that is otherwise in place and is applied to it in the very way Cai prescribes. Furthermore, this is the use of a known technique to improve a similar method in the same way, the rationale of MPEP § 2143(C).
Claim 18 is substantially similar in scope and spirit as claim 9. Therefore the rejection of claim 9 is applied accordingly.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALAN CHEN whose telephone number is (571)272-4143. The examiner can normally be reached M-F 10-7.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALAN CHEN/Primary Examiner, Art Unit 2125