DETAILED ACTION
This action is responsive to the claims filed on 03/05/2026. Claims 1-13, 15-16, 18-26, and 28-30 are pending for examination.
This action is Final.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Response to Applicant’s § 101 Arguments:
Step 2A, Prong 1 — mental process / mathematical concept (Remarks, page 13)
The Examiner respectfully disagrees with Applicant’s assertion that claim 1 does not recite a mental process or mathematical concept. As set forth in the rejection, claim 1 recites selecting a reconstruction level set based on previously decoded indices, decoding a quantization index, dequantizing to a reconstruction level, determining a set of reconstruction levels based on a state, updating that state based on a previously decoded index, and decoding using arithmetic coding with a probability model dependent on that state. These limitations describe evaluation, selection, and mathematical manipulation of data, i.e., operations that can be carried out conceptually or with pen and paper at the level recited, even if a practical implementation would ordinarily be performed by a computer for speed or scale. Applicant’s argument that neural network parameters may be high-dimensional or numerous is not persuasive because the claim does not recite any particular hardware architecture, timing constraint, or implementation detail that would remove the recited steps from the category of mental evaluation and mathematical processing identified in the rejection. The claim instead recites the algorithmic rules themselves.
Step 2A, Prong 2 — alleged practical application / technological improvement (Remarks, page 14)
The Examiner respectfully disagrees with Applicant’s contention that the claims integrate the alleged abstract idea into a practical application. The claim is limited to the field of decoding neural network parameters from a data stream, but that field-of-use limitation does not, by itself, integrate the judicial exception into a practical application. The processor is recited only generically as performing the claimed selecting, decoding, dequantizing, and state-updating operations, and the rejection already explains that the processor is merely invoked as a tool to carry out the abstract steps rather than being improved in its own functioning. Likewise, Applicant’s asserted benefits of improved memory utilization, transmission efficiency, and computational efficiency are not recited in the claim as a particular machine improvement or other concrete technological application; rather, they are alleged results of performing the claimed mathematical and evaluative operations. The claim therefore remains directed to the abstract idea itself, implemented on a generic processor, rather than to a practical application of that idea.
Processor not a specific technical implementation (Remarks, page 15)
The Examiner respectfully disagrees with Applicant’s argument that the claim recites how the processor operates in a manner sufficient for eligibility. Although claim 1 recites a sequence of operations, the recited “how” is still only the abstract algorithm: selection of a set, decoding of an index, dequantization using that index, and state-based probability modeling. The claim does not recite a specialized decoder architecture, a non-generic memory arrangement, a particular circuit structure, or any other concrete implementation detail that would amount to an improvement in computer technology itself. Instead, the claim recites a processor configured to execute the abstract rules. As such, the processor remains generic and serves only as a tool for performing the recited mental and mathematical operations.
Step 2B — no significantly more / ordered combination (Remarks, page 15)
The Examiner respectfully disagrees with Applicant’s argument that the ordered combination supplies significantly more. Considered as an ordered combination, the claim still amounts to collecting and processing encoded parameter information according to mathematical rules for selection, state update, and probability-based decoding, using only a generic processor. The additional limitations do not add any unconventional hardware, transformation of physical subject matter, or other inventive concept beyond the abstract idea itself. Applicant’s characterization of the claimed flow as a “specific technical architecture” is not persuasive because the claim recites only a sequence of abstract data-processing rules and not a particular technological implementation beyond a generic processor. Accordingly, the claim does not amount to significantly more than the judicial exception, and the § 101 rejection is maintained.
Response to Applicant’s § 103 Arguments for Claims 1 and 22
Kasner as analogous art (Remarks, pages 21-22)
The Examiner respectfully disagrees with Applicant’s argument that Kasner is non-analogous art. Han is directed to compressing neural network weights by quantization and index-based representation in order to reduce storage and bandwidth while preserving accuracy, and thus squarely concerns compact representation and coding of sequences of parameter values. Kasner is likewise directed to quantization and coding of scalar sequences using state-dependent subset selection in a trellis framework, and expressly teaches that “TRELLIS CODED quantization (TCQ) has been shown to be an effective technique for quantizing memoryless sources with low to moderate complexity.” (Kasner, page 1, col. 1, paragraph 1) Both references are reasonably pertinent to the problem of how to represent quantized values more efficiently in a coded bitstream. This undermines Applicant’s position that a person of ordinary skill would not have looked to signal-coding references such as Kasner when addressing DNN compression.
Kasner does not require ECTCQ to be imported into Han (Remarks, page 22)
The Examiner respectfully disagrees with Applicant’s argument that Kasner’s teaching is inapplicable because it supposedly depends on ECTCQ. Kasner expressly presents UTCQ as an alternative that “requires neither stored codebooks nor a computationally intense codebook design algorithm.” (Kasner, abstract) Thus, the rejection does not rely on importing an ECTCQ system into Han; rather, it relies on Kasner’s own teaching that a trellis state can determine which superset of reconstruction levels is available and that a sequence of indices, together with the initial state, is sufficient to reproduce the chosen codewords. That is precisely the aspect relied upon in the rejection. Applicant’s argument regarding ECTCQ therefore does not address the actual teaching from Kasner that is being applied.
Motivation to combine Kasner with Han (Remarks, pages 22-23)
The Examiner respectfully disagrees with Applicant’s argument that the references cannot be combined because Han retrains centroids using one approach and Kasner describes codeword training differently. The rejection does not propose a wholesale substitution of one reference into the other, nor does claim 1 recite any specific training or retraining procedure. Rather, Han is relied upon for neural-network-weight compression using quantization and index-based representation of shared weights, and expressly teaches that “for each weight, we then need to store only a small index into a table of shared weights,” (Han, page 3, section 3, paragraph 2) while Kasner is relied upon for a state-driven subset-selection framework in which the current state determines the eligible reconstruction-level set and the chosen index drives the next state. A person of ordinary skill in the art would have understood that these teachings can be combined at the claimed level of generality without bodily incorporating every implementation detail of either reference. Applicant’s training-based argument is therefore not commensurate with the scope of the claim.
Applicability of Kasner (Remarks, pages 23-24)
The Examiner respectfully disagrees with Applicant’s argument that Kasner’s techniques would not have been applicable because neural network weights may have statistical properties different from image data. Obviousness does not require the references to disclose identical source distributions, only that a person of ordinary skill in the art would have had reason to apply the known coding technique to the claimed data-representation problem with a reasonable expectation that it would function for its intended purpose. Here, Kasner teaches a general trellis-coded quantization framework for scalar sequences in which state determines the available subset and indices permit reconstruction, and further teaches that “Training a modest subset of all codewords on the actual data being quantized allows UTCQ to adapt to source statistics and greatly improves performance.” (Kasner, page 1686, col. 2, paragraph 2) Han teaches representing neural network weights by quantized shared values and stored indices. Combining a known state-driven subset-selection quantization framework with a known index-based neural-network compression framework would have been a predictable use of known techniques to improve coding efficiency. Applicant’s argument regarding differing source statistics is therefore unpersuasive.
Sze lack of disclosure of a probability model dependent on a state (Remarks pages, 25)
The Examiner respectfully disagrees with Applicant’s argument that Sze does not disclose the claimed probability model because Sze’s “state” relates to probability estimation rather than to the trellis state that determines reconstruction-level subsets. This argument attacks Sze individually and does not address the rejection as a combination. The rejection relies on Kasner for the state-transition process that determines the applicable reconstruction-level set and updates the state based on the previously decoded index, and relies on Sze, as set forth in the rejection, for arithmetic coding using a probability model selected based on context/state information, where Sze explains that “for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate.” (Sze paragraph 6) The references need not disclose identical notions of “state” in isolation; it is sufficient that one reference teaches the state-driven quantization framework and another teaches state/context-dependent arithmetic coding. Thus, the rejection properly relies on the combination, not on Sze alone, and Applicant’s distinction is not persuasive.
Non-Obviousness and Advantages (Remarks, page 27)
The Examiner respectfully disagrees with Applicant’s assertion that there would have been no reasonable expectation of success or that the claimed arrangement achieves a unique advantage not suggested by the art. Han already teaches that neural-network weights may be quantized and represented by indices into shared values, and that doing so can “reduce the storage requirement of neural networks by 35× to 49× without affecting their accuracy.” (Han, abstract) Kasner teaches a state-driven quantization framework in which the current trellis state governs subset selection and index decoding enables reconstruction of the output sequence. The additional use of context/state-dependent arithmetic coding, as set forth in the rejection, would have been a predictable enhancement directed to the same goal of reducing coded bitstream size. Applicant’s alleged advantages therefore amount, at most, to expected benefits of applying known compression tools to the known problem of efficient coding of quantized neural-network parameters, and do not outweigh the rationale for combination set forth in the rejection.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 02/19 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claims 15-16, 18, and 28-29 are objected to because of the following informalities:
Claims 15 and 16 depend from cancelled claim 14.
Claim 18 depends from cancelled claim 17.
Claims 28 and 29 depend from cancelled claim 27.
Appropriate correction is required.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1-8, 11-13, and 15-22 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 3-8, 16, 18-21, 24-25, 34-27, and 46, filed on 11/11/25 of U.S. Patent Application No. 17843772. Although the claims at issue are not identical, they are not patentably distinct from each other because they recite limitations substantially similar to those recited in the copending application.
This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented.
Instant Application
Application No. 17843772
1. Apparatus for decoding neural network parameters, which define a neural network, from a data stream, comprising a processor configured to sequentially decode the neural network parameters by selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets depending on quantization indices decoded from the data stream for previous neural network parameters, decoding a quantization index for the current neural network parameter from the data stream, wherein the quantization index indicates one reconstruction level out of the selected set of reconstruction levels for the current neural network parameter, dequantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels that is indicated by the quantization index for the current neural network parameter.
1. (Currently Amended) Apparatus for decoding neural network parameters, which define a neural network, from a data stream, configured to sequentially decode the neural network parameters by selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets depending on quantization indices decoded from the data stream for previous neural network parameters, decoding a quantization index for the current neural network parameter from the data stream, wherein the quantization index indicates one reconstruction level out of the selected set of reconstruction levels for the current neural network parameter, dequantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels that is indicated by the quantization index for the current neural network parameter,
wherein the apparatus is configured to select, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets by means of a state transition process by determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter, and decode the quantization index for the current neural network parameter from the data stream using arithmetic coding using a probability model which depends on the state for the current neural network parameter and on the quantization index of previously decoded neural network parameters.
select, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets by means of a state transition process by determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter, and decode the quantization index for the current neural network parameter from the data stream using arithmetic coding using a probability model which depends on the state for the current neural network parameter.
2. Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two.
3. (Original) Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two.
3. Apparatus of claim 1, configured to parametrize the plurality of reconstruction level sets by way of a predetermined quantization step size and derive information on the predetermined quantization step size from the data stream.
4. (Original) Apparatus of claim 1, configured to parametrize the plurality of reconstruction level sets by way of a predetermined quantization step size and derive information on the predetermined quantization step size from the data stream.
4. Apparatus of claim 1, wherein the neural network comprises a one or more NN layers and the apparatus is configured to derive, for each NN layer, information on a predetermined quantization step size for the respective NN layer from the data stream, and parametrize, for each NN layer, the plurality of reconstruction level sets using the predetermined quantization step size derived for the respective NN layer so as to be used for dequantizing the neural network parameters belonging to the respective NN layer.
5. (Original) Apparatus of claim 1, wherein the neural network comprises a one or more NN layers and the apparatus is configured to derive, for each NN layer, information on a predetermined quantization step size for the respective NN layer from the data stream, and parametrize, for each NN layer, the plurality of reconstruction level sets using the predetermined quantization step size derived for the respective NN layer so as to be used for dequantizing the neural network parameters belonging to the respective NN layer.
5. Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two and the plurality of reconstruction level sets comprises a first reconstruction level set that comprises zero and even multiples of a predetermined quantization step size, and a second reconstruction level set that comprises zero and odd multiples of the predetermined quantization step size.
6. (Original) Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two and the plurality of reconstruction level sets comprises a first reconstruction level set that comprises zero and even multiples of a predetermined quantization step size, and a second reconstruction level set that comprises zero and odd multiples of the predetermined quantization step size.
6. Apparatus of claim 1, wherein all reconstruction levels of all reconstruction level sets represent integer multiples of a predetermined quantization step size, and the apparatus is configured to dequantize the neural network parameters by deriving, for each neural network parameter, an intermediate integer value depending on the selected reconstruction level set for the respective neural network parameter and the entropy decoded quantization index for the respective neural network parameter, andmultiplying, for each neural network parameter, the intermediate value for the respective neural network parameter with the predetermined quantization step size for the respective neural network parameter.
7. (Original) Apparatus of claim 1, wherein all reconstruction levels of all reconstruction level sets represent integer multiples of a predetermined quantization step size, and the apparatus is configured to dequantize the neural network parameters by deriving, for each neural network parameter, an intermediate integer value depending on the selected reconstruction level set for the respective neural network parameter and the entropy decoded quantization index for the respective neural network parameter, and multiplying, for each neural network parameter, the intermediate value for the respective neural network parameter with the predetermined quantization step size for the respective neural network parameter.
7. Apparatus of claim 6, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two and the apparatus is configured to derive the intermediate value for each neural network parameter by, if the selected reconstruction level set for the respective neural network parameter is a first set, multiply the quantization index for the respective neural network parameter by two to acquire the intermediate value for the respective neural network parameter; and if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is equal to zero, set the intermediate value for the respective neural network parameter equal to zero; and if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is greater than zero, multiply the quantization index for the respective neural network parameter by two and subtract one from the result of the multiplication to acquire the intermediate value for the respective neural network parameter; and if the selected reconstruction level set for a current neural network parameter is a second set and the quantization index for the respective neural network parameter is less than zero, multiply the quantization index for the respective neural network parameter by two and add one to the result of the multiplication to acquire the intermediate value for the respective neural network parameter.
8. (Currently Amended) Apparatus of claim 7, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two and the apparatus is configured to derive the intermediate value for each neural network parameter by, if the selected reconstruction level set for the respective neural network parameter is a first set, multiply the quantization index for the respective neural network parameter by two to acquire the intermediate value for the respective neural network parameter; and if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is equal to zero, set the intermediate value for the respective neural network parameter equal to zero; and if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is greater than zero, multiply the quantization index for the respective neural network parameter by two and subtract one from the result of the multiplication to acquire the intermediate value for the respective neural network parameter; and if the selected reconstruction level set for a current neural network parameter is a second set and the quantization index for the respective neural network parameter is less than zero, multiply the quantization index for the respective neural network parameter by two and add one to the result of the multiplication to acquire the intermediate value for the respective neural network parameter.
8. Apparatus of claim 1, wherein the apparatus is configured to select, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets by means of a state transition process by determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter.
16. (Original) Apparatus of claim 1, wherein the apparatus is configured to select, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets by means of a state transition process by determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter.
10. Apparatus of claim 8, configured to update the state for the subsequent neural network parameter using a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter.
18. (Original) Apparatus of claim 16, configured to update the state for the subsequent neural network parameter using a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter.
11. Apparatus of claim 8, wherein the state transition process is configured to transition between four or eight possible states.
19. (Original) Apparatus of claim 16, wherein the state transition process is configured to transition between four or eight possible states.
12. Apparatus of claim 8, configured to transition, in the state transition process, between an even number of possible states and the number of reconstruction level sets of the plurality of reconstruction level sets is two, wherein the determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets depending on the state associated with the current neural network parameter determines a first reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a first half of the even number of possible states, and a second reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a second half of the even number of possible states.
20. (Currently Amended) Apparatus of claim 16, configured to transition, in the state transition process, between an even number of possible states and the number of reconstruction level sets of the plurality of reconstruction level sets is two, wherein the determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on the state associated with the current neural network parameter determines a first reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a first half of the even number of possible states, and a second reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a second half of the even number of possible states.
13. Apparatus of claim 8, configured to perform the update of the state by means of a transition table which maps a combination of the state and a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter onto a further state associated with the subsequent neural network parameter.
21. Apparatus of claim 16, configured to perform the update of the state by means of a transition table which maps a combination of the state and a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter onto a further state associated with the subsequent neural network parameter.
15. Apparatus of claim 14, configured to decode the quantization index for the current neural network parameter from the data stream using binary arithmetic coding by using the probability model which depends on the state for the current neural network parameter for at least one bin of a binarization of the quantization index.
24. (Original) Apparatus of claim 23, configured to decode the quantization index for the current neural network parameter from the data stream using binary arithmetic coding by using the probability model which depends on the state for the current neural network parameter for at least one bin of a binarization of the quantization index.
16. Apparatus of claim 14, wherein the at least one bin comprises a significance bin indicative of the quantization index of the current neural network parameter being equal to zero or not.
25. (Original) Apparatus of claim 23, wherein the at least one bin comprises a significance bin indicative of the quantization index of the current neural network parameter being equal to zero or not.
17. Apparatus of claim 14, wherein the probability model additionally depends on the quantization index of previously decoded neural network parameters.
34. (Currently Amended) Apparatus of claim 23, wherein the probability model additionally depends on the quantization index of previously decoded neural network parameters.
18. Apparatus of claim 17, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, a subset of probability models out of a plurality of probability models and select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters.
35. (Original) Apparatus of claim 34, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, a subset of probability models out of a plurality of probability models and select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters.
19. Apparatus of claim 18, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, the subset of probability models out of the plurality of probability models in a manner so that a subset preselected for a first state or reconstruction levels set is disjoint to a subset preselected for any other state or reconstruction levels set.
36. Apparatus of claim 35, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, the subset of probability models out of the plurality of probability models in a manner so that a subset preselected for a first state or reconstruction levels set is disjoint to a subset preselected for any other state or reconstruction levels set.
20. Apparatus of claim 18, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to.
37. (Original) Apparatus of claim 35, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to.
21. Apparatus of claim 18, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on a characteristic of the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, the characteristic comprising on or more of the signs of non-zero quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, the number of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to a difference between a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to,and the number of quantization indices of the previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero.
38. (Original) Apparatus of claim 35, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on a characteristic of the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, the characteristic comprising on or more of the signs of non-zero quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, the number of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to a difference between a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and the number of quantization indices of the previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero.
22. Apparatus for encoding neural network parameters, which define a neural network, into a data stream, comprising a processor configured to sequentially encode the neural network parameters by selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets depending on quantization indices encoded into the data stream for previously encoded neural network parameters ,quantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels, and encoding a quantization index for the current neural network parameter that indicates the one reconstruction level onto which the quantization index for the current neural network parameter is quantized into the data stream.
46. (Original) Apparatus for encoding neural network parameters, which define a neural network, into a data stream, configured to sequentially encode the neural network parameters by selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets depending on quantization indices encoded into the data stream for previously encoded neural network parameters, quantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels, and encoding a quantization index for the current neural network parameter that indicates the one reconstruction level onto which the quantization index for the current neural network parameter is quantized into the data stream.
wherein the apparatus is configured to select, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets by means of a state transition process by determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, and updating the state for a subsequent neural network parameter depending on the quantization index encoded into the data stream for the immediately preceding neural network parameter, and encode the quantization index for the current neural network parameter into the data stream using arithmetic coding using a probability model which depends on the state for the current neural network parameter and on the quantization index of previously decoded neural network parameters.
select, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets by means of a state transition process by determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, and updating the state for a subsequent neural network parameter depending on the quantization index encoded into the data stream for the immediately preceding neural network parameter, and encode the quantization index for the current neural network parameter into the data stream using arithmetic coding using a probability model which depends on the state for the current neural network parameter.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-13, 15-16, 18-26, and 28-30 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Statutory Categories
Claims 1-13, 15-16, 18-21 are directed to an apparatus.
Claim 22 is directed to an apparatus.
Claims 23-26 and 28-30 are directed to an apparatus.
Independent Claim 1
Step 2A Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes. Independent claim 1 recites limitations that are abstract ideas in the form of mental processes:
Claim 1 recites:
sequentially decode the neural network parameters by selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets depending on quantization indices decoded from the data stream for previous neural network parameters, (simply selecting a set is interpreted to be a mental proves of evaluation which can reasonably be performed in human mind)
decoding a quantization index for the current neural network parameter from the data stream, wherein the quantization index indicates one reconstruction level out of the selected set of reconstruction levels for the current neural network parameter, (this limitation merely comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), page 27, section 2.2.2 of this application’s specification outlines the mathematical procedure for this step)
dequantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels that is indicated by the quantization index for the current neural network parameter. (this limitation merely comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), page 27, section 2.2.2 of this application’s specification outlines the mathematical procedure for this step)
select, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets by means of a state transition process by determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, and (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter, and (an updating process that is being considered as a mathematical implementation, see pages 52-3, for support for the mathematical process of this limitation)
decode the quantization index for the current neural network parameter from the data stream using arithmetic coding using a probability model which depends on the state for the current neural network parameter and on the quantization index of previously encoded neural network parameters. (a decoding process that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 27, section 2.2.2 for support for the mathematical process of this limitation)
This claim further recites the following additional elements for the purposes of Step 2A Prong Two analysis:
Apparatus for decoding neural network parameters, which define a neural network, from a data stream, comprising a processor configured to (this limitation invokes computers or other machinery merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
wherein the apparatus is configured to (this limitation invokes computers or other machinery merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
The additional limitations fail step 2A Prong 2 of the 101 analysis because they do not transform the claim into a practical application. These limitations are too abstract or lack technical improvement that would make the concept practically useful. Without clear utility or integration into a specific field, the claim does not relate to any particular application. It does not meet the requirements of Step 2A Prong 2, as it fails to make the concept meaningfully applicable in practice. Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea.
This claim recites the following additional elements for the purposes of Step 2B analysis:
Apparatus for decoding neural network parameters, which define a neural network, from a data stream, comprising a processor configured to (this limitation invokes computers or other machinery merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
wherein the apparatus is configured to (this limitation invokes computers or other machinery merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
The claim also fails Step 2B of the analysis because the additional limitations do not amount to significantly more than the abstract idea itself. The additional limitations do not enhance the claim in a way that would move it beyond its abstract ideas as they minimally elaborate on the core concept without adding any inventive or technical substance. Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Dependents of Claim 1
The remaining dependent claims corresponding to independent claim 1 do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. The analysis of which is shown below:
The claims below recite additional limitations which fail step 2A Prong 2 of the 101 analysis because they do not transform the claim into a practical application. These limitations are too abstract or lack technical improvement that would make the concept practically useful. Without clear utility or integration into a specific field, the claim does not relate to any particular application. It does not meet the requirements of Step 2A Prong 2, as it fails to make the concept meaningfully applicable in practice.
The claims also fails Step 2B of the analysis because the additional limitations do not amount to significantly more than the abstract idea itself. The additional limitations do not enhance the claim in a way that would move it beyond its abstract ideas as they minimally elaborate on the core concept without adding any inventive or technical substance. The claims are unpatentable.
Claim 2 recites the further limitation of:
Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two. (further specifying a parameter for aforementioned mathematical process is still being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations. Page 23, section 2.1, provides support for the mathematical process of this claim.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 3 recites the further limitation of:
Apparatus of claim 1, configured to parametrize the plurality of reconstruction level sets by way of a predetermined quantization step size and derive information on the predetermined quantization step size from the data stream. (further defining parameterization and step size for this process is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, page 23, section 2.1, provides support for the mathematical process of this claim)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 4 recites the further limitation of:
Apparatus of claim 1, wherein the neural network comprises a one or more NN layers and the apparatus is configured to (under analysis of step 2A prong II and step 2B, a neural network comprising one or more neural network layers is being considered as mere instructions to apply an exception, see MPEP 2106.05(f))
derive, for each NN layer, information on a predetermined quantization step size for the respective NN layer from the data stream, (computing/obtaining per-layer step sizes is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, page 23, section 2.1, provides support for the mathematical process of this claim)
and parametrize, for each NN layer, the plurality of reconstruction level sets using the predetermined quantization step size derived for the respective NN layer so as to be used for dequantizing the neural network parameters belonging to the respective NN layer. (a parameterization per layer that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, page 27, section 2.2.2, provides support for the mathematical process of this claim)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 5 recites the further limitation of:
Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two and the plurality of reconstruction level sets comprises (further specifying a parameter for aforementioned mathematical process is still being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations. Page 23, section 2.1, provides support for the mathematical process of this claim.)
a first reconstruction level set that comprises zero and even multiples of a predetermined quantization step size, (further defining these construction sets is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations. Page 33, rule c, provides support for the mathematical process of this claim.)
and a second reconstruction level set that comprises zero and odd multiples of the predetermined quantization step size. (further defining these construction sets is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations. Page 33, rule c, provides support for the mathematical process of this claim.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 6 recites the further limitation of:
Apparatus of claim 1, wherein all reconstruction levels of all reconstruction level sets represent integer multiples of a predetermined quantization step size, and the apparatus is configured to dequantize the neural network parameters by (further defining step sizes is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 33, rules a-d for support for the mathematical process of this limitation)
deriving, for each neural network parameter, an intermediate integer value depending on the selected reconstruction level set for the respective neural network parameter and the entropy decoded quantization index for the respective neural network parameter, (the process of derivation is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 33, rules a-d for support for the mathematical process of this limitation)
and multiplying, for each neural network parameter, the intermediate value for the respective neural network parameter with the predetermined quantization step size for the respective neural network parameter. (a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 25, line 5-6, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 7 recites the further limitation of:
Apparatus of claim 6, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two and the apparatus is configured to derive the intermediate value for each neural network parameter by, if the selected reconstruction level set for the respective neural network parameter is a first set, multiply the quantization index for the respective neural network parameter by two to acquire the intermediate value for the respective neural network parameter; and(an arithmetic rule that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations)
if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is equal to zero, set the intermediate value for the respective neural network parameter equal to zero; and (an arithmetic rule that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations)
if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is greater than zero, multiply the quantization index for the respective neural network parameter by two and subtract one from the result of the multiplication to acquire the intermediate value for the respective neural network parameter; and (an arithmetic rule that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations)
if the selected reconstruction level set for a current neural network parameter is a second set and the quantization index for the respective neural network parameter is less than zero, multiply the quantization index for the respective neural network parameter by two and add one to the result of the multiplication to acquire the intermediate value for the respective neural network parameter. (an arithmetic rule that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 8 recites the further limitation of:
Apparatus of claim 1, wherein the apparatus is configured to select, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets by means of a state transition process by determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, and (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter. (an updating process, stated at a high level of generality, that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 27, section 2.2.2, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 9 recites the further limitation of:
Apparatus of claim 8, configured to update the state for the subsequent neural network parameter using a binary function of the quantization index decoded from the data stream for the immediately preceding neural network parameter. (an updating process that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 26, section 2.2.1, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 10 recites the further limitation of:
Apparatus of claim 8, configured to update the state for the subsequent neural network parameter using a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter. (an updating process that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 27, section 2.2.2, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 11 recites the further limitation of:
Apparatus of claim 8, wherein the state transition process is configured to transition between four or eight possible states. (further defining the state process is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 52-53 of the spec, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 12 recites the further limitation of:
Apparatus of claim 8, configured to transition, in the state transition process, between an even number of possible states and the number of reconstruction level sets of the plurality of reconstruction level sets is two, (further parameterizing the state transition process is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see pages 52-53, for support for the mathematical process of this limitation)
wherein the determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction levels sets depending on the state associated with the current neural network parameter determines a first reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a first half of the even number of possible states, and a second reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a second half of the even number of possible states. (a determination that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 13 recites the further limitation of:
Apparatus of claim 8, configured to perform the update of the state by means of a transition table which maps a combination of the state and a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter onto a further state associated with the subsequent neural network parameter. (an updating process that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see pages 52-3, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 15 recites the further limitation of:
Apparatus of claim 14, configured to decode the quantization index for the current neural network parameter from the data stream using binary arithmetic coding by using the probability model which depends on the state for the current neural network parameter for at least one bin of a binarization of the quantization index. (a decoding process that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 27, section 2.2.2., for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 16 recites the further limitation of:
Apparatus of claim 14, wherein the at least one bin comprises a significance bin indicative of the quantization index of the current neural network parameter being equal to zero or not. (further defining a significant bit for this process is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see pages 26, section 2.2.1, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 18 recites the further limitation of:
Apparatus of claim 17, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, a subset of probability models out of a plurality of probability models and (a simple preselection of a set of models is being considered as a mental process of evaluation which can reasonably be performed in one’s mind)
select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters. (a selection of a model based on a ascertained value is being considered as a mental process of evaluation which can reasonably be performed in one’s mind)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 19 recites the further limitation of:
Apparatus of claim 18, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, the subset of probability models out of the plurality of probability models in a manner so that a subset preselected for a first state or reconstruction levels set is disjoint to a subset preselected for any other state or reconstruction levels set (a selection of a model between inherently different sets is being considered as a mental process of evaluation which can reasonably be performed in one’s mind)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 20 recites the further limitation of:
Apparatus of claim 18, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to. (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 21 recites the further limitation of:
Apparatus of claim 18, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on a characteristic of the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, the characteristic comprising on or more of the signs of non-zero quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
the number of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero (a counting statistic that (a counting statistic that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations see page 27, section 2.2.2, for support for the mathematical process of this limitation)
a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to (a sum that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations)
a difference between a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and the number of quantization indices of the previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero. (a difference that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Independent Claim 22
Step 2A Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes. Independent claim 22 recites limitations that are abstract ideas in the form of mental processes:
Claim 22 recites:
sequentially encode the neural network parameters by selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets depending on quantization indices encoded into the data stream for previously encoded neural network parameters, (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
quantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels, and (this limitation merely comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see page 26, section 2.2.1, for support for the mathematical process of this limitation)
encoding a quantization index for the current neural network parameter that indicates the one reconstruction level onto which the quantization index for the current neural network parameter is quantized into the data stream. (this limitation merely comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see page 26, section 2.2.1, for support for the mathematical process of this limitation)
select, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets by means of a state transition process by determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, and (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter, and (an updating process that is being considered as a mathematical implementation, see pages 52-3, for support for the mathematical process of this limitation)
encode the quantization index for the current neural network parameter from the data stream using arithmetic coding using a probability model which depends on the state for the current neural network parameter and on the quantization index of previously decoded neural network parameters. (a decoding process that is being considered as a mathematical implementation involving mathematical concepts, algorithms, or calculations, see page 27, section 2.2.2 for support for the mathematical process of this limitation)
This claim further recites the following additional elements for the purposes of Step 2A Prong Two analysis:
Apparatus for encoding neural network parameters, which define a neural network, into a data stream, comprising a processor configured to (this limitation invokes computers or other machinery merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
The additional limitations fail step 2A Prong 2 of the 101 analysis because they do not transform the claim into a practical application. These limitations are too abstract or lack technical improvement that would make the concept practically useful. Without clear utility or integration into a specific field, the claim does not relate to any particular application. It does not meet the requirements of Step 2A Prong 2, as it fails to make the concept meaningfully applicable in practice.
This claim recites the following additional elements for the purposes of Step 2B analysis:
Apparatus for encoding neural network parameters, which define a neural network, into a data stream, comprising a processor configured to (this limitation invokes computers or other machinery merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
The claim also fails Step 2B of the analysis because the additional limitations do not amount to significantly more than the abstract idea itself. The additional limitations do not enhance the claim in a way that would move it beyond its abstract ideas as they minimally elaborate on the core concept without adding any inventive or technical substance. The claim is unpatentable.
Independent Claim 23
Step 2A Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes. Independent claim 23 recites limitations that are abstract ideas in the form of mental processes:
Claim 23 recites:
derive first neural network parameters for a first reconstruction layer to yield, per neural network parameter, a first- reconstruction-layer neural network parameter value, (a derivation stated at such a level of generality is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
decode second neural network parameters for a second reconstruction layer from a data stream to yield, per neural network parameter, a second-reconstruction-layer neural network parameter value, (this limitation merely comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see page 26, section 2.2.1, for support for the mathematical process of this limitation)
and reconstruct the neural network parameters by, for each neural network parameter, combining the first-reconstruction-layer neural network parameter value and the second-reconstruction-layer neural network parameter value. (this limitation merely comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see page 27, section 2.2.2, for support for the mathematical process of this limitation)
decode the second-reconstruction-layer neural network parameter value from the data stream by context-adaptive entropy decoding, (a decoding process using a probability model is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see pages 27, section 2.2.2, for support for the mathematical process of this limitation)
selecting a probability context set out of a collection of probability context sets depending on the first-reconstruction-layer neural network parameter value, (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
selecting a probability context to be used out of the selected probability context set depending on the first-reconstruction-layer neural network parameter value. (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
This claim further recites the following additional elements for the purposes of Step 2A Prong Two analysis:
Apparatus for reconstructing neural network parameters, which define a neural network, comprising a processor configured to (this limitation invokes computers or other machinery merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
The additional limitations fail step 2A Prong 2 of the 101 analysis because they do not transform the claim into a practical application. These limitations are too abstract or lack technical improvement that would make the concept practically useful. Without clear utility or integration into a specific field, the claim does not relate to any particular application. It does not meet the requirements of Step 2A Prong 2, as it fails to make the concept meaningfully applicable in practice.
This claim recites the following additional elements for the purposes of Step 2B analysis:
Apparatus for reconstructing neural network parameters, which define a neural network, comprising a processor configured to (this limitation invokes computers or other machinery merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
The claim also fails Step 2B of the analysis because the additional limitations do not amount to significantly more than the abstract idea itself. The additional limitations do not enhance the claim in a way that would move it beyond its abstract ideas as they minimally elaborate on the core concept without adding any inventive or technical substance. The claim is unpatentable.
Dependents of Claim 23
The remaining dependent claims corresponding to independent claim 23 do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. The analysis of which is shown below:
The claims below recite additional limitations which fail step 2A Prong 2 of the 101 analysis because they do not transform the claim into a practical application. These limitations are too abstract or lack technical improvement that would make the concept practically useful. Without clear utility or integration into a specific field, the claim does not relate to any particular application. It does not meet the requirements of Step 2A Prong 2, as it fails to make the concept meaningfully applicable in practice.
The claims also fails Step 2B of the analysis because the additional limitations do not amount to significantly more than the abstract idea itself. The additional limitations do not enhance the claim in a way that would move it beyond its abstract ideas as they minimally elaborate on the core concept without adding any inventive or technical substance. The claims are unpatentable.
Claim 24 recites the further limitation of:
Apparatus of claim 23, configured to Decode the first neural network parameters for the first reconstruction layer from the data stream or from a separate data stream, and (this limitation merely comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see page 27, section 2.2.2, for support for the mathematical process of this limitation)
decode the second neural network parameters for the second reconstruction layer from the data stream by context-adaptive entropy decoding using separate probability contexts for the first and second reconstruction layers. (this limitation merely comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see page 27, section 2.2.2, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 25 recites the further limitation of:
Apparatus of claim 23, configured to reconstruct the neural network parameters by a parameter wise sum (a sum that comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see pages 23-24, and section 2.2.1-2.2.2, for support for the mathematical process of this limitation)
or parameter wise product of, per neural network parameter, the first-reconstruction-layer neural network parameter value and the second-reconstruction-layer neural network parameter value. (a product that comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see pages 23-24, and section 2.2.1-2.2.2, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 26 recites the further limitation of:
Apparatus of claim 23, configured to Decode the second-reconstruction-layer neural network parameter value from the data stream by context-adaptive entropy decoding using a probability model which depends on the first-reconstruction-layer neural network parameter value. (a decoding process using a probability model is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see pages 27, section 2.2.2, for support for the mathematical process of this limitation)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 28 recites the further limitation of:
Apparatus of claim 27, wherein the collection of probability context sets comprises three probability context sets, and the apparatus is configured to (a numerical parameterization that is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see pages 30-31, for support for the mathematical process of this limitation)
select a first probability context set out of the collection of probability context sets as the selected probability context set if the first-reconstruction-layer neural network parameter value is negative, (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
select a second probability context set out of the collection of probability context sets as the selected probability context set if the first-reconstruction-layer neural network parameter value is positive, (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
select a third probability context set out of the collection of probability context sets as the selected probability context set if the first-reconstruction-layer neural network parameter value is zero. (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 29 recites the further limitation of:
Apparatus of claim 27, wherein the collection of probability context sets comprises two probability context sets, and the apparatus is configured to (a numerical parameterization that is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see pages 30-31, for support for the mathematical process of this limitation)
select a first probability context set out of the collection of probability context sets as the selected probability context set if the first-reconstruction-layer neural network parameter value is greater than a predetermined value, (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
and select a second probability context set out of the collection of probability context sets as the selected probability context set if the first-reconstruction-layer neural network parameter value is not greater than the predetermined value, (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
or select the first probability context set out of the collection of probability context sets as the selected probability context set if an absolute value of the first-reconstruction-layer neural network parameter value is greater than the predetermined value, (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
and select the second probability context set out of the collection of probability context sets as the selected probability context set if the absolute value of the first-reconstruction-layer neural network parameter value is not greater than the predetermined value (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim 30 recites the further limitation of:
Apparatus for encoding neural network parameters, which define a neural network, by using first neural network parameters for a first reconstruction layer which comprise, per neural network parameter, a first- reconstruction-layer neural network parameter value, and the apparatus comprising a processor being configured to (Under step 2A prong II and step 2B, this limitation invokes computers or other machinery merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
encode second neural network parameters for a second reconstruction layer into a data stream, which comprise, per neural network parameter, a second-reconstruction-layer neural network parameter value, (an encoding process that comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see pages 26, section 2.2.1, for support for the mathematical process of this limitation)
wherein the neural network parameters are reconstructible by, for each neural network parameter, combining the first-reconstruction-layer neural network parameter value and the second-reconstruction-layer neural network parameter value. (a sum that comprises a mathematical analysis of data and is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see page 27, section 2.2.2., for support for the mathematical process of this limitation)
encode the second-reconstruction-layer neural network parameter value from the data stream by context-adaptive entropy encoding, (a encoding process using a probability model is being considered as directed to a mathematical concept, see MPEP 2106.04(a), see pages 27, section 2.2.2, for support for the mathematical process of this limitation)
selecting a probability context set out of a collection of probability context sets depending on the first-reconstruction-layer neural network parameter value, (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
selecting a probability context to be used out of the selected probability context set depending on the first-reconstruction-layer neural network parameter value. (a selection that is being considered as a mental process of evaluation which can reasonably be performed in mind or with aid of pen and paper)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 2, 3, 4, 6, 8, 11, 15-16, 18-20 and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Han et al., (Han, S., Mao, H., & Dally, W. J. (2015) Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149.), hereafter referred to as Han, in view of Kasner et al. (Kasner, J. H., Marcellin, M. W., & Hunt, B. R. (1999). Universal trellis coded quantization. IEEE Transactions on Image Processing, 8(12), 1677-1687.), hereafter referred to as Kasner and in further view of Sze et al. (US20130272389A1), hereafter referred to as Sze.
Claim 1: Han teaches the following limitations:
Apparatus for decoding neural network parameters, which define a neural network, from a data stream, comprising a processor configured to sequentially decode the neural network parameters (Han, abstract, “Our method first prunes the network by learning only the important connections. Next, we quantize the weights to enforce weight sharing, finally, we apply Huffman coding. After the first two steps we retrain the network to fine tune the remaining connections and the quantized centroids.”, Han describes decoding quantized indices from neural network connections.)
Kasner, in the same field of quantization methods, teaches the following limitations which the above fails to teach:
by selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets depending on quantization indices decoded from the data stream for previous neural network parameters, (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
. Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen”, In Kasner, each trellis state is associated with access to one of two supersets of codewords (each superset being a subset of the overall uniform codebook). At any given state, the next reconstruction value must be chosen from one of these two supersets; and the state itself is updated based on the previously encoded index sequence. Thus, the superset of codewords at the current state is what is being interpreted as the “set of reconstruction levels (one reconstruction-level set)”, the collection of supersets over the trellis is the “plurality of reconstruction level sets,” and the fact that the next state (and thus which superset applies to the current sample) depends on the previously chosen indices is what is being interpreted as “selecting … a set of reconstruction levels … depending on quantization indices decoded from the data stream for previous neural network parameters.” )
decoding a quantization index for the current neural network parameter from the data stream, wherein the quantization index indicates one reconstruction level out of the selected set of reconstruction levels for the current neural network parameter, (Kasner, page 1678, col. 2, paragraph 1, “The UTCQ quantizer returns the S0 indices and the negative of the S1 indices, allowing one probability model to be used for entropy coding. The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly”, In Kasner, for each source sample, the UTCQ quantizer returns an index within the currently active superset; the decoder, given the current state and the received index, looks up exactly one reconstruction codeword in that superset. Hence, the superset’s codewords are being interpreted as the “reconstruction levels,” and the index within that superset is the claimed “quantization index” that uniquely points to one codeword (one reconstruction level) out of the set associated with the current state.)
dequantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels that is indicated by the quantization index for the current neural network parameter. (Kasner, page 1678, col. 2, paragraph 2, “During dequantization, two types of reconstruction levels are employed, uniform and trained. For
PNG
media_image2.png
18
97
media_image2.png
Greyscale
, uniform levels are used (i.e., the codeword is the center of the quantization cell). The remaining codewords are trained on the source data itself, except CW0 which is typically set to 0. The trained codeword
PNG
media_image3.png
18
118
media_image3.png
Greyscale
, is determined by taking the sample mean of all source symbols that map to
PNG
media_image4.png
17
45
media_image4.png
Greyscale
and the negative of all source symbols mapping to
PNG
media_image5.png
15
58
media_image5.png
Greyscale
”, Kasner’s decoder in whichever selected subset (S0 or S1), performs a lookup of the exact reconstruction level corresponding to the decoded index i or -i.)
wherein the apparatus is configured to select, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets by means of a state transition process (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work… Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
. Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen [3].”, Kasner’s TCQ’s trellis is a state transition process that selects which codebook subset (superset) to use for each coefficient based on the current state. Kasner explicitly shows an eight-state trellis and explains that “at any given trellis state, the next codeword must come from one of two supersets,” and that given an initial state and a sequence of indices the decoder can reconstruct the sequence of chosen codewords. The movement from state to state along the trellis edges as each new symbol/index is processed is the claimed “state transition process.” At each step, the current state determines which superset (reconstruction-level set) is used, and the next state is determined by the current state and the chosen index, so the trellis operation as a whole is being interpreted as “selecting … the set of quantization levels … by means of a state transition process.” )
by determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
.”, The excerpt shows that which superset (the reconstruction level set comprising the set of quantization levels) is valid depends on the current state.)
and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter, (Kasner, page 1677, col. 2, paragraph 2, “Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen [3].”, the decoder updates its state after each decoded index so that the next state (and thus the next superset) is determined by the previous index.)
A person of ordinary skill in the art (POSITA) before the effective filing date of the claimed invention would have recognized that Kasner’s trellis-coded quantization is equally applicable to any sequence of scalar parameters, including the neural network weights of Han. Han’s Deep Compression already teaches a quantization technique that uses a learned codebook of centroids per layer and represents each weight by an index into that codebook. Kasner’s TCQ similarly uses codebooks portioned into subsets and represents each sample by a subset-specific index. To combine the two methods one would swap the codebook of Kasner’s TCQ with the codebook of Han and then, during decoding, use Kasner’s trellis-state (updated by prior weight indices) to select the correct subset before looking up the index for the decoded value. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Han (i.e. deep neural network quantization methos) by incorporating the teachings of Kasner (i.e. index based Trellis-Coded Quantization methods). A motivation of which is to provide a codebook quantization technique without needing additional storage and training of codebooks. (Kasner, page 1687, “The advantages of UTCQ are simplicity and flexibility. Unlike previous ECTCQ systems, no prior codebook training is needed and no codebooks are stored. We have shown that the distortion-rate performance of UTCQ is comparable with that of optimal ECTCQ for memoryless sources at most encoding rates.”)
Sze, in the same field of data encoding, teaches the following limitations which the above fails to teach:
and decode the quantization index for the current neural network parameter from the data stream using arithmetic coding using a probability model which depends on the state for the current neural network parameter and on the quantization index of the previously decoded neural network parameters. (Sze, paragraph 6, “CABAC has multiple probability modes for different contexts. It first converts all non-binary symbols to binary symbols referred to as bins. Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate. Arithmetic coding is then applied to compress the data.”, CABAC’s arithmetic coder selects a probability model (context) based on the current state, and then arithmetically decodes each bin under that model, teaching state-dependent arithmetic decoding.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wang (i.e. deep neural network quantization methos) and Kasner by incorporating the teachings of Sze (i.e. Context-Aware Binary Arithmetic Coding methods). Kasner further recognizes that different sets of indices (corresponding to different supersets) may be entropy coded with different coding models (e.g., different arithmetic models or different Huffman tables), supporting selection of the probability model based on the state/superset associated with the current parameter. A motivation of which is to provide an entropy coding technique (arithmetic coding with context/probability models) that would further reduce the overall bitstream size for the quantization indices.(Sze, paragraph 6, “CABAC is an inherently lossless compression technique notable for providing considerably better compression than most other encoding algorithms used in video encoding at the cost of increased complexity.”, by adopting CABAC for entropy-coding Han’s quantization indices, a POSITA would directly reduce the overall bitstream size of a neural networks parameters.)
Claim 2: Han, Kasner, and Sze teaches the limitations of claim 1, Kasner further teaches:
Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two. (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
. Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen”, at each step exactly two supersets (reconstruction level sets) are available.)
Claim 3: Han, Kasner, and Sze teaches the limitations of claim 1, Kasner further teaches:
Apparatus of claim 1, configured to parametrize the plurality of reconstruction level sets by way of a predetermined quantization step size and derive information on the predetermined quantization step size from the data stream. (Kasner, page 1678, col. 2, paragraph 1, “The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly…. The quantization thresholds are simply the midpoints between the reconstruction levels within a subset. This allows for fast computation of superset indices requiring only scaling and rounding. No thresholds need to be precomputed, nor is a binary tree search necessary. For a given trellis, the encoder is completely characterized by the stepsize parameter Δ .", this discloses the reconstruction level sets are explicitly parameterized using a quantization step size gained from a stream of data (index stream).
“During dequantization, two types of reconstruction levels are employed, uniform and trained. For
PNG
media_image2.png
18
97
media_image2.png
Greyscale
, uniform levels are used (i.e., the codeword is the center of the quantization cell)… The remaining codewords are trained on the source data itself”, the reconstruction level sets (which comprise the step size) are derived from the index stream.)
Claim 4: Han, Kasner, and Sze teaches the limitations of claim 1, Kasner further teaches:
Apparatus of claim 1, wherein the neural network comprises a one or more NN layers and the apparatus is configured to derive, for each NN layer, information on a predetermined quantization step size for the respective NN layer from the data stream, (Kasner, page 1678, col. 2, paragraph 1, “The UTCQ quantizer returns the S0 indices and the negative of the S1 indices, allowing one probability model to be used for entropy coding. The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly”, Kasner states that, “for a given trellis, the encoder is completely characterized by the stepsize parameter Δ,” and that UTCQ uses uniform thresholds and reconstruction levels based on this Δ. The value of Δ (or its encoded representation) is the “information on a predetermined quantization step size” that the decoder must know or derive from the bitstream to reconstruct the uniform codebook. In the combined mapping with Wang, Wang supplies the multiple NN layers; for each such layer, the trellis-based quantizer operates with a chosen Δ, and the Δ value signaled or implied in the data stream for that layer is what is being interpreted as “information on a predetermined quantization step size for the respective NN layer … derived from the data stream.” The decoder can recover the index stream by tracking state and negating codewords, and derives, quantization information like step sizes from this data stream.)
and parametrize, for each NN layer, the plurality of reconstruction level sets using the predetermined quantization step size derived for the respective NN layer so as to be used for dequantizing the neural network parameters belonging to the respective NN layer. (Kasner, page 1678, col. 2, paragraph 1, “The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly…. The quantization thresholds are simply the midpoints between the reconstruction levels within a subset. This allows for fast computation of superset indices requiring only scaling and rounding. No thresholds need to be precomputed, nor is a binary tree search necessary. For a given trellis, the encoder is completely characterized by the stepsize parameter Δ .",
“During dequantization, two types of reconstruction levels are employed, uniform and trained. For
PNG
media_image2.png
18
97
media_image2.png
Greyscale
, uniform levels are used (i.e., the codeword is the center of the quantization cell)… The remaining codewords are trained on the source data itself”, the quantization thresholds are midpoints between reconstruction levels and that the encoder is completely characterized by the step size parameter Δ. This directly teaches that the reconstruction level sets are parameterized using the predetermined step size, which enables dequantization tailored per NN layer.)
Claim 6: Han, Kasner, and Sze teaches the limitations of claim 1, Kasner further teaches:
Apparatus of claim 1, wherein all reconstruction levels of all reconstruction level sets represent integer multiples of a predetermined quantization step size, and the apparatus is configured to dequantize the neural network parameters by (Kasner, page 1678, col. 2, paragraph 1, “UTCQ uses uniform thresholds and reconstruction levels at the encoder. The quantization thresholds are simply the midpoints between the reconstruction levels within a subset. This allows for fast computation of superset indices requiring only scaling and rounding. No thresholds need to be precomputed, nor is a binary tree search necessary. For a given trellis, the encoder is completely characterized by the stepsize parameter Δ.”, Kasner explains that UTCQ uses uniform thresholds and reconstruction levels and that the encoder is characterized by a stepsize parameter Δ. In a uniform scalar quantizer, the reconstruction values are positioned at regularly spaced points (centers of quantization cells) on the real line, each separated by Δ; these regular positions can be expressed as k·Δ for integer k. Thus, the uniform reconstruction levels in the UTCQ codebook are being interpreted as the claimed “reconstruction levels,” Δ is the “predetermined quantization step size,” and the fact that these levels lie on a uniform grid defined by Δ is what supports the statement that all reconstruction levels of all reconstruction level sets represent integer multiples of Δ. A uniform threshold scheme places reconstruction levels at exact multples of the uniform step size Δ (i.e. midpoints at k * Δ) so all codewords in every subset are integer-multiples of Δ.)
PNG
media_image6.png
87
275
media_image6.png
Greyscale
Figure 2 of Kasner
deriving, for each neural network parameter, an intermediate integer value depending on the selected reconstruction level set for the respective neural network parameter and the entropy decoded quantization index for the respective neural network parameter, (Kasner, page 1679, col. 2, paragraph 3, “Given a source sample to quantize, subset quantization indices may be computed directly. Given a quantization index, the reconstruction level (for uniform codewords) may be computed.”,
Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
. Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen”, In UTCQ, Fig. 2 shows a uniform codebook partitioned into subsets, and Kasner explains that “a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen” (Kasner, p.1677). Kasner also states that “given a quantization index, the reconstruction level (for uniform codewords) may be computed” (Kasner, p.1679). Under a broadest reasonable interpretation, the index that is entropy-decoded for the active superset is the claimed “entropy decoded quantization index.” Because the underlying uniform codebook is ordered and parameterized by the step size Δ, that index corresponds to a particular uniform codeword location on the real line, which is treated as the claimed “intermediate integer value” that identifies which Δ-spaced reconstruction level is to be used.)
and multiplying, for each neural network parameter, the intermediate value for the respective neural network parameter with the predetermined quantization step size for the respective neural network parameter. (Kasner, page 1678, col. 2, paragraph 2, “During dequantization, two types of reconstruction levels are employed, uniform and trained. For
PNG
media_image2.png
18
97
media_image2.png
Greyscale
, uniform levels are used (i.e., the codeword is the center of the quantization cell). The remaining codewords are trained on the source data itself, except CW0 which is typically set to 0. The trained codeword
PNG
media_image3.png
18
118
media_image3.png
Greyscale
, is determined by taking the sample mean of all source symbols that map to
PNG
media_image4.png
17
45
media_image4.png
Greyscale
and the negative of all source symbols mapping to
PNG
media_image5.png
15
58
media_image5.png
Greyscale
”, Kasner’s UTCQ uses a uniform quantizer with stepsize Δ, so once the decoder determines which codeword index k in the uniform codebook applies (the “intermediate integer value”), the corresponding reconstruction value is k·Δ (or an empirically trained offset close to that). Under a broad, uniform-quantizer interpretation, this is equivalent to multiplying the integer index by the step size to obtain the reconstructed sample. Thus, the uniform codeword index k is being treated as the “intermediate integer value,” Δ as the “predetermined quantization step size,” and the multiplication k·Δ is what is being interpreted as “multiplying … the intermediate value … with the predetermined quantization step size.” )
Claim 8: Han, Kasner, and Sze teaches the limitations of claim 1, Kasner further teaches:
Apparatus of claim 1, wherein the apparatus is configured to select, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets by means of a state transition process by (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work… Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
. Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen [3].”, Kasner’s TCQ’s trellis is a state transition process that selects which codebook subset (superset) to use for each coefficient based on the current state. Kasner explicitly shows an eight-state trellis and explains that “at any given trellis state, the next codeword must come from one of two supersets,” and that given an initial state and a sequence of indices the decoder can reconstruct the sequence of chosen codewords. The movement from state to state along the trellis edges as each new symbol/index is processed is the claimed “state transition process.” At each step, the current state determines which superset (reconstruction-level set) is used, and the next state is determined by the current state and the chosen index, so the trellis operation as a whole is being interpreted as “selecting … the set of quantization levels … by means of a state transition process.” )
determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work… Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
.”, the trellis state corresponds to the state associated with the current neural network parameter when combined with Han.)
and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter. (Kasner, page 1678, col. 2, paragraph 1, “Equation (1) states that if we take the codeword associated with index i ∈ S₀, and that corresponding to –i ∈ S₁, the codewords will be the negative of one another. Equation (2) states that the probability of the codeword with index i ∈ S₀ equals the probability of the codeword with index –i ∈ S₁. These relationships allow the use of a single variable-rate code for both supersets [8]. The UTCQ quantizer returns the S₀ indices and the negative of the S₁ indices, allowing one probability model to be used for entropy coding. The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly.”, Kasner teaches that after decoding each index, the decoder “keeps track of the current state” by noting whether the index came from superset S₀ (positive codeword) or S₁ (negative codeword) and then applies that state when interpreting the next index. When combined with Han’s sequential decoding of neural-network weight indices, each decoded weight index becomes the “quantization index” that drives Kasner’s sign-based state update. Thus, for every neural-network parameter in Han’s stream, Kasner’s rule updates the internal trellis state based on the immediately preceding decoded index.)
Claim 11: Han, Kasner, and Sze teaches the limitations of claim 8, Kasner further teaches:
Apparatus of claim 8 wherein the state transition process is configured to transition between four or eight possible states. (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work.”, Kasner showcases an 8 trellis state model
Kasner, page 1678, col. 1, paragraph 1, “If such a codeword is required, the trellis must switch to state four by choosing a nonzero codeword from D2. If the following source samples require a string of zero reconstruction levels, the trellis must work its way back to state zero.”, Trellis Coded Quantization (TCQ) must switch (transition) between states as an inherent process.)
Claim 22: Kasner teaches:
Apparatus for encoding neural network parameters, which define a neural network, into a data stream, configured to sequentially encode the neural network parameters (Han, abstract, “Our method first prunes the network by learning only the important connections. Next, we quantize the weights to enforce weight sharing, finally, we apply Huffman coding.”, Han explicitly applies quantized weights to neural network parameters, thereby encoding them to a compressed stream.)
by selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets depending on quantization indices encoded into the data stream for previously encoded neural network parameters, (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work… Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
. Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen [3].”, Kasner’s TCQ trellis (Fig. 1) uses the previously encoded index to transition the trellis state, which in turn selects superset S₀ (even multiples) or S₁ (odd multiples) for the next parameter. Han’s neural‐network parameters take the place of Kasner’s “coefficients,” so this teaches using previous indices to choose the reconstruction‐level set. The superset of codewords at the current state is what is being interpreted as the “set of reconstruction levels”, the collection of supersets over the trellis is the “plurality of reconstruction level sets,” and the fact that the next state (and thus which superset applies to the current sample) depends on the previously chosen indices is what is being interpreted as “selecting … a set of reconstruction levels … depending on quantization indices decoded from the data stream for previous neural network parameters.” )
quantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels, (Kasner, page 1677, col. 2, paragraph 2, “During quantization, the Viterbi algorithm [6] is used to pick the sequence of codewords allowed by the trellis structure that minimizes the cumulative MSE between the input data and output reconstruction.”, Kasner’s encoder actually quantizes each parameter value by finding the closest codeword in the active superset via Viterbi. In combination with Han’s weight‐sharing (where each shared weight is treated as a codebook entry), this teaches quantizing a neural-network parameter onto its selected reconstruction level. The superset’s codewords are being interpreted as the “reconstruction levels,” and the index within that superset is the claimed quantization index that uniquely points to one codeword (one reconstruction level) out of the set associated with the current state.)
and encoding a quantization index for the current neural network parameter that indicates the one reconstruction level onto which the quantization index for the current neural network parameter is quantized into the data stream. (Kasner, page 1678, col. 2, paragraph 1, “The UTCQ quantizer returns the S0 indices and the negative of the S1 indices, allowing one probability model to be used for entropy coding. The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly”, Kasner describes how each quantization index—which inherently indicates a specific reconstruction level in its superset—is entropy-coded into the output stream (here, using sign-shifted indices). By combining with Han’s teaching that these indices represent neural network weights, one of ordinary skill would recognize that Kasner’s index bit-stream serves as the encoded quantization index for each neural network parameter.)
wherein the apparatus is configured to select, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets by means of a state transition process (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work… Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
. Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen [3].”, Kasner’s trellis is both a state machine and a selection mechanism: (i) at each symbol time, the current trellis state determines which of the two supersets (reconstruction-level sets) is eligible, and (ii) the transition to the next state is determined by the index chosen within that superset. That is, stateₙ → choice of superset for sample n, and (stateₙ, indexₙ) → stateₙ₊₁. This coupling of superset selection to the evolving trellis state is what is being interpreted as “selecting … the set of quantization levels … by means of a state transition process”—the selection and the transitions are two aspects of the same trellis operation. Thus, Kasner’s TCQ’s trellis is a state transition process that selects which codebook subset (superset) to use for each coefficient based on the current state.)
by determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
.”, The excerpt shows that which superset (the reconstruction level set comprising the set of quantization levels) is valid depends on the current state.)
and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter, (Kasner, page 1677, col. 2, paragraph 2, “Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen [3].”, the decoder updates its state after each decoded index so that the next state (and thus the next superset) is determined by the previous index.)
The motivation to combine Han with Kasner is substantially similar to that applied for claim 1 above.
Sze, in the same field of data encoding, teaches the following limitations which the above fails to teach:
and encode the quantization index for the current neural network parameter from the data stream using arithmetic coding using a probability model which depends on the state for the current neural network parameter and on the quantization index of previously decoded neural network parameters. (Sze, paragraph 6, “CABAC has multiple probability modes for different contexts. It first converts all non-binary symbols to binary symbols referred to as bins. Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate. Arithmetic coding is then applied to compress the data.”, CABAC’s arithmetic coder selects a probability model (context) based on the current state, and then arithmetically encodes each bin under that model, teaching state-dependent arithmetic encoding.)
The motivation to combine Han and Kasner with Sze is substantially similar to that applied for claim 1 above.
Claim 15: Han, Kasner, and Sze teaches the limitations of claim 14. Sze further teaches:
Apparatus of claim 14, configured to decode the quantization index for the current neural network parameter from the data stream using binary arithmetic coding by using the probability model which depends on the state for the current neural network parameter for at least one bin of a binarization of the quantization index. (Sze, paragraph 6, “In CABAC, bins can be either context coded or bypass coded. Bypass coded bins do not require context selection which allows these bins to be processed at a much high throughput than context coded bins.”, CABAC’s binary arithmetic coding of “bins” (the binarized bits of the index) using a state-dependent context model for at least one bin directly maps to the claims bin-wise arithmetic decoding.)
Claim 16: Han, Kasner, and Sze teaches the limitations of claim 14. Sze further teaches:
Apparatus of claim 14, wherein the at least one bin comprises a significance bin indicative of the quantization index of the current neural network parameter being equal to zero or not. (Sze, paragraph 6, “The theory and operation of CABAC coding for H.264/AVC is defined in the International Telecommunication Union, Telecommunication Standardization Sector (ITU-T) standard “Advanced video coding for generic audiovisual services” H.264, revision March 2005 or later, which is incorporated by reference herein. General principles are explained in “Context-Based Adaptive Binary Arithmetic Coding in the H.264/AVC Video Compression Standard,” Detlev Marpe, July 2003, which is incorporated by reference herein.”, Sze incorporates the H.264/AVC CABAC scheme by reference, where for transform coefficients a significant_flag (or significant_coeff_flag) is the first bin in the binarization signaling whether the coefficient is zero or non-zero. In that context, the bin representing significant_flag is the claimed “significance bin indicative of the quantization index … being equal to zero or not”—a bin value of 0 indicates a zero coefficient, and a bin value of 1 indicates a non-zero coefficient. The text in Sze pointing to CABAC’s standard operation is thus being interpreted as importing this significance-bin behavior.)
Claim 18: Han, Kasner, and Sze teaches the limitations of claim 17. Sze further teaches:
Apparatus of claim 17, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, a subset of probability models out of a plurality of probability models (Sze, paragraph 36, “Referring now to the CABAC encoder of FIG. 2A, the binarizer 200 converts syntax elements into strings of one or more binary symbols. The binarizer 200 directs each bin to either the context coding 206 or the bypass coding 208 of the bin encoder 204 based on a bin type determined by the context modeler 202.”,
Sze, paragraph 6, “Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate.”, In Sze’s CABAC encoder (Fig. 2A; paragraph, 36), the binarizer 200 “directs each bin to either the context coding 206 or the bypass coding 208 of the bin encoder 204 based on a bin type determined by the context modeler 202.” Sze also explains that, “for each bin, the coder selects which probability model to use” (Paragraph, 6). Under a broadest reasonable interpretation, the “plurality of probability models” are the various context probability models used on the context-coded branch together with the (effectively fixed) model used on the bypass branch. When the binarizer/context modeler decides whether a bin is sent to the context coding 206 path or the bypass 208 path, that decision preselects which subset of the available probability models will be used for that bin (the context-coded subset vs. the bypass subset), corresponding to “preselect[ing] … a subset of probability models out of a plurality of probability models.”)
and select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters. (Sze, paragraph 6, “Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate”, the selection of probability model (the context model as describe further in the patent) is based on the values of nearby elements. By combining with Han’s neural network quantization a POSITA would have applied the nearby elements paradigm to its neural network parameters.)
Claim 19: Han, Kasner, and Sze teaches the limitations of claim 18. Kasner further teaches:
Apparatus of claim 18, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, the subset of probability models out of the plurality of probability models in a manner so that a subset preselected for a first state or reconstruction levels set is disjoint to a subset preselected for any other state or reconstruction levels set. (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work… Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets
PNG
media_image1.png
18
185
media_image1.png
Greyscale
. Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen [3].”, As discussed for claim 18 above, Sze teaches having a plurality of probability models (context models and bypass mode) and preselecting a subset of those probability models for a given bin (e.g., by directing the bin to the context-coded branch or to the bypass branch of the CABAC engine). Kasner, in turn, teaches that at any given trellis state “the next codeword must come from one of two supersets” S₀ or S₁ (the reconstruction level sets from which to select), and that a sequence of indices identifying which codeword from the appropriate superset was chosen allows the decoder to reproduce the sequence of codewords (Kasner, p.1677, col.2). Thus, Kasner’s trellis state determines which of two disjoint reconstruction-level supersets (S₀ vs. S₁) is active at each step. It is interpreted by the examiner that the choice of S₀ vs. S₁ serves as an additional input to Sze’s context modeler so that, when S₀ (a first reconstruction level set selected) is active, one subset of Sze’s probability models is preselected, and when S₁ is active (a second reconstruction level set selected), a different, non-overlapping subset is preselected. In this combined system, the subset of probability models preselected for a first state / reconstruction-level set (S₀) is disjoint from the subset preselected for another state / reconstruction-level set (S₁), as required by the claim.)
Claim 20: Han, Kasner, and Sze teaches the limitations of claim 18. Sze further teaches:
Apparatus of claim 18, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to. (Sze, paragraph 6, “Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate”, Sze (via CABAC) states that, for each bin, the probability model (context) is selected using information from nearby elements—for transform coefficients, this includes previously decoded significant_flags, levels, and positions in the same block. When this is applied to neural-network parameters (via Han/Wang), the “nearby elements” correspond to previously decoded quantized parameters in a neighboring portion of the network (e.g., adjacent weights or nodes). Thus, CABAC’s rule of choosing a context based on the already decoded neighboring coefficients is what is being interpreted as “selecting the probability model … depending on the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring the portion of the current parameter.” )
Claims 5, 9, 10, 12, and 13are rejected under 35 U.S.C. 103 as being unpatentable over Han in view of Kasner and in further view of Sze and Coban et al., (US 11,451,840 B2), hereafter referred to as Coban.
Claim 5: Han, Kasner, and Sze teaches the limitations of claim 1. Coban, in the same field of trellis coded quantization, teaches the following limitations which the above fails to teach:
Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two (Coban, col. 12, line 54, “FIG . 3 is a conceptual diagram illustrating an example of how two scalar quantizers can be used to perform quantization… a first quantizer ( e.g. , Q1 ) may be configured with a first set of quantization parameters and a second quantizer ( e.g. , Q2 ) may be configured with a second set of quantization parameters that are different in value from the first set”, two reconstruction level sets Q1 and Q2 are present in Coban.)
and the plurality of reconstruction level sets comprises a first reconstruction level set that comprises zero and even multiples of a predetermined quantization step size, (Coban, col. 13, line 1, “In particular , when using two scalar quantizers , a first quantizer Q0 may map transform coefficient levels ( numbers below the points , e.g. , absolute values ) to even integer multiples of quantization step size Δ.”)
and a second reconstruction level set that comprises zero and odd multiples of the predetermined quantization step size. (Coban, col. 13, line 4, “The second quantizer Q1 may map the transform coefficient levels to odd integer multiples of the quantization step size Δ or to zero .”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Han (i.e. deep neural network quantization methos) and Kasner by incorporating the teachings of Coban (i.e. Trellis-Coded Quantization using parity based indexing). A motivation of which is to provide a codebook quantization technique using Trellis Coded Quantization as to increase computation efficiency. (Coban, col. 33, line 26, “In this way , video encoder 200 and / or video decoder 300 may quantize or inverse quantize a set of syntax elements representing the remaining levels ( e.g. , gt3 418A - 418N ) of the plurality of coefficient levels for transform coefficients of residual data for the block of video data without grouping all bypass coded bins for simpler parsing , thereby improving a computation efficiency of video encoder 200 and / or video decoder 300 .”)
Claim 9: Han, Kasner, and Sze teaches the limitations of claim 8 Coban, in the same field of trellis coded quantization, teaches the following limitations which the above fails to teach:
Apparatus of claim 8 configured to update the state for the subsequent neural network parameter using a binary function of the quantization index decoded from the data stream for the immediately preceding neural network parameter. (Coban, col. 13, line 1, “In particular , when using two scalar quantizers , a first quantizer Q0 may map transform coefficient levels ( numbers below the points , e.g. , absolute values ) to even integer multiples of quantization step size Δ. The second quantizer Q1 may map the transform coefficient levels to odd integer multiples of the quantization step size Δ or to zero .”, Coban drives state transitions by the parity (even/odd) of the previous decoded coefficient level. It is interpreted that this form of Trellis-Coded Quantization would then be applied to the neural network parameters of Han. It is interpreted by the examiner that the parity (even/odd) function used by Coban is analogous to that of a binary function. Coban distinguishes Q0 (mapping to even integer multiples of Δ) from Q1 (mapping to odd integer multiples or zero). The decision of which quantizer (and hence which reconstruction-level set) to use can thus be implemented as a function that only examines whether the previous level/ index is even or odd. In other words, one can define a binary function f(index) ∈ {0,1} that outputs one value for even indices and another for odd indices; that binary parity result is then used to drive the state machine’s next state or set selection. In this sense, Coban’s reliance on even vs. odd integer multiples is what is being interpreted as “updating the state … using a binary function of the quantization index” (where the binary function is the parity test).).
Claim 10: Han, Kasner, and Sze teaches the limitations of claim 8. Coban, in the same field of trellis coded quantization, teaches the following limitations which the above fails to teach:
Apparatus of claim 8, configured to update the state for the subsequent neural network parameter using a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter. (Coban, col. 13, line 1, “In particular , when using two scalar quantizers , a first quantizer Q0 may map transform coefficient levels ( numbers below the points , e.g. , absolute values ) to even integer multiples of quantization step size Δ. The second quantizer Q1 may map the transform coefficient levels to odd integer multiples of the quantization step size Δ or to zero .”, Coban drives state transitions by the parity (even/odd) of the previous decoded coefficient level. Because Coban explicitly organizes reconstruction levels into even multiples (Q0) and odd multiples (Q1) of Δ, the quantizer’s behavior is governed by the parity (even/odd) of the relevant integer index or level. Whenever the scheme chooses between Q0 and Q1 based on whether a level (or its integer multiple index) is even or odd, it is effectively using the parity of the quantization index as the update variable for the state machine. That even-versus-odd test is what is being interpreted as “using a parity of the quantization index … to update the state for the subsequent neural network parameter.” It is interpreted that this form of Trellis-Coded Quantization would then be applied to the neural network parameters of Han.)
Claim 12: Han, Kasner, and Sze teaches the limitations of claim 8. Coban, in the same field of trellis coded quantization, teaches the following limitations which the above fails to teach:
Apparatus of claim 8, configured to transition, in the state transition process, between an even number of possible states and the number of reconstruction level sets of the plurality of reconstruction level sets is two, wherein the determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets depending on the state associated with the current neural network parameter determines a first reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a first half of the even number of possible states, and a second reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a second half of the even number of possible states. (Coban, col. 13, line 17, “In this example , coefficients in state 0 and state 1 use the Q0 ( even integer multiples of stepsize ) quantizer , and coefficients in states 2 and 3 use the Q1 ( odd quantizer , and coefficients in states 2 and 3 use the Q1 ( odd sets can be achieved by changing the parity of the level of integer multiples of stepsize ) quantizer.”, Coban’s 4-state example (states 0–3) shows that states 0 and 1 use Q0 (even-multiple quantizer) while states 2 and 3 use Q1 (odd-multiple quantizer). Here, the total number of states (4) is even; the first two states (0,1) constitute the “first half” of the states and are associated with the first reconstruction-level set (Q0), while the remaining two states (2,3) constitute the “second half” and are associated with the second reconstruction-level set (Q1). Thus, Coban’s mapping of Q0 to states in the first half and Q1 to states in the second half of an even-sized state set is what is being interpreted as the claimed “if the state belongs to a first half … use the first reconstruction level set … if the state belongs to a second half … use the second reconstruction level set.”)
Claim 13: Han, Kasner, and Sze teaches the limitations of claim 8. Coban, in the same field of trellis coded quantization, teaches the following limitations which the above fails to teach:
Apparatus of claim 8, configured to perform the update of the state by means of a transition table which maps a combination of the state and a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter onto a further state associated with the subsequent neural network parameter. (Coban, col 13, line 14, “FIG . 4 is a conceptual diagram illustrating an example state transition scheme for two scalar quantizers used to perform quantization . In this example , coefficients in state 0 and state 1 use the Q0 ( even integer multiples of stepsize quantizer , and coefficients in states 2 and 3 use the Q1 ( odd integer multiples of stepsize ) quantizer .”, figure 4 of Coban clearly discloses a state transition scheme which maps the combination of the state and parity of the quantization index and shows transitions to further states. It is interpreted by the examiner that the neural network quantization of Han and its data stream would similarly apply here and provide a data stream of neural network related parameters to perform the trellis-coded quantization of Coban.)
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Han in view of Kasner and Sze in further view of Coban and Schwarz et al., (Schwarz, H., Nguyen, T., Marpe, D., & Wiegand, T. (2019, March). Hybrid video coding with trellis-coded quantization. In 2019 Data Compression Conference (DCC) (pp. 182-191). IEEE.), hereafter referred to as Schwarz.
Claim 7: Han, Kasner, and Sze teaches the limitations of claim 1. Coban, in the same field of trellis coded quantization, teaches the following limitations which the above fails to teach:
Apparatus of claim 7, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two and the apparatus is configured to derive the intermediate value for each neural network parameter by, if the selected reconstruction level set for the respective neural network parameter is a first set, multiply the quantization index for the respective neural network parameter by two to acquire the intermediate value for the respective neural network parameter; (Coban, col. 13, line 1, “In particular , when using two scalar quantizers , a first quantizer Q0 may map transform coefficient levels ( numbers below the points , e.g. , absolute values ) to even integer multiples of quantization step size Δ.”, mapping each index to an even integer multiple of Δ is exactly the same as computing an intermediate value of the index multiplied by 2.)
and if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is equal to zero, set the intermediate value for the respective neural network parameter equal to zero; (Coban, col. 13, line 4, “The second quantizer Q1 may map the transform coefficient levels to odd integer multiples of the quantization step size Δ or to zero .”, the case when the decoded index is 0 is covered by Q1 outputting 0 as the intermediate value.)
It would have been obvious to one of ordinary skill in the art before the effective filing date
of the claimed invention to have modified the teachings of Han, Kasner, and Sze by incorporating the teachings of Coban (i.e. Trellis-Coded Quantization using parity based indexing). A motivation of which is to provide a codebook quantization technique using Trellis Coded Quantization as to increase computation efficiency.
Schwarz, in the same field of trellis coded quantization, teaches the following limitations which Han, Kasner, Sze and Coban fail to teach:
and if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is greater than zero, multiply the quantization index for the respective neural network parameter by two and subtract one from the result of the multiplication to acquire the intermediate value for the respective neural network parameter; (Schwarz, page 184, paragraph 3, “The structure of the two scalar quantizers Q0 and Q1 used in our approach is shown in Fig. 1. The reconstruction values for Q0 are given by the even multiples of the quantization step size Δ; the reconstruction values for Q1 are given by the odd multiples of Δ and, in addition, the value of zero… For both quantizers, the reconstruction values are indicated by integer quantization indexes q, where the quantization index equal to zero corresponds to the reconstruction value equal to zero.”,
Page 185, paragraph 1, “Given the N quantization indexes qk of a block, with k indicating the coding order, the associated reconstructed transform coefficients t k can be obtained by the following simple algorithm:
PNG
media_image7.png
90
254
media_image7.png
Greyscale
”, In Schwarz, the TCQ design defines two scalar quantizers Q0 and Q1, where Q0 uses even multiples of Δ and Q1 uses odd multiples of Δ (plus zero), with each reconstruction value associated with an integer quantization index q. For each coefficient, the decoder computes the reconstructed value as t_k = (2·q_k − (s_k >> 1)·sgn(q_k))·Δ. For states corresponding to the second quantizer (Q1), (s_k >> 1) = 1, so the formula specializes to t_k = (2·q_k − sgn(q_k))·Δ. For a positive quantization index q_k > 0, sgn(q_k) = 1, and thus t_k = (2·q_k − 1)·Δ. Under the claim’s terminology, the quantization index is q_k, the “intermediate value” is the integer (2·q_k − 1), and multiplying that intermediate value by the predetermined quantization step size Δ yields the odd reconstruction level. This directly corresponds to “multiply the quantization index … by two and subtract one … to acquire the intermediate value” when the selected reconstruction-level set is the second set and the quantization index is greater than zero.)
and if the selected reconstruction level set for a current neural network parameter is a second set and the quantization index for the respective neural network parameter is less than zero, multiply the quantization index for the respective neural network parameter by two and add one to the result of the multiplication to acquire the intermediate value for the respective neural network parameter. (Schwarz, page 184, paragraph 3, “The structure of the two scalar quantizers Q0 and Q1 used in our approach is shown in Fig. 1. The reconstruction values for Q0 are given by the even multiples of the quantization step size Δ; the reconstruction values for Q1 are given by the odd multiples of Δ and, in addition, the value of zero… For both quantizers, the reconstruction values are indicated by integer quantization indexes q, where the quantization index equal to zero corresponds to the reconstruction value equal to zero.”,
Page 185, paragraph 1, “Given the N quantization indexes qk of a block, with k indicating the coding order, the associated reconstructed transform coefficients t k can be obtained by the following simple algorithm:
PNG
media_image7.png
90
254
media_image7.png
Greyscale
”, As above, Schwarzs’ TCQ scheme uses two quantizers Q0 and Q1; Q1’s reconstruction values are odd integer multiples of Δ plus zero, with each value addressed by an integer quantization index q. The decoder reconstructs each coefficient via t_k = (2·q_k − (s_k >> 1)·sgn(q_k))·Δ. For the second quantizer Q1, (s_k >> 1) = 1, giving t_k = (2·q_k − sgn(q_k))·Δ. When the quantization index is negative (q_k < 0), sgn(q_k) = −1, so the expression becomes t_k = (2·q_k + 1)·Δ. The factor (2·q_k + 1) is a negative odd integer (… −5, −3, −1, …), which, when multiplied by Δ, yields the negative odd-multiple reconstruction levels required for Q1. Under the claim’s terminology, the quantization index is q_k, the “intermediate value” is the integer (2·q_k + 1), and multiplying that intermediate value by the predetermined quantization step size Δ produces the final reconstruction level. This matches “multiply the quantization index … by two and add one … to acquire the intermediate value” for the second reconstruction-level set when the quantization index is less than zero.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to incorporate the trellis-coded quantization scheme of Schwarz into the existing Han/Kasner/Coban combination in order to improve coding efficiency and rate–distortion performance when quantizing and reconstructing coefficients. Schwarz describes replacing conventional scalar quantization by trellis-coded quantization and reports that “the coding efficiency of transform coding can be improved by replacing scalar quantization with trellis-coded quantization (TCQ) and using advanced entropy coding techniques for coding the quantization indexes” (Abstract) and that their implementation in the VVC test model “yielded average bit-rate savings of 4.9% for intra-only coding and 3.3% for typical random access configurations” (Abstract) relative to a scalar-quantization baseline. A person of ordinary skill in the art implementing Han’s neural-network quantization with Kasner/Coban-style trellis coding would therefore have been motivated to adopt Schwarz explicit TCQ reconstruction rule—using intermediate values of the form 2·q_k − sgn(q_k) (i.e., 2i−1 and 2i+1 for positive and negative indices)—to obtain the same kind of bit-rate savings and improved coding efficiency for neural-network parameters.
Claims 21 are rejected under 35 U.S.C. 103 as being unpatentable over Han in view of Kasner and in further view of Sze and Marpe et al., (Marpe, D., Schwarz, H., & Wiegand, T. (2003). Context-based adaptive binary arithmetic coding in the H. 264/AVC video compression standard. IEEE Transactions on circuits and systems for video technology, 13(7), 620-636.), hereafter referred to as Marpe.
Claim 21: Kasner, Han, and Sze teaches the limitations of claim 18. Marpe, in the same field of Context-Aware Binary Arithmetic Coding further teaches:
Apparatus of claim 18, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on a characteristic of the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, (Marpe, page 629, col. 1, paragraph 2, “Characteristic Features: For the coding of residual data within the H.264/AVC standard specifically designed syntax elements are used in CABAC entropy coding mode. These elements and their related coding scheme are characterized by the following distinct features… Context models for coding of nonzero transform coefficients are chosen based on the number of previously transmitted nonzero levels within the reverse scanning path.”, H.264’s CABAC selects a context model for each bin by examining previously coded motion-vector differences (MVD’s), which is interpreted as analogous to choosing a probability model based on previously decoded weight indices from neighboring parameters when combined with Han.)
the characteristic comprising on or more of the signs of non-zero quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, (Marpe, page 629, col. 2, paragraph 2, “Significance Map: If the coded_block_flag indicates that a block has significant coefficients, a binary-valued significance map is encoded. For each coefficient in scanning order, a one-bit symbol significant_coeff_flag is transmitted.”, Marpe’s CABAC description encodes, for each coefficient, both (i) a significant_coeff_flag indicating zero vs. non-zero, and (ii) a coeff_sign_flag representing the sign of each non-zero coefficient. Contexts for these flags are conditioned on neighboring coefficients in the same block. When mapped to neural networks, each transform coefficient corresponds to a parameter; the neighboring coefficients in the block correspond to a neighboring portion of the neural network, and the coeff_sign_flag bits for those neighbors are the “signs of non-zero quantization indices of previously decoded neural network parameters.” Thus, using those sign flags in context templates is what is being interpreted as using the sign characteristics of neighboring non-zero indices to drive probability-model selection.)
the number of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero (Marpe, page 625, col. 2, paragraph 1, “In contrast to all other types of context models, both types depend on the context categories of different block types, as specified below… modeling functions are specified that involve the evaluation of the accumulated number of encoded (decoded) levels with a specific value prior to the current level bin to encode (decode).”, Marpe, page 630, col. 2, last paragraph, “When a level with an absolute value greater than 1 has been encoded, i.e., when NumLgt1 is greater than 0, a context index increment of 4 is used for all remaining levels of the regarded block.”, Marpe specifies context models that depend on counters like NumLgt1(i) and NumT1(i), which are accumulated numbers of previously encoded non-zero coefficients with certain absolute values before the current bin. These counters effectively count how many prior coefficients in the same block (neighboring region) are non-zero. When mapped onto neural-network parameters, that block of coefficients is the neighboring portion of the network, and the count NumLgt1/NumT1 is what is being interpreted as the “number of quantization indices of previously decoded neural network parameters … in the neighboring portion … which are non-zero.”)
a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to a difference between (Marpe, page 630, col. 2, last paragraph, “When a level with an absolute value greater than 1 has been encoded, i.e., when NumLgt1 is greater than 0, a context index increment of 4 is used for all remaining levels of the regarded block.”, Marpe defines counters such as NumLgt1 that accumulate how many previously encoded coefficients in the block have an absolute level greater than 1, and uses conditions like “when NumLgt1 is greater than 0, a context index increment of 4 is used for all remaining levels of the regarded block” (Marpe, p.630, col.2). Each coefficient with |level|>1 contributes at least 2 to the total sum of absolute values across the block, so the condition NumLgt1>0 implies that the sum of absolute values of previously encoded coefficients in that block exceeds a positive minimum. Using NumLgt1 as a context-selection variable is therefore being interpreted as using a characteristic derived from the sum of absolute values of previously decoded neighboring quantization indices when selecting a probability model, corresponding to the claimed “sum of the absolute values … of previously decoded … neighboring parameters.” Using NumLgt1(i) as a context-selection condition is therefore being interpreted, under a broad reading, as using a function of the sum of absolute values of previously decoded neighboring quantization indices when selecting a probability model, corresponding to the claimed “sum of the absolute values … of previously decoded … neighboring parameters.”)
a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and the number of quantization indices of the previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero. (Marpe, page 630, col. 2, last paragraph, “Then, the context for the first bin of coeff_abs_level_minus1 is determined by the current value NumT1, where the following additional rules apply. If more than three past coded coefficients have an absolute value of 1, the context index increment of three is always chosen. When a level with an absolute value greater than 1 has been encoded, i.e., when NumLgt1 is greater than 0, a context index increment of 4 is used for all remaining levels of the regarded block.”, Marpe’s decision logic jointly examines NumT1 (the count of prior levels equal to 1) and NumLgt1 (the count/sum of prior levels >1) to choose between two context-index increments (3 vs. 4). In effect, it is checking whether the difference between the total magnitude (sum) and the simple count exceeds a threshold, mapping directly to use a difference between “sum of absolute values of quantization indices” and “number of quantization indices” for model selection.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Han (i.e. deep neural network quantization methos), Kasner, and Sze by incorporating the teachings of Marpe (i.e. Context-Aware Binary Arithmetic Coding in H.264/AVC). A motivation of which is to provide a quantization technique that would further reduce the overall bitstream size. (Marpe, page 632, col. 2, paragraph 3, “The CABAC entropy coding method is part of the Main profile of H.264/AVC [1] and may find its way into video streaming, broadcast, or storage applications within this profile. Experimental results have shown the superior performance of CABAC in comparison to the baseline entropy coding method of VLC/CAVLC. For typical test sequences in broadcast applications, averaged bit-rate savings of 9% to 14% corresponding to a range of acceptable video quality of about 30–38 dB were obtained.”, Marpe’s CABAC method provides explicit superior performance in terms of amount of compression yielding acceptable quality loss.)
Claims 23, and 25 are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Gao et al.,(Gao, M., Zhao, D., Wang, Q., & Gao, W. (2014, October). Hierarchical dependency context model based arithmetic coding for DCT video compression. In 2014 IEEE International Conference on Image Processing (ICIP) (pp. 3136-3140). IEEE.), hereafter referred to as Gao, in view of Nguyen et al., (US 9036710 B2), hereafter referred to as Nguyen.
Claim 23: Wang teaches the following:
Apparatus for reconstructing neural network parameters, which define a neural network, comprising a processor configured to derive first neural network parameters for a first reconstruction layer to yield, per neural network parameter, a first- reconstruction-layer neural network parameter value, (Wang, page 512, section 2, paragraph 3, “To address this problem, we adopt the scalable coding concept in image/video coding… and represent the weights hierarchically, i.e., each weight is represented by a base-layer component and several enhancement-layer components;”, In Wang, each weight (i.e., neural network parameter) is represented by a “base-layer component” and enhancement components. The “base-layer” component corresponds to the claimed “first reconstruction layer”, and because “each weight is represented” in this way, there is one base-layer value per parameter. Thus Wang teaches deriving first neural network parameters for a first reconstruction layer that yield, per neural network parameter, a first-reconstruction-layer value.)
decode second neural network parameters for a second reconstruction layer from a data stream to yield, per neural network parameter, a second-reconstruction-layer neural network parameter value, (Wang, page 512, section 2, paragraph 4, “This yields the 1-bit base-layer approximations of all weights. Next, we perform another K-means clustering with K = 2 on all quantization errors, and record the corresponding cluster indices, centroids, and quantization errors. The gives us the 1-bit first-enhancement-layer representations of all weights.”, In Wang, the first-enhancement-layer representations (cluster indices and centroids obtained from K-means on the quantization errors) are the claimed “second neural network parameters for a second reconstruction layer.” Wang, hierarchical quantization produces “1-bit first-enhancement-layer representations of all weights,” and the weight is then expressed as w ≈ b₁ + e₁ + … + eₙ₋₁, where b₁ and eᵢ are the centroids of the base layer and i-th enhancement layer (Wang, p.512, eq. (1)). Given that the codebook “stores the centroid and cluster index information of all quantization layers” (Wang, p.512, col. 2, paragraph 2), the centroids and cluster indices corresponding to the first enhancement layer, when read from that stored codebook for a given weight and used as e₁ in equation 1, are being interpreted as the “second neural network parameters for a second reconstruction layer” and as yielding, per neural network parameter, a second-reconstruction-layer neural network parameter value.
That process of reading the enhancement-layer indices/centroids from the codebook stream and using them to obtain eᵢ for each weight is what is being interpreted as “decoding the second neural network parameters … from a data stream … to yield, per neural network parameter, a second-reconstruction-layer … value.” They are stored in a codebook/bitstream that is transmitted and decoded, so decoding those enhancement-layer parameters from the stream yields, for each weight, a second-layer value, as recited in this claim’s limitation.)
and reconstruct the neural network parameters by, for each neural network parameter, combining the first-reconstruction-layer neural network parameter value and the second-reconstruction-layer neural network parameter value. (Wang, page 512, col. 1, last paragraph, “By repeating this procedure, we can obtain a n-layer hierarchical representation of a weight, i.e.,
PNG
media_image8.png
23
237
media_image8.png
Greyscale
”, Wang reconstructs each weight using a hierarchical representation in which
w ≈ b₁ + e₁ + … + eₙ₋₁, where b₁ and eᵢ are centroids of the base layer and i-th enhancement layer respectively. For a two-layer case, each neural network parameter is reconstructed by combining b₁ (first-layer value) and e₁ (second-layer value). This directly matches “for each neural network parameter, combining the first-reconstruction-layer… value and the second-reconstruction-layer… value.”).
Gao, in the same field of context model entropy coding, teaches the following limitation which Wang fails to teach:
decode the second-reconstruction-layer neural network parameter value from the data stream by context-adaptive entropy decoding, (Gao, page 3138, col. 1, paragraph 2, “To code these coding elements, binary arithmetic coding based on HDCM (HDCMBAC) is designed”,
Abstract, “Experimental results demonstrate that HDCMBAC can achieve the similar coding performance as CABAC at low and high QPs. Meanwhile the context modeling and arithmetic decoding in HDCMBAC can be carried out in parallel, since the context dependency only exists among different parts of basic coding elements in HDCM.”, Gao’s HDCMBAC uses binary arithmetic coding with a hierarchical dependency context model (HDCM) for transform-coefficient elements (significant_flag, bin0, bin13). The encoder codes these elements using context-dependent probability models; the decoder performs the inverse operation, arithmetic decoding using the same context modeling based on previously decoded elements and statistics. This is exactly context-adaptive entropy decoding: the probability of each decoded bin is conditioned on context variables (position, number of non-zero coefficients, etc.), and the binary arithmetic decoder adapts the probabilities accordingly. When mapped to the claim, the “second-reconstruction-layer neural network parameter value” is treated like a coefficient level coded via HDCMBAC, and recovering it from the bitstream using HDCMBAC is what is being interpreted as “decode … by context-adaptive entropy decoding.”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have combined Wang with Gao’s HDCMBAC scheme (“Hierarchical dependency context model based arithmetic coding for DCT video compression”), a motivation of which would have been to reuse a proven hierarchical dependency context model and binary arithmetic coder to exploit statistical dependencies in coefficient-like parameters, thereby improving entropy-coding efficiency for Wang’s quantized neural network parameters. Gao, abstract, states: “This paper presents an arithmetic coding scheme for DCT coefficients in video compression, in which the number of non-zero coefficients, significant map and level information for a DCT block are used as coding elements. To exploit the statistical correlations, an hierarchical dependency context model (HDCM) is proposed, where the number of non-zero coefficients and scanned position are used to capture the magnitude varying tendency of DCT coefficients. Then a new binary arithmetic coding using HDCM (HDCMBAC) is proposed to code the coding elements.”, Given that Wang already quantizes NN parameters into discrete levels and sends them through an entropy coder, a POSITA would naturally apply Gao’s hierarchical dependency context modeling to those parameters (treated analogously to transform coefficients) to better capture correlations and thus implement the claimed context-adaptive entropy decoding for the second-layer parameters.
Nguyen, in the same field of context encoding/decoding, teaches the following limitations which Wang and Gao fail to teach:
selecting a probability context set out of a collection of probability context sets depending on the first-reconstruction-layer neural network parameter value, (Nguyen, col. 12, lines 21-30, “There may further be multiple context sets that are selected based on neighboring or nearby significant-coefficient flags M, N, and P… context_set = min(2, M+N+P)”, Nguyen explicitly describes multiple context sets and selecting a particular context set (context_set) based on a function of previously decoded significant-coefficient flags M, N, P. Those flags encode already-reconstructed transform-coefficient information, i.e., reconstruction-layer values. Thus Nguyen teaches selecting a probability context set (the context_set index) out of a collection of sets, depending on values derived from previously reconstructed parameters.)
selecting a probability context to be used out of the selected probability context set depending on the first-reconstruction-layer neural network parameter value. (Nguyen, col. 12, line 17-18, “There may be contexts… selected based on the sum of the neighboring or nearby greater-than-one flags… context = context_set*4 + min(3, a+b+c+d)”, For level coding, Nguyen first chooses a context set (context_set), then selects a specific context within that set using the sum of neighboring flags a, b, c, d. Those flags encode information about already-reconstructed coefficient levels (magnitude/significance), i.e., first-layer reconstruction values. This directly corresponds to “selecting a probability context … out of the selected probability context set depending on” values derived from earlier reconstruction-layer parameters.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have combined Wang and Gao with Nguyen, a motivation of which would have been to refine Gao’s generic context modeling by using adaptive thresholds and level information from previously reconstructed groups when selecting how many level flags or contexts to use, thereby further improving entropy-coding efficiency of Wang’s layered neural network parameters. Nguyen’s abstract explains: “Methods and devices for reconstructing coefficient levels from a bitstream of encoded video data for a coefficient group in a transform unit, using adaptive-threshold-based level coding. Threshold is set based upon level information from one or more previously-reconstructed coefficient groups in the transform unit. Threshold may be maximum number of level flags to decode for the coefficient group. Level information may include number of level flags decoded in previous coefficient groups.”, Given Gao’s HDCM already uses neighboring/statistical information for context modeling, a POSITA would see Nguyen’s adaptive-threshold mechanism as a way to condition context/flag coding not just on position but also on previously decoded level information. In Wang’s neural-network setting, the “coefficient groups” map naturally to groups of weights/parameters, so using Nguyen’s adaptive thresholds and context choice to drive the second-layer entropy decoding is a reasonably expected optimization.
Claim 25: Wang, Gao, and Nguyen teaches the limitations of claim 23, Wang further teaches:
Apparatus of claim 23, configured to reconstruct the neural network parameters by a parameter wise sum or parameter wise product of, per neural network parameter, the first-reconstruction-layer neural network parameter value and the second-reconstruction-layer neural network parameter value. (Wang, page 512, col. 1, last paragraph, “By repeating this procedure, we can obtain a n-layer hierarchical representation of a weight, i.e.,
PNG
media_image8.png
23
237
media_image8.png
Greyscale
where w is a uncompressed weight, b1 and ei are the centroid of the base layer and the i-th enhancement layer respectively.”, Wang’s formula w ≈ b₁ + e₁ + … shows that each weight is reconstructed by adding the base-layer centroid b₁ and enhancement-layer centroid(s) eᵢ. With b₁ as the first-reconstruction-layer value and e₁ as the second-reconstruction-layer value, this is a parameter-wise sum of first and second layer values for each neural network parameter.)
Claims 24, and 26 are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Gao and Nguyen and in further view of Lee et al.,(WO2006112642A1), hereafter referred to as Lee.
Claim 24: Wang, Gao, and Nguyen teaches the limitations of claim 23. Wang further teaches:
Apparatus of claim 23, configured to decode the first neural network parameters for the first reconstruction layer from the data stream or from a separate data stream, (Wang, page 512, col. 2, paragraph 2, “After the hierarchical quantization, we can build a codebook that stores the centroid and cluster index information of all quantization layers.”, Wang explains that, after hierarchical quantization, “we can build a codebook that stores the centroid and cluster index information of all quantization layers” and that a weight w is represented hierarchically as w ≈ b₁ + e₁ + … + eₙ₋₁, where b₁ is the centroid of the base layer (Wang, p.512). Under a broadest reasonable interpretation, the centroid and cluster index information in the codebook for the base layer are the “first neural network parameters for the first reconstruction layer,” and when those stored base-layer centroids and indices are read and used in the hierarchical representation to obtain b₁ for each weight, that corresponds to “decoding the first neural network parameters for the first reconstruction layer from the data stream (or a separate data stream).” Those base-layer centroids/indices are the claimed “first neural network parameters for the first reconstruction layer,” and the act of reading them from the codebook bitstream and reconstructing b₁ per weight is what is being interpreted as “decoding the first neural network parameters … from the data stream (or a separate data stream).” These values are transmitted and decoded when reconstructing the network. The base-layer centroids/indices are the claimed “first neural network parameters for the first reconstruction layer”; decoding them from the codebook stream.)
Lee, in the same field of context model entropy coding, teaches the following limitation which Wang fails to teach:
and decode the second neural network parameters for the second reconstruction layer from the data stream by context-adaptive entropy decoding using separate probability contexts for the first and second reconstruction layers. (Lee, paragraph 22, “According to still yet another aspect of the present invention, there is provided a decoder for decoding a residual prediction flag indicating whether residual data for an enhancement layer block of a multi-layered video signal is predicted from residual data for a lower layer block corresponding to the residual data for the enhancement layer block,”, Lee’s multi-layer decoder describes coding a residual prediction flag for an enhancement-layer block using CABAC (Context-Adaptive Binary Arithmetic Coding, a form of entropy encoding), where the coding method (i.e., which probability model / context is used) depends on the energy of the corresponding lower-layer (base-layer) residual. In CABAC, different coding methods correspond to distinct context sets. Wang supplies first and second reconstruction layers for NN parameters; Lee supplies the idea that enhancement-layer symbols (second layer) are coded with CABAC using context models chosen based on lower-layer information, and these models are distinct from those used for lower-layer symbols. This is what is being interpreted as “context-adaptive entropy decoding using separate probability contexts for the first and second reconstruction layers.”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have combined Wang, Gao, and Nguyen with Lee’s multi-layer video coding, a motivation of which would have been to exploit redundancy between a base representation and an enhancement representation by using information from the base layer to guide entropy coding of the enhancement layer, thereby improving overall rate–distortion efficiency in a layered reconstruction scheme. Lee explicitly frames a base/enhancement split and inter-layer prediction “to increase coding efficiency” (Lee, paragraph 95, “Because a context model is selected for each slice according to the above-described CABAC, probability values of probability models are initialized to a table of constant values for the slice. CABAC provides better coding efficiency than conventional VLC when a predetermined amount of information is accumulated because a context model has to be continuously updated based on statistics of recently-coded data symbols.”)
Claim 26: Wang, Gao, and Nguyen teaches the limitations of claim 23.
Lee, in the same field of context model entropy coding, teaches the following limitation which Wang, Gao, and Nguyen fails to teach:
Apparatus of claim 23, configured to decode the second-reconstruction-layer neural network parameter value from the data stream by context-adaptive entropy decoding using a probability model which depends on the first-reconstruction-layer neural network parameter value. (Lee, paragraph 22, “According to still yet another aspect of the present invention, there is provided a decoder for decoding a residual prediction flag indicating whether residual data for an enhancement layer block of a multi-layered video signal is predicted from residual data for a lower layer block corresponding to the residual data for the enhancement layer block,”, Lee explains that, for an enhancement-layer block, the decoder (i) calculates the energy of the residual data for the lower layer block, and (ii) determines a coding method for the enhancement-layer residual prediction flag according to that energy—if the energy is low, one coding method is used; if high, another. In CABAC terms, that “coding method” is a choice of probability model / context, and the selection depends on a function of the lower-layer residual data. When mapped to Wang’s framework, the lower-layer residual/energy corresponds to the first reconstruction-layer value, and the enhancement-layer coding corresponds to the second reconstruction layer. Thus, Lee is being interpreted as “using a probability model which depends on the first-reconstruction-layer neural network parameter value” when decoding second-layer parameters by context-adaptive entropy decoding.)
Claims 28 are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Gao, Nguyen and in further view of Iwahashi et al., (Iwahashi, M., Yoshida, T., Mokhtar, N. B., & Kiya, H. (2015). Bit-depth scalable lossless coding for high dynamic range images. EURASIP Journal on Advances in Signal Processing, 2015(1), 22.), hereafter referred to as Iwahashi.
Claim 28: Wang, Gao, and Nguyen teaches the limitations of claim 27. Nguyen further teaches:
Apparatus of claim 27, wherein the collection of probability context sets comprises three probability context sets, (Nguyen, col. 12, line 25, “Three different context sets may be defined, each context set containing the four contexts selected using the greater-than-one flags a, b, c, and d.”, Nguyen explicitly states “three different context sets”, each holding several contexts used for entropy coding.)
and the apparatus is configured to select a first probability context set out of the collection of probability context sets as the selected probability context set (Nguyen, col. 12, line 21, “There may further be multiple context sets that are selected based on neighboring or nearby significant-coefficient flags M, N, and P.” And: “context_set = min(2, M+N+P)”, Nguyen’s context_set variable indexes one of three context sets. For given values of M, N, P (derived from already-decoded data), the encoder/decoder selects a specific context set (e.g., context_set = 0 can be regarded as the “first” set). This teaches selection of a first probability context set out of a collection of three context sets.)
select a second probability context set out of the collection of probability context sets as the selected probability context set (Nguyen, col. 12, line 25, “Three different context sets may be defined, each context set containing the four contexts selected using the greater-than-one flags a, b, c, and d.”, Because Nguyen defines three different context sets, and context_set can take different index values (0,1,2), the system inherently supports selecting a second context set (e.g., context_set = 1) as the active set in some cases, matching selection of “a second probability context set” from the same collection.)
select a third probability context set out of the collection of probability context sets as the selected probability context set (Nguyen, col. 12, line 25, “Three different context sets may be defined, each context set containing the four contexts selected using the greater-than-one flags a, b, c, and d.”, By defining three context sets and selecting among them via context_set = min(2, M+N+P), Zhang necessarily allows cases where the index points to a third context set (e.g., context_set = 2). This corresponds to selecting a “third probability context set” from the collection.)
Iwahashi, in the same field of context model entropy coding, teaches the following limitation which Wang, Gao, and Nguyen fails to teach:
if the first-reconstruction-layer neural network parameter value is negative, (Iwahashi, page 3, col. 1, paragraph 3, “In contrast, xH in Equation 6 for a type A image can be negative, zero, or positive.”, Iwahashi’s HDR pixel value x_H can explicitly be negative, zero, or positive. These signed reconstructed pixel values play the role of first-layer reconstructed parameters. Thus Iwahashi expressly supports a reconstruction-layer value taking a negative value.)
if the first-reconstruction-layer neural network parameter value is positive, (Iwahashi, page 3, col. 1, paragraph 3, “In contrast, xH in Equation 6 for a type A image can be negative, zero, or positive.”, The Iwahashi passage explicitly includes positive x_H values as one of the possible cases. These signed HDR pixel values correspond to first-layer reconstruction values that may be positive.)
if the first-reconstruction-layer neural network parameter value is zero. (Iwahashi, page 3, col. 1, paragraph 3, “In contrast, xH in Equation 6 for a type A image can be negative, zero, or positive… Note that pixel values that are less than or equal to zero are first clipped to the minimum positive pixel value in the image.”, Iwahashi explicitly includes zero as a valid value of x_H, and further describes special handling of pixel values “less than or equal to zero,” which includes zero. This demonstrates that reconstructed base-layer values can be exactly zero and that the system recognizes and treats that case separately.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have combined Wang, Gao, and Nguyen with Iwahashi, a motivation of which would have been to handle first-layer reconstruction values that can be negative, zero, or positive in a layered coding system, and to use that sign information (negative / zero / positive) as a natural discriminator when selecting probability context sets for the second-layer coding as in claim 28. Iwahashi show both a two-layer reconstruction framework and a first-layer quantity x_H that may be negative, zero, or positive, which directly supports using the sign of the first reconstruction layer’s parameter (negative/zero/positive) as a selector for different probability context sets in the combined Wang–Gao–Nguyen neural-network coding system.
Claims 29 are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Gao, Nguyen, and in further view of Ji et al., (US 20140003488 A1), hereafter referred to as Ji.
Claim 29: Wang, Gao, and Nguyen teaches the limitations of claim 27. Ji, in the same field of context model entropy coding, teaches the following limitation which Wang, Gao, and Nguyen fails to teach:
Apparatus of claim 27, wherein the collection of probability context sets comprises two probability context sets, (Ji, claim 4, “selecting a context set comprises selecting a context set from among a plurality of context sets including a first context set and a second context set, and wherein a mapping of contexts in the first context set to the position of the greater-than-one flag and the position of the last significant coefficient is different from a mapping of contexts in the second context set.”, Ji explicitly states that the decoder selects a context set from “a plurality of context sets including a first context set and a second context set.” This matches a collection of probability context sets that “comprises two probability context sets” (the first and second context sets).)
and the apparatus is configured to select a first probability context set out of the collection of probability context sets as the selected probability context set (Ji, claim 4, “The method claimed in claim 3, wherein selecting a context set comprises selecting a context set from among a plurality of context sets including a first context set and a second context set…”, Ji’s decoder selects a context set from multiple sets, one of which is explicitly called the “first context set.” When that first context set is chosen, it functions as the selected context set for decoding. This directly corresponds to “select[ing] a first probability context set… as the selected probability context set.”)
if the first-reconstruction-layer neural network parameter value is greater than a predetermined value, (Ji, paragraph 62, “The likelihood of greater-than-one coefficients is correlated to whether a greater-than-one coefficient was encountered in the preceding coefficient group. In some cases, the correlation is based on how many greater-than-one coefficients were encountered in the preceding coefficient group. For example, if more than a threshold number of greater-than-one flags were observed in the previous coefficient group, then it may result in selection of a different context set than would otherwise be used.”, Ji discloses a decoder that has at least two context sets and selects between them using a threshold test on a scalar parameter derived from previously coded coefficients. Ji explains that a “factor or parameter or threshold test” (Paragraph 62) is used to pick the context set, and gives the concrete example that “if more than a threshold number of greater-than-one flags were observed in the previous coefficient group, then it may result in selection of a different context set” (Paragraph 63). Under a broadest reasonable interpretation, the claimed “first-reconstruction-layer neural network parameter value” corresponds to such a scalar parameter input to Ji’s threshold test, the “predetermined value” corresponds to Ji’s “threshold number,” and Ji’s teaching of selecting a different context set when the parameter exceeds that threshold directly matches the conditional “select a first probability context set … if the … value is greater than a predetermined value.”)
and select a second probability context set out of the collection of probability context sets as the selected probability context set if the first-reconstruction-layer neural network parameter value is not greater than the predetermined value, (Ji, paragraph 62, “For example, if more than a threshold number of greater-than-one flags were observed in the previous coefficient group, then it may result in selection of a different context set than would otherwise be used… [63] It will be understood that in determining context for greater-than-one flag coding there may be two or more context sets… The context set used for coding the greater-than-one flags of [a] coefficient group is selected using a factor or parameter or threshold test.”, Ji teaches: (1) two context sets (“first context set and a second context set”), and (2) selecting between them using a “factor or parameter or threshold test” involving “more than a threshold number” of greater-than-one flags. If the count is above the threshold, one context set (e.g., the first) is chosen; when it is not above the threshold, a different context set (e.g., the second) is used. This exactly corresponds to “select … a second probability context set … if the … value is not greater than the predetermined value” in the claim.)
or select the first probability context set out of the collection of probability context sets as the selected probability context set if an absolute value of the first-reconstruction-layer neural network parameter value is greater than the predetermined value, and select the second probability context set out of the collection of probability context sets as the selected probability context set if the absolute value of the first-reconstruction-layer neural network parameter value is not greater than the predetermined value. (Ji, paragraph 43, “The magnitudes for those non-zero coefficients may then be encoded. In some standards, magnitudes (i.e. levels) are encoded by encoding one or more level flags. If additional information is required to signal the magnitude of a quantized transform domain coefficient, then remaining-level data may be encoded. In one example implementation, the levels may be encoded by first encoding a sequence of greater-than-one flags indicating which non-zero coefficients having an absolute value level greater than one. Greater-than-two flags may then be encoded to indicate which non-zero coefficients have a level greater than two. Remaining level data may then be encoded for any of the coefficients having an absolute value greater than two. The value encoded in the remaining-level integer may be the actual value minus three. The sign of each of the non-zero coefficients is also encoded. Each non-zero coefficient has a sign bit indicating whether the level of that non-zero coefficient is negative or positive.”, In this alternative half of claim 29, the only difference from the first half is that the comparison is made on “an absolute value” of the first-reconstruction-layer parameter. The quoted passage shows that the greater-than-one flag is explicitly based on “non-zero coefficients having an absolute value level greater than one,” i.e., a decision on |level| > 1. Thus, the absolute value of the first-reconstruction-layer neural network parameter value in the claim corresponds to the absolute value level in the reference, and the predetermined value corresponds to the constant 1, so the same first/second context-set selection already mapped for the first half is now driven by this absolute-value comparison.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have combined Wang, Gao, and Nguyen with Ji’s context-set-selection work (e.g., “Methods and devices for context set selection”), a motivation of which would have been to further specialize the context models by introducing distinct context sets and selecting among them based on region or significance information, thereby reducing redundancy when entropy-coding second-layer neural network parameters in different “regions” or neighborhoods of the network. Ji, abstract, explicitly teaches multiple context sets and region-based selection: “Methods of encoding and decoding for video data for encoding or decoding a sequence of greater-than-one flags for a coefficient group are provided. Context-based encoding and decoding selects a context for encoding or decoding the greater-than-one flag based upon the position of the greater-than-one flag in the sequence of greater-than-one flags. Selection of the context may also be based upon the position of the last-non-zero coefficient in the coefficient group.”, A POSITA already using Gao’s HDCM and Nguyen’s adaptive thresholds would see Ji’s distinct context sets and region-dependent selection as a natural extension: partition weights/parameters into regions (layers, channels, spatial neighborhoods, etc.) and use different context sets for each, as taught for transform coefficients.
Claim 30 is rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Nguyen and in further view of Lee.
Claim 30: Wang teaches:
Apparatus for encoding neural network parameters, which define a neural network, by using first neural network parameters for a first reconstruction layer which comprise, per neural network parameter, a first-reconstruction-layer neural network parameter value, and the apparatus comprising a processor being configured to (Wang, page 512, section 2, paragraph 3, “To address this problem, we adopt the scalable coding concept in image/video coding [16, 13, 15], and represent the weights hierarchically, i.e., each weight is represented by a base-layer component and several enhancement-layer components;”, Wang explicitly decomposes each DNN weight into a base-layer component plus enhancement components. The base-layer component is a first-layer representation per weight, corresponding to the claimed first-reconstruction-layer neural network parameters and their per-parameter values.)
encode second neural network parameters for a second reconstruction layer into a data stream, which comprise, per neural network parameter, a second-reconstruction-layer neural network parameter value, (Wang, page 512, col. 1, last paragraph, “By repeating this procedure, we can obtain a n-layer hierarchical representation of a weight, i.e.,
PNG
media_image8.png
23
237
media_image8.png
Greyscale
where w is a uncompressed weight, b1 and ei are the centroid of the base layer and the i-th enhancement layer respectively.”,
Page 512, col. 2, paragraph 2, “After the hierarchical quantization, we can build a codebook that stores the centroid and cluster index information of all quantization layers. For a network layer of DNN with N weights, there are 2n centroids, and the number of cluster indices is Nn. If each uncompressed weight or centroid is represented by b bits (b=32 for single-precision floatingpoint number), the compression rate of the n-bit hierarchical quantization scheme is
PNG
media_image9.png
34
213
media_image9.png
Greyscale
”, In Wang, each uncompressed weight w is represented hierarchically as w ≈ b₁ + e₁ + … + eₙ₋₁, where b₁ and eᵢ are the centroids of the base layer and i-th enhancement layer, respectively (Wang, p.512, eq. (1)). Wang further states that, after hierarchical quantization, “we can build a codebook that stores the centroid and cluster index information of all quantization layers” and that these centroids and indices are represented in b-bit form to achieve a particular compression rate (Wang, p.512). Under a broadest reasonable interpretation, representing the enhancement-layer centroids eᵢ and their cluster indices as bits in this codebook is an encoding of “second neural network parameters for a second reconstruction layer into a data stream,” with one such second-layer value (eᵢ) associated with each neural network parameter. They are encoded as centroids and indices in the bitstream (data stream), directly corresponding to the claimed “second neural network parameters for a second reconstruction layer” encoded per parameter.)
wherein the neural network parameters are reconstructible by, for each neural network parameter, combining the first-reconstruction-layer neural network parameter value and the second-reconstruction-layer neural network parameter value. (Wang, page 512, col. 1, last paragraph, “By repeating this procedure, we can obtain a n-layer hierarchical representation of a weight, i.e.,
PNG
media_image8.png
23
237
media_image8.png
Greyscale
where w is a uncompressed weight, b1 and ei are the centroid of the base layer and the i-th enhancement layer respectively.”, Wang reconstructs each weight w by combining the base-layer value b_1 with one or more enhancement-layer values e_i. Even with just one enhancement layer, w ≈ b_1 + e_1, which is exactly the claimed behavior of reconstructing a neural network parameter by combining a first-reconstruction-layer value with a second-reconstruction-layer value.)
Nguyen, in the same field of context encoding/decoding, teaches the following limitations which Wang fails to teach:
selecting a probability context set out of a collection of probability context sets depending on the first-reconstruction-layer neural network parameter value, (Nguyen, col. 12, lines 21-30, “There may further be multiple context sets that are selected based on neighboring or nearby significant-coefficient flags M, N, and P… context_set = min(2, M+N+P)”, Nguyen explicitly describes multiple context sets and selecting a particular context set (context_set) based on a function of previously decoded significant-coefficient flags M, N, P. Those flags encode already-reconstructed transform-coefficient information, i.e., reconstruction-layer values. Thus Nguyen teaches selecting a probability context set (the context_set index) out of a collection of sets, depending on values derived from previously reconstructed parameters.)
selecting a probability context to be used out of the selected probability context set depending on the first-reconstruction-layer neural network parameter value. (Nguyen, col. 12, line 17-18, “There may be contexts… selected based on the sum of the neighboring or nearby greater-than-one flags… context = context_set*4 + min(3, a+b+c+d)”, For level coding, Nguyen first chooses a context set (context_set), then selects a specific context within that set using the sum of neighboring flags a, b, c, d. Those flags encode information about already-reconstructed coefficient levels (magnitude/significance), i.e., first-layer reconstruction values. This directly corresponds to “selecting a probability context … out of the selected probability context set depending on” values derived from earlier reconstruction-layer parameters.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have combined Wang with Nguyen, a motivation of which would have been to refine Gao’s generic context modeling by using adaptive thresholds and level information from previously reconstructed groups when selecting how many level flags or contexts to use, thereby further improving entropy-coding efficiency of Wang’s layered neural network parameters. Nguyen’s abstract explains: “Methods and devices for reconstructing coefficient levels from a bitstream of encoded video data for a coefficient group in a transform unit, using adaptive-threshold-based level coding. Threshold is set based upon level information from one or more previously-reconstructed coefficient groups in the transform unit. Threshold may be maximum number of level flags to decode for the coefficient group. Level information may include number of level flags decoded in previous coefficient groups.”, Given Gao’s HDCM already uses neighboring/statistical information for context modeling, a POSITA would see Nguyen’s adaptive-threshold mechanism as a way to condition context/flag coding not just on position but also on previously decoded level information. In Wang’s neural-network setting, the “coefficient groups” map naturally to groups of weights/parameters, so using Nguyen’s adaptive thresholds and context choice to drive the second-layer entropy decoding is a reasonably expected optimization.
Lee, in the same field of context model entropy coding, teaches the following limitation which Wang and Nguyen fails to teach:
encode the second-reconstruction-layer neural network parameter value into the data stream by context-adaptive entropy encoding, (Lee, paragraph 19, “According to yet another aspect of the present invention, there is provided a method of decoding a residual prediction flag indicating whether residual data for an enhancement layer block of a multi-layered video signal is predicted from residual data for a lower layer block corresponding to the residual data for the enhancement layer block, the method comprising calculating a value of a CBP of the lower layer block, determining a decoding method for the residual prediction flag according to the calculated value of the CBP, and decoding the residual prediction flag using the determined decoding method.”, Lee teaches encoding/decoding information for an enhancement layer of a multi-layered signal using context-adaptive entropy coding selected based on lower-layer information. Under a broadest reasonable interpretation, Lee’s enhancement-layer coded information corresponds to the claimed second-reconstruction-layer neural network parameter value being encoded into the data stream by context-adaptive entropy encoding.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have combined Wang and Nguyen with Lee’s multi-layer video coding, a motivation of which would have been to exploit redundancy between a base representation and an enhancement representation by using information from the base layer to guide entropy coding of the enhancement layer, thereby improving overall rate–distortion efficiency in a layered reconstruction scheme. Lee explicitly frames a base/enhancement split and inter-layer prediction “to increase coding efficiency” (Lee, paragraph 95, “Because a context model is selected for each slice according to the above-described CABAC, probability values of probability models are initialized to a table of constant values for the slice. CABAC provides better coding efficiency than conventional VLC when a predetermined amount of information is accumulated because a context model has to be continuously updated based on statistics of recently-coded data symbols.”)
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chen, S., Wang, W., & Pan, S. J. (2019, July). Deep neural network quantization via layer-wise optimization using limited training data. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 33, No. 01, pp. 3329-3336).
US 8,989,265 B2 - Method for modeling coding information of video signal for compressing/decompressing coding information
Schwarz, H., Nguyen, T., Marpe, D., & Wiegand, T. (2019, March). Hybrid video coding with trellis-coded quantization. In 2019 Data Compression Conference (DCC) (pp. 182-191). IEEE.
Choi, Y., El-Khamy, M., & Lee, J. (2020). Universal deep neural network compression. IEEE Journal of Selected Topics in Signal Processing, 14(4), 715-726.
Gong, Y., Liu, L., Yang, M., & Bourdev, L. (2014). Compressing deep convolutional networks using vector quantization. arXiv preprint arXiv:1412.6115.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HYUNGJUN B YI whose telephone number is (703)756-4799. The examiner can normally be reached M-F 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/H.B.Y./Examiner, Art Unit 2146
/USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146