Prosecution Insights
Last updated: October 02, 2026
Application No. 17/843,772

Concepts for Coding Neural Networks Parameters

Final Rejection §103
Filed
Jun 17, 2022
Priority
Dec 20, 2019 — EU 19218862.1 +2 more
Examiner
YI, HYUNGJUN B
Art Unit
2146
Tech Center
2100 — Computer Architecture & Software
Assignee
Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V.
OA Round
3 (Final)
32%
Grant Probability
At Risk
4-5
OA Rounds
0m
Est. Remaining
58%
With Interview

Examiner Intelligence

Grants only 32% of cases
32%
Career Allowance Rate
9 granted / 28 resolved
-22.9% vs TC avg
Strong +26% interview lift
Without
With
+25.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
20 currently pending
Career history
64
Total Applications
across all art units

Statute-Specific Performance

§101
28.1%
-11.9% vs TC avg
§103
54.9%
+14.9% vs TC avg
§102
11.8%
-28.2% vs TC avg
§112
4.4%
-35.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 28 resolved cases

Office Action

§103
DETAILED ACTION This action is responsive to the claims filed on 06/09/2026. Claims 1-8, 16, 18-21, 24-25, 28-29, 34-46, 106-107, 110, and 112-113 are pending for examination. This action is Final. Information Disclosure Statement The information disclosure statements (IDS) submitted on 01/27/2026, 02/19/2026, 05/18/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Response to Arguments Applicant’s amendments to claims 24, 25, 28, and 34 have been considered. Applicant amended these claims so that they no longer depend from canceled claim 23. Accordingly, the prior objections to claims 24, 25, 28, and 34, based upon their dependency from canceled claim 23, are withdrawn. Applicant’s arguments regarding the statutory double-patenting rejection of claims 1, 24-25, and 34-38 have been fully considered and are persuasive. In view of the amendments to claim 1 and the resulting differences between the presently claimed subject matter and the claims of the reference application, the claims are no longer considered to be directed to identical or coextensive subject matter. Accordingly, the statutory double-patenting rejection of claims 1, 24-25, and 34-38 is withdrawn. Applicant has also submitted terminal disclaimers in response to the provisional nonstatutory double-patenting rejections of claims 46, 106, 107, 112, and 113 and has expressly requested reconsideration of those rejections. The terminal disclaimers have been considered and are sufficient to obviate the provisional nonstatutory double-patenting rejections. Accordingly, the provisional nonstatutory double-patenting rejections of claims 46, 106, 107, 112, and 113 are withdrawn. Applicant’s arguments regarding the rejection under 35 U.S.C. § 101 have been fully considered and are persuasive. Upon reconsideration of the claims as a whole, the claimed coordinated state-transition process, state-dependent reconstruction-level-set selection, state updating across a sequence of neural-network parameters, and arithmetic coding using a probability model selected depending on the state constitute a sufficiently specific implementation of neural-network parameter encoding and decoding. Accordingly, the rejection of claims 1-8, 16, 18-21, 24-25, 28-29, 34-46, 106-107, 110, and 112-113 under 35 U.S.C. § 101 is withdrawn. Applicant argues that “the cited combination fails to teach or suggest the claimed use of a shared state used for both reconstruction level selection and arithmetic coding.” (Remarks, pages 34-35). Applicant’s argument has been fully considered but does not overcome the rejection as presently set forth. In view of the amendments, Coban is now relied upon for the coordinated state-dependent limitations. Coban teaches that “video encoder 200 and/or video decoder 300 may select the quantizer based on a state machine that is driven by its previous state and the parity of the level (e.g., absolute value) of the previously coded coefficient,” and further teaches that coefficients in states 0 and 1 use quantizer Q0, while coefficients in states 2 and 3 use quantizer Q1 (Coban, US 11,451,840 B2, col. 13, lines 1-24). Thus, the state associated with the current coefficient determines which quantizer, and therefore which set of reconstruction levels, is used. Coban further teaches that “the significance (greater than zero or ‘gt0’) map and greater than 1 (‘gt1’) map context used in context based arithmetic coding depends on the state that is driven by the previous state and the parity of the coefficient level” (Coban, col. 13, line 32 through col. 14, line 4). Accordingly, Coban expressly teaches coordinated use of the evolving state for both reconstruction-level-set selection and context-based arithmetic coding. Applicant argues that “Sze does not disclose or suggest using such a state to control quantization or reconstruction level selection,” that “Kasner does not disclose or suggest using this state for entropy coding, nor does it disclose any arithmetic coding process,” and that “neither Kasner nor Sze provides any teaching or suggestion to reuse a state from a quantization process as the controlling state for probability model selection in arithmetic coding.” (Remarks, page 35) These arguments have been fully considered but do not overcome the present rejection. The amended shared-state limitation is not presently mapped to Sze, and Sze is not relied upon to establish that its internal probability-estimation state controls reconstruction-level-set selection. Likewise, Kasner is not relied upon alone to establish the coordinated arithmetic-coding limitation. Instead, Coban is relied upon for the amended relationship between the state-selected quantizer and the state-selected arithmetic-coding context. Coban expressly teaches that the coding or parsing approach for a quantized coefficient may depend on whether Q0 or Q1 produced the coefficient, that “one set of context models may be used for the quantized coefficients by Q0 and another set may be used for those by Q1,” and that this adaptation is possible when “the same state machine” is implemented in transform-coefficient coding so that “the current state (i.e., the right quantizer) is determined while coding/parsing the current quantized coefficient” (Coban, col. 15, lines 38-52). Sze remains applied only to the separate CABAC, binarization, context-selection, and probability-model-update limitations of the dependent claims for which it is expressly mapped. Applicant argues that “the amended claims do not merely require that both quantization and entropy coding are state-dependent. Rather, they require that the probability model used for arithmetic coding is selected depending on the same state that is used to determine the reconstruction level set,” and that “[a] combination of Kasner and Sze would, at most, result in two independent mechanisms.” (Remarks, pages 35-36). The Examiner respectfully disagrees that the present rejection results in two independent mechanisms. Coban teaches that “[i]n the TCQ scheme of this disclosure, video encoder 200 and video decoder 300 may use separate context models for significance and greater than one coefficient contexts based on the current state (even or odd reconstruction level quantizer) used,” and further explains that “if the current state is 0 or 1, then context set 0 is used, whereas if the current state is 2 or 3, then context set 1 is used” (Coban, col. 14, lines 9-20). The states 0 and 1 are the same states that select Q0, while states 2 and 3 are the same states that select Q1. Coban further states that “[s]witching between the context sets may also change the quantizer, hence the reconstruction levels” (Coban, col. 14, lines 20-23). Therefore, Coban teaches that the arithmetic-coding context model and the reconstruction-level quantizer are selected according to the same state, rather than according to two unrelated state mechanisms. Applicant argues that “[t]he cited references do not provide any teaching or motivation to unify these mechanisms into a single shared state as required by the claims” and that “[t]he Office action does not provide an articulated reasoning with rational underpinning why a person of ordinary skill in the art would have modified the cited teachings to arrive at the claimed shared-state architecture.” (Remarks, page 36) The Examiner respectfully disagrees. Coban itself teaches the allegedly missing coordination and therefore does not require a person of ordinary skill in the art to independently invent a new mechanism for unifying unrelated states. Coban teaches modifying transform-coefficient coding so that the coding approach depends on the active quantizer, using one context-model set for Q0 and another context-model set for Q1, and implementing the same state machine during coding or parsing so that the current quantizer state is available when the current coefficient is coded (Coban, col. 15, lines 38-52). Han teaches quantizing neural-network weights and representing the weights using stored shared values and corresponding indices, while Kasner teaches trellis-based selection of reconstruction-level supersets. A person of ordinary skill in the art would have been motivated to incorporate Coban’s known coordinated quantizer/context-model arrangement into the Han/Kasner neural-network compression system so that the arithmetic-coding probability model corresponds to the reconstruction-level set being used for the current parameter, thereby adapting the entropy-coding operation to the statistical characteristics of the active quantizer and improving coding efficiency. Applicant argues that “[i]mplementing such a shared state would require modifying the interaction between quantization and entropy coding in a manner that is not suggested by the prior art and would require more than a routine or predictable use of prior art elements according to their established functions.” (Remarks, page 36). The Examiner respectfully disagrees. Coban already teaches the interaction Applicant characterizes as absent. Specifically, Coban teaches that “video encoder 200 and video decoder 300 may use separate context models for significance and greater than one coefficient contexts based on the current state (even or odd reconstruction level quantizer) used,” and further explains that the encoder and decoder “may maintain a separate context model for even and odd quantizers.” Coban additionally teaches that “if the current state is 0 or 1, then context set 0 is used, whereas if the current state is 2 or 3, then context set 1 is used,” and that “[s]witching between the context sets may also change the quantizer, hence the reconstruction levels” (Coban, col. 14, lines 10-23). Thus, Coban expressly teaches selecting different arithmetic-coding context-model sets according to the current state identifying the active even or odd reconstruction-level quantizer. Coban further teaches implementing the same state-machine information during coefficient coding, explaining that “one set of context models may be used for the quantized coefficients by Q0 and another set may be used for those by Q1,” and that this may be accomplished by “implement[ing] the same state machine during transform coefficient coding” such that “the current state (i.e., the right quantizer) is determined while coding/parsing the current quantized coefficient” (Coban, col. 15, lines 38-52). Accordingly, the proposed combination does not require altering the established functions of the quantizer, state machine, or arithmetic-coding context models. Rather, it applies Coban’s known coordinated state-dependent coding arrangement to Han’s known sequence of quantized neural-network parameters and Kasner’s known trellis-coded reconstruction-level selection. Each component continues to perform its established function, and the resulting arrangement predictably uses the active quantizer state to select an appropriate probability model for coding the corresponding quantization index. Applicant argues that “[e]ven if the cited references were combined, the result would not yield the claimed coordinated use of a single state across both quantization and entropy coding.” (Remarks, page 36). This argument is not persuasive because Coban expressly discloses the coordinated result. Coban teaches that “coefficients in state 0 and state 1 use the Q0 (even integer multiples of stepsize) quantizer, and coefficients in states 2 and 3 use the Q1 (odd integer multiples of stepsize) quantizer” (Coban, col. 13, lines 14-19). Coban further teaches, with respect to entropy coding, that “if the current state is 0 or 1, then context set 0 is used, whereas if the current state is 2 or 3, then context set 1 is used,” and expressly states that “[s]witching between the context sets may also change the quantizer, hence the reconstruction levels” (Coban, col. 14, lines 15-23). Therefore, the same state pairs that select Q0 or Q1 and their respective reconstruction levels also select context set 0 or context set 1 for context-based arithmetic coding. Coban additionally teaches that the state machine is “driven by its previous state and the parity of the level . . . of the previously coded coefficient” and provides specific transitions from each current state according to whether the parity equals zero or one (Coban, col. 13, lines 7-31). Thus, in the applied combination, the evolving state determines the reconstruction-level set for the current neural-network parameter, the preceding quantization index supplies parity information used to update the state for the subsequent neural-network parameter, and that same state selects the context model constituting the probability model for arithmetic coding. Applicant’s remarks therefore do not demonstrate that the amended limitations are absent from the combined teachings of Han, Kasner, and Coban. Claim Objections Claim 25 is objected to because of the following informality: claim 25 recites “wherein the at least one bin comprises a significance bin,” but claim 25 presently depends directly from claim 1, and claim 1 does not previously introduce “at least one bin.” The term “at least one bin” is instead introduced in claim 24. Accordingly, “the at least one bin” in claim 25 lacks sufficient antecedent basis. Appropriate correction is required, such as amending the dependency of claim 25 or otherwise introducing the referenced “at least one bin” before its use in claim 25. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-6, 7, 16, 18-21, 36, 39-41, 44, 46, 106-107, 110 and 112-113 are rejected under 35 U.S.C. 103 as being unpatentable over Han et al., (Han, S., Mao, H., & Dally, W. J. (2015) Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149.), hereafter referred to as Han, in view of Kasner et al. (Kasner, J. H., Marcellin, M. W., & Hunt, B. R. (1999). Universal trellis coded quantization. IEEE Transactions on Image Processing, 8(12), 1677-1687.), hereafter referred to as Kasner, and in further view of Coban et al. (US 11451840 B2), hereafter referred to as Coban. Claim 1: Han teaches the following limitations: Apparatus for decoding neural network parameters, which define a neural network, from a data stream, configured to sequentially decode the neural network parameters (Han, abstract, “Our method first prunes the network by learning only the important connections. Next, we quantize the weights to enforce weight sharing, finally, we apply Huffman coding. After the first two steps we retrain the network to fine tune the remaining connections and the quantized centroids.”, Han describes decoding quantized indices from neural network connections.) Kasner, in the related field of scalar quantization and entropy coding, teaches the following limitations which the above fails to teach: decoding a quantization index for the current neural network parameter from the data stream, wherein the quantization index indicates one reconstruction level out of the selected set of reconstruction levels for the current neural network parameter, (Kasner, page 1678, col. 2, paragraph 1, “The UTCQ quantizer returns the S0 indices and the negative of the S1 indices, allowing one probability model to be used for entropy coding. The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly”, In Kasner, for each source sample, the UTCQ quantizer returns an index within the currently active superset; the decoder, given the current state and the received index, looks up exactly one reconstruction codeword in that superset. Hence, the superset’s codewords are being interpreted as the “reconstruction levels,” and the index within that superset is the claimed “quantization index” that uniquely points to one codeword (one reconstruction level) out of the set associated with the current state.) dequantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels that is indicated by the quantization index for the current neural network parameter. (Kasner, page 1678, col. 2, paragraph 2, “During dequantization, two types of reconstruction levels are employed, uniform and trained. For PNG media_image1.png 18 97 media_image1.png Greyscale , uniform levels are used (i.e., the codeword is the center of the quantization cell). The remaining codewords are trained on the source data itself, except CW0 which is typically set to 0. The trained codeword PNG media_image2.png 18 118 media_image2.png Greyscale , is determined by taking the sample mean of all source symbols that map to PNG media_image3.png 17 45 media_image3.png Greyscale and the negative of all source symbols mapping to PNG media_image4.png 15 58 media_image4.png Greyscale ”, Kasner’s decoder in whichever selected subset (S0 or S1), performs a lookup of the exact reconstruction level corresponding to the decoded index i or -i.) A person of ordinary skill in the art (POSITA) before the effective filing date of the claimed invention would have recognized that Kasner’s trellis-coded quantization is applicable to sequential scalar parameters, and a POSITA would have recognized it as reasonably applicable to Han’s sequential neural-network parameter indices. Han’s Deep Compression already teaches a quantization technique that uses a learned codebook of centroids per layer and represents each weight by an index into that codebook. Kasner’s TCQ similarly uses codebooks portioned into subsets and represents each sample by a subset-specific index. To combine the two methods, a POSITA would have applied Kasner’s trellis/state-based subset-selection mechanism to the sequence of quantized neural-network parameters/indices used in Han. Specifically, the decoder would maintain a state as taught by Kasner and, for each parameter, would use the current state to select the applicable reconstruction-level set (superset/subset) before interpreting the decoded index as the reconstruction level within that selected set. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Han (i.e. deep neural network quantization methos) by incorporating the teachings of Kasner (i.e. index based Trellis-Coded Quantization methods). A motivation of which is to provide a codebook quantization technique without needing additional storage and training of codebooks. (Kasner, page 1687, “The advantages of UTCQ are simplicity and flexibility. Unlike previous ECTCQ systems, no prior codebook training is needed and no codebooks are stored. We have shown that the distortion-rate performance of UTCQ is comparable with that of optimal ECTCQ for memoryless sources at most encoding rates.”)Coban, in the same field of trellis-coded quantization and entropy coding, further teaches the following amended limitations which the above fails to teach: selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets by means of a state transition process (Coban, col. 13, lines 1-31, “In particular, when using two scalar quantizers, a first quantizer Q0 may map transform coefficient levels (numbers below the points, e.g., absolute values) to even integer multiples of quantization step size Δ. The second quantizer Q1 may map the transform coefficient levels to odd integer multiples of the quantization step size Δ or to zero. Transform coefficients may be quantized in coding order, and video encoder 200 and/or video decoder 300 may select the quantizer based on a state machine that is driven by its previous state and the parity of the level (e.g., absolute value) of the previously coded coefficient. The state machine may be represented by a data structure. An example of the state machine is described in FIG. 4. FIG. 4 is a conceptual diagram illustrating an example state transition scheme for two scalar quantizers used to perform quantization…. As shown, the data structure of FIG. 4 transitions from state 0 to state 0, state 1 to state 2, state 2 to state 1, and state 3 to state 3 when a parity equals 0. In this example, the data structure of FIG. 4 transitions from state 0 to state 2, state 1 to state 0, state 2 to state 3, and state 3 to state 1 when the parity equals 1. As shown, the data structure of FIG. 4 may include a state machine.” Coban teaches selecting between two quantizers Q0 and Q1, which provide respective sets of reconstruction levels, by means of a state-transition process implemented by the disclosed state machine. When Coban’s state-transition quantizer selection is applied to sequentially processed neural-network parameters, Coban teaches selecting a reconstruction-level set for a current neural-network parameter by means of the state-transition process.) by determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, (Coban, col. 13, lines 14-24, “FIG. 4 is a conceptual diagram illustrating an example state transition scheme for two scalar quantizers used to perform quantization. In this example, coefficients in state 0 and state 1 use the Q0 (even integer multiples of stepsize) quantizer, and coefficients in states 2 and 3 use the Q1 (odd integer multiples of stepsize) quantizer. That is, video encoder 200 and/or video decoder 300 may determine a quantizer Q0 when the second state comprises state 0 or state 1. In some examples, video encoder 200 and/or video decoder 300 may determine a quantizer Q1 when the second state comprises state 2 or state 3.” Coban’s Q0 includes reconstruction levels corresponding to even integer multiples of the quantization step size, while Q1 includes reconstruction levels corresponding to odd integer multiples of the quantization step size or zero. Coban therefore teaches that states 0 and 1 determine the Q0 reconstruction-level set and states 2 and 3 determine the Q1 reconstruction-level set. Thus, when applied to Han’s neural-network parameters, the state associated with the current neural-network parameter determines which reconstruction-level set is used for that current neural-network parameter.) and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter, (Coban, col. 13, lines 7-11 and 25-31, “Transform coefficients may be quantized in coding order, and video encoder 200 and/or video decoder 300 may select the quantizer based on a state machine that is driven by its previous state and the parity of the level (e.g., absolute value) of the previously coded coefficient…. As shown, the data structure of FIG. 4 transitions from state 0 to state 0, state 1 to state 2, state 2 to state 1, and state 3 to state 3 when a parity equals 0. In this example, the data structure of FIG. 4 transitions from state 0 to state 2, state 1 to state 0, state 2 to state 3, and state 3 to state 1 when the parity equals 1.”; Coban, col. 33, lines 1-18, “Video encoder 200 and/or video decoder 300 update the data structure to a second state according to the first state and a parity of a partial set of syntax elements representing a partial set of a plurality of coefficient levels for the previous transform coefficient…. In this example, video encoder 200 and/or video decoder 300 may determine the second state using, with the parity as an input, the state transition scheme of FIG. 4.” Coban teaches processing coefficients in coding order and updating the state used for the current or subsequent coefficient based upon the previous state and parity information obtained from the previously coded or decoded coefficient level. The decoded coefficient level is represented by its corresponding quantization index, and its parity is therefore information obtained from that quantization index. Accordingly, when Coban’s sequential state update is applied to Han’s sequentially decoded neural-network parameters, Coban teaches updating the state for the subsequent neural-network parameter depending on the quantization index decoded for the immediately preceding neural-network parameter.) decoding the quantization index for the current neural network parameter from the data stream using arithmetic coding using a probability model which is selected depending on the state associated with the current neural network parameter. (Coban, col. 13, line 32 through col. 14, line 23, “In this scheme, coefficients may be quantized in coefficient coding order. In the JVET-J0014 scheme, the entire coefficient level must be coded before moving to the coding of the next coefficient. The reason for this is that the significance (greater than zero or ‘gt0’) map (the absolute value of coefficient level greater than 0) and greater than 1 (gt1) map context used in context based arithmetic coding depends on the state that is driven by the previous state and the parity of the coefficient level…. In the TCQ scheme of this disclosure, video encoder 200 and video decoder 300 may use separate context models for significance and greater than one coefficient contexts based on the current state (even or odd reconstruction level quantizer) used. Basically, video encoder 200 and video decoder 300 may maintain a separate context model for even and odd quantizers. In one example, if the current state is 0 or 1, then context set 0 is used, whereas if the current state is 2 or 3, then context set 1 is used. Switching between the context sets can be achieved by changing the parity of the level of the previous coefficients that minimizes the RD cost. Switching between the context sets may also change the quantizer, hence the reconstruction levels.” Coban expressly teaches context-based arithmetic coding in which the context model used for a current coefficient depends on the current state. A context model provides the probability model used in context-based arithmetic coding. Coban further teaches selecting context set 0 when the state is 0 or 1 and context set 1 when the state is 2 or 3, while those same respective state pairs select the even reconstruction-level quantizer Q0 and the odd reconstruction-level quantizer Q1. Accordingly, Coban teaches decoding using an arithmetic-coding probability model selected depending on the same state associated with the current parameter that determines the applicable reconstruction-level set.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combined neural-network quantization and trellis-coded quantization teachings of Han and Kasner by incorporating Coban’s coordinated state-based quantizer and context-model selection. Han teaches quantizing and coding sequential neural-network parameters, and Kasner teaches trellis-coded quantization using reconstruction-level subsets and quantization indices. Coban further teaches using the current state of the trellis process both to determine which reconstruction-level quantizer applies to the current coefficient and to determine which context model is used for context-based arithmetic coding of that coefficient. A person of ordinary skill in the art would have recognized that using Coban’s coordinated state-based quantizer and probability-model selection in the Han/Kasner neural-network compression system would predictably select a probability model corresponding to the active reconstruction-level set, thereby adapting the entropy-coding operation to the quantizer being used and improving coding efficiency. Claim 2: Han, Kasner, and Coban teaches the limitations of claim 1, Han further teaches: Apparatus of claim 1, wherein the neural network parameters relate to weights of neuron interconnections of the neural network. (Han, abstract, “To address this limitation, we introduce “deep compression”, a three stage pipeline: pruning, trained quantization and Huffman coding, that work together to reduce the storage requirement of neural networks by 35× to 49× without affecting their accuracy. Our method first prunes the network by learning only the important connections. Next, we quantize the weights to enforce weight sharing, finally, we apply Huffman coding.”, Han explicitly treats the neural network parameters as weights.) Claim 3: Han, Kasner, and Coban teaches the limitations of claim 1, Kasner further teaches: Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two. (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets PNG media_image5.png 18 185 media_image5.png Greyscale . Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen”, at each step exactly two supersets (reconstruction level sets) are available.) Claim 4: Han, Kasner, and Coban teaches the limitations of claim 1, Kasner further teaches: Apparatus of claim 1, configured to parametrize the plurality of reconstruction level sets by way of a predetermined quantization step size and derive information on the predetermined quantization step size from the data stream. (Kasner, page 1678, col. 2, paragraph 1, “The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly…. The quantization thresholds are simply the midpoints between the reconstruction levels within a subset. This allows for fast computation of superset indices requiring only scaling and rounding. No thresholds need to be precomputed, nor is a binary tree search necessary. For a given trellis, the encoder is completely characterized by the stepsize parameter Δ .", this discloses the reconstruction level sets are explicitly parameterized using a quantization step size gained from a stream of data (index stream). “During dequantization, two types of reconstruction levels are employed, uniform and trained. For PNG media_image1.png 18 97 media_image1.png Greyscale , uniform levels are used (i.e., the codeword is the center of the quantization cell)… The remaining codewords are trained on the source data itself”, the reconstruction level sets (which comprise the step size) are derived from the index stream.) Claim 5: Han, Kasner, and Coban teaches the limitations of claim 1, Kasner further teaches: Apparatus of claim 1, wherein the neural network comprises a one or more NN layers and the apparatus is configured to derive, for each NN layer, information on a predetermined quantization step size for the respective NN layer from the data stream, (Kasner, page 1678, col. 2, paragraph 1, “The UTCQ quantizer returns the S0 indices and the negative of the S1 indices, allowing one probability model to be used for entropy coding. The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly”, Kasner states that, “for a given trellis, the encoder is completely characterized by the stepsize parameter Δ,” and that UTCQ uses uniform thresholds and reconstruction levels based on this Δ. The value of Δ (or its encoded representation) is the “information on a predetermined quantization step size” that the decoder must know or derive from the bitstream to reconstruct the uniform codebook. In the combined mapping with Wang, Wang supplies the multiple NN layers; for each such layer, the trellis-based quantizer operates with a chosen Δ, and the Δ value signaled or implied in the data stream for that layer is what is being interpreted as “information on a predetermined quantization step size for the respective NN layer … derived from the data stream.” The decoder can recover the index stream by tracking state and negating codewords, and derives, quantization information like step sizes from this data stream.) and parametrize, for each NN layer, the plurality of reconstruction level sets using the predetermined quantization step size derived for the respective NN layer so as to be used for dequantizing the neural network parameters belonging to the respective NN layer. (Kasner, page 1678, col. 2, paragraph 1, “The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly…. The quantization thresholds are simply the midpoints between the reconstruction levels within a subset. This allows for fast computation of superset indices requiring only scaling and rounding. No thresholds need to be precomputed, nor is a binary tree search necessary. For a given trellis, the encoder is completely characterized by the stepsize parameter Δ .", “During dequantization, two types of reconstruction levels are employed, uniform and trained. For PNG media_image1.png 18 97 media_image1.png Greyscale , uniform levels are used (i.e., the codeword is the center of the quantization cell)… The remaining codewords are trained on the source data itself”, the quantization thresholds are midpoints between reconstruction levels and that the encoder is completely characterized by the step size parameter Δ. This directly teaches that the reconstruction level sets are parameterized using the predetermined step size, which enables dequantization tailored per NN layer.) Claim 6: Han, Kasner, and Coban teaches the limitations of claim 1. Coban further teaches: Apparatus of claim 1, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two (Coban, col. 12, line 54, “FIG . 3 is a conceptual diagram illustrating an example of how two scalar quantizers can be used to perform quantization… a first quantizer ( e.g. , Q1 ) may be configured with a first set of quantization parameters and a second quantizer ( e.g. , Q2 ) may be configured with a second set of quantization parameters that are different in value from the first set”, two reconstruction level sets Q1 and Q2 are present in Coban.) and the plurality of reconstruction level sets comprises a first reconstruction level set that comprises zero and even multiples of a predetermined quantization step size, (Coban, col. 13, line 1, “In particular , when using two scalar quantizers , a first quantizer Q0 may map transform coefficient levels ( numbers below the points , e.g. , absolute values ) to even integer multiples of quantization step size Δ.”) and a second reconstruction level set that comprises zero and odd multiples of the predetermined quantization step size. (Coban, col. 13, line 4, “The second quantizer Q1 may map the transform coefficient levels to odd integer multiples of the quantization step size Δ or to zero .”) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Han (i.e. deep neural network quantization methos), Kasner, and Sze by incorporating the teachings of Coban (i.e. Trellis-Coded Quantization using parity based indexing). A motivation of which is to provide a codebook quantization technique using Trellis Coded Quantization as to increase computation efficiency. (Coban, col. 33, line 26, “In this way , video encoder 200 and / or video decoder 300 may quantize or inverse quantize a set of syntax elements representing the remaining levels ( e.g. , gt3 418A - 418N ) of the plurality of coefficient levels for transform coefficients of residual data for the block of video data without grouping all bypass coded bins for simpler parsing , thereby improving a computation efficiency of video encoder 200 and / or video decoder 300 .”) Claim 7: Han, Kasner, and Coban teaches the limitations of claim 1, Kasner further teaches: Apparatus of claim 1, wherein all reconstruction levels of all reconstruction level sets represent integer multiples of a predetermined quantization step size, and the apparatus is configured to dequantize the neural network parameters by (Kasner, page 1678, col. 2, paragraph 1, “UTCQ uses uniform thresholds and reconstruction levels at the encoder. The quantization thresholds are simply the midpoints between the reconstruction levels within a subset. This allows for fast computation of superset indices requiring only scaling and rounding. No thresholds need to be precomputed, nor is a binary tree search necessary. For a given trellis, the encoder is completely characterized by the stepsize parameter Δ.”, Kasner explains that UTCQ uses uniform thresholds and reconstruction levels and that the encoder is characterized by a stepsize parameter Δ. In a uniform scalar quantizer, the reconstruction values are positioned at regularly spaced points (centers of quantization cells) on the real line, each separated by Δ; these regular positions can be expressed as k·Δ for integer k. Thus, the uniform reconstruction levels in the UTCQ codebook are being interpreted as the claimed “reconstruction levels,” Δ is the “predetermined quantization step size,” and the fact that these levels lie on a uniform grid defined by Δ is what supports the statement that all reconstruction levels of all reconstruction level sets represent integer multiples of Δ. A uniform threshold scheme places reconstruction levels at exact multples of the uniform step size Δ (i.e. midpoints at k * Δ) so all codewords in every subset are integer-multiples of Δ.) PNG media_image6.png 87 275 media_image6.png Greyscale Figure 2 of Kasner deriving, for each neural network parameter, an intermediate integer value depending on the selected reconstruction level set for the respective neural network parameter and the entropy decoded quantization index for the respective neural network parameter, (Kasner, page 1679, col. 2, paragraph 3, “Given a source sample to quantize, subset quantization indices may be computed directly. Given a quantization index, the reconstruction level (for uniform codewords) may be computed.”, Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets PNG media_image5.png 18 185 media_image5.png Greyscale . Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen”, In UTCQ, Fig. 2 shows a uniform codebook partitioned into subsets, and Kasner explains that “a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen” (Kasner, p.1677). Kasner also states that “given a quantization index, the reconstruction level (for uniform codewords) may be computed” (Kasner, p.1679). Under a broadest reasonable interpretation, the index that is entropy-decoded for the active superset is the claimed “entropy decoded quantization index.” Because the underlying uniform codebook is ordered and parameterized by the step size Δ, that index corresponds to a particular uniform codeword location on the real line, which is treated as the claimed “intermediate integer value” that identifies which Δ-spaced reconstruction level is to be used.) and multiplying, for each neural network parameter, the intermediate value for the respective neural network parameter with the predetermined quantization step size for the respective neural network parameter. (Kasner, page 1678, col. 2, paragraph 2, “During dequantization, two types of reconstruction levels are employed, uniform and trained. For PNG media_image1.png 18 97 media_image1.png Greyscale , uniform levels are used (i.e., the codeword is the center of the quantization cell). The remaining codewords are trained on the source data itself, except CW0 which is typically set to 0. The trained codeword PNG media_image2.png 18 118 media_image2.png Greyscale , is determined by taking the sample mean of all source symbols that map to PNG media_image3.png 17 45 media_image3.png Greyscale and the negative of all source symbols mapping to PNG media_image4.png 15 58 media_image4.png Greyscale ”, once the decoder has its intermediate integer value (the subset quantization index), it directly looks up the corresponding reconstruction level in the selected subset. It is interpreted by the examiner that the subset employing uniform levels (e.g. the codeword is the center of the quantization cell) has to multiply the step size (Δ) with the intermediate value (the quantization index) in order to preserve the uniform level nature of the reconstruction sets.) Claim 16: Han, Kasner, and Coban teaches the limitations of claim 1, Kasner further teaches: Apparatus of claim 1, wherein the apparatus is configured to select, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets by means of a state transition process by (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work… Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets PNG media_image5.png 18 185 media_image5.png Greyscale . Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen [3].”, Kasner’s TCQ’s trellis is a state transition process that selects which codebook subset (superset) to use for each coefficient based on the current state. Kasner explicitly shows an eight-state trellis and explains that “at any given trellis state, the next codeword must come from one of two supersets,” and that given an initial state and a sequence of indices the decoder can reconstruct the sequence of chosen codewords. The movement from state to state along the trellis edges as each new symbol/index is processed is the claimed “state transition process.” At each step, the current state determines which superset (reconstruction-level set) is used, and the next state is determined by the current state and the chosen index, so the trellis operation as a whole is being interpreted as “selecting … the set of quantization levels … by means of a state transition process.” ) determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work… Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets PNG media_image5.png 18 185 media_image5.png Greyscale .”, the trellis state corresponds to the state associated with the current neural network parameter when combined with Han.) and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter. (Kasner, page 1678, col. 1, paragraph 1, “Equation (1) states that if we take the codeword associated with index i ∈ S₀, and that corresponding to –i ∈ S₁, the codewords will be the negative of one another. Equation (2) states that the probability of the codeword with index i ∈ S₀ equals the probability of the codeword with index –i ∈ S₁. These relationships allow the use of a single variable-rate code for both supersets [8]. The UTCQ quantizer returns the S₀ indices and the negative of the S₁ indices, allowing one probability model to be used for entropy coding. The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly.”, Kasner teaches that after decoding each index, the decoder “keeps track of the current state” by noting whether the index came from superset S₀ (positive codeword) or S₁ (negative codeword) and then applies that state when interpreting the next index. When combined with Han’s sequential decoding of neural-network weight indices, each decoded weight index becomes the “quantization index” that drives Kasner’s sign-based state update. Thus, for every neural-network parameter in Han’s stream, Kasner’s rule updates the internal trellis state based on the immediately preceding decoded index.) Claim 18: Han, Kasner, and Coban teaches the limitations of claim 16. Coban further teaches: Apparatus of claim 16, configured to update the state for the subsequent neural network parameter using a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter. (Coban, col. 13, line 1, “In particular , when using two scalar quantizers , a first quantizer Q0 may map transform coefficient levels ( numbers below the points , e.g. , absolute values ) to even integer multiples of quantization step size Δ. The second quantizer Q1 may map the transform coefficient levels to odd integer multiples of the quantization step size Δ or to zero .”, Coban drives state transitions by the parity (even/odd) of the previous decoded coefficient level. Because Coban explicitly organizes reconstruction levels into even multiples (Q0) and odd multiples (Q1) of Δ, the quantizer’s behavior is governed by the parity (even/odd) of the relevant integer index or level. Whenever the scheme chooses between Q0 and Q1 based on whether a level (or its integer multiple index) is even or odd, it is effectively using the parity of the quantization index as the update variable for the state machine. That even-versus-odd test is what is being interpreted as “using a parity of the quantization index … to update the state for the subsequent neural network parameter.” It is interpreted that this form of Trellis-Coded Quantization would then be applied to the neural network parameters of Han.) Claim 19: Han, Kasner, and Coban teaches the limitations of claim 1, Kasner further teaches: Apparatus of claim 16, wherein the state transition process is configured to transition between four or eight possible states. (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work.”, Kasner showcases an 8 trellis state model Kasner, page 1678, col. 1, paragraph 1, “If such a codeword is required, the trellis must switch to state four by choosing a nonzero codeword from D2. If the following source samples require a string of zero reconstruction levels, the trellis must work its way back to state zero.”, Trellis Coded Quantization (TCQ) must switch (transition) between states as an inherent process.) Claim 20: Han, Kasner, and Coban teaches the limitations of claim 16. Coban further teaches: Apparatus of claim 16, configured to transition, in the state transition process, between an even number of possible states and the number of reconstruction level sets of the plurality of reconstruction level sets is two, wherein the determining, for the current neural network parameter, the set of quantization levels out of the plurality of reconstruction level sets depending on the state associated with the current neural network parameter determines a first reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a first half of the even number of possible states, and a second reconstruction level set out of the plurality of reconstruction level sets if the state belongs to a second half of the even number of possible states. (Coban, col. 13, line 17, “In this example , coefficients in state 0 and state 1 use the Q0 ( even integer multiples of stepsize ) quantizer , and coefficients in states 2 and 3 use the Q1 ( odd quantizer , and coefficients in states 2 and 3 use the Q1 ( odd sets can be achieved by changing the parity of the level of integer multiples of stepsize ) quantizer.”, Coban’s 4-state example (states 0–3) shows that states 0 and 1 use Q0 (even-multiple quantizer) while states 2 and 3 use Q1 (odd-multiple quantizer). Here, the total number of states (4) is even; the first two states (0,1) constitute the “first half” of the states and are associated with the first reconstruction-level set (Q0), while the remaining two states (2,3) constitute the “second half” and are associated with the second reconstruction-level set (Q1). Thus, Coban’s mapping of Q0 to states in the first half and Q1 to states in the second half of an even-sized state set is what is being interpreted as the claimed “if the state belongs to a first half … use the first reconstruction level set … if the state belongs to a second half … use the second reconstruction level set.”) Claim 21: Han, Kasner, and Coban teaches the limitations of claim 16. Coban further teaches: Apparatus of claim 16, configured to perform the update of the state by means of a transition table which maps a combination of the state and a parity of the quantization index decoded from the data stream for the immediately preceding neural network parameter onto a further state associated with the subsequent neural network parameter. (Coban, col 13, line 14, “FIG . 4 is a conceptual diagram illustrating an example state transition scheme for two scalar quantizers used to perform quantization . In this example , coefficients in state 0 and state 1 use the Q0 ( even integer multiples of stepsize quantizer , and coefficients in states 2 and 3 use the Q1 ( odd integer multiples of stepsize ) quantizer .”, figure 4 of Coban clearly discloses a state transition scheme which maps the combination of the state and parity of the quantization index and shows transitions to further states. It is interpreted by the examiner that the neural network quantization of Han and its data stream would similarly apply here and provide a data stream of neural network related parameters to perform the trellis-coded quantization of Coban.) Claim 41: Han, Kasner, and Coban teaches the limitations of claim 1. Han further teaches: Apparatus of claim 1, configured to decode the quantization indices for the neural network parameters and perform the dequantization of the neural network parameters along a common sequential order among the neural network parameters. (Han, abstract, “Our method first prunes the network by learning only the important connections. Next, we quantize the weights to enforce weight sharing, finally, we apply Huffman coding.”, Han presents a decoding in a strict pipeline of pruning, retraining, quantizing and Huffman coding. Huffman coding is interpreted at the decoding process for each quantized indices of the neural network parameters and inherently does so in a sequential manner.) Claim 44: Han, Kasner, and Coban teaches the limitations of claim 1. Kasner further teaches: Apparatus of claim 1, wherein the neural network parameters relate to one reconstruction layer of reconstruction layers using which the neural network is represented, and the apparatus is in configured to reconstruct the neural network by combining the neural network parameters, neural network parameter wise, with corresponding neural network parameters of one or mor further reconstruction layers. (Kasner, page 1679, col. 2, section 4, paragraph 3, “Given a quantization index, the reconstruction level (for uniform codewords) may be computed.”, Kasner teaches the generation of reconstructed values based on decoded quantization indices. Kasner, page 1678, col. 2, paragraph 2, “During dequantization, two types of reconstruction levels are employed, uniform and trained.”, during dequantization (the reconstruction of the neural network) the reconstruction sets are used to determine each dequantized parameter value. When combined with Han these would be applied to neural network parameters values.) Claim 46: Han teaches: Apparatus for encoding neural network parameters, which define a neural network, into a data stream, configured to sequentially encode the neural network parameters (Han, abstract, “Our method first prunes the network by learning only the important connections. Next, we quantize the weights to enforce weight sharing, finally, we apply Huffman coding.”, Han explicitly applies quantized weights to neural network parameters, thereby encoding them to a compressed stream.) Kasner, in the same field of quantization methods, teaches the following limitations which the above fails to teach: quantizing the current neural network parameter onto the one reconstruction level of the selected set of reconstruction levels, (Kasner, page 1677, col. 2, paragraph 2, “During quantization, the Viterbi algorithm [6] is used to pick the sequence of codewords allowed by the trellis structure that minimizes the cumulative MSE between the input data and output reconstruction.”, Kasner’s encoder actually quantizes each parameter value by finding the closest codeword in the active superset via Viterbi. In combination with Han’s weight‐sharing (where each shared weight is treated as a codebook entry), this teaches quantizing a neural-network parameter onto its selected reconstruction level. The superset’s codewords are being interpreted as the “reconstruction levels,” and the index within that superset is the claimed quantization index that uniquely points to one codeword (one reconstruction level) out of the set associated with the current state.) and encoding a quantization index for the current neural network parameter that indicates the one reconstruction level onto which the quantization index for the current neural network parameter is quantized into the data stream. (Kasner, page 1678, col. 2, paragraph 1, “The UTCQ quantizer returns the S0 indices and the negative of the S1 indices, allowing one probability model to be used for entropy coding. The decoder may uniquely recover the index stream by simply keeping track of the current state, and negating codewords accordingly”, Kasner describes how each quantization index—which inherently indicates a specific reconstruction level in its superset—is entropy-coded into the output stream (here, using sign-shifted indices). By combining with Han’s teaching that these indices represent neural network weights, one of ordinary skill would recognize that Kasner’s index bit-stream serves as the encoded quantization index for each neural network parameter.) The motivation to combine Han with Kasner is substantially similar to that applied for claim 1 above. Coban, in the same field of trellis-coded quantization and entropy coding, further teaches the following corresponding amended limitations which the above fails to teach: selecting, for a current neural network parameter, a set of reconstruction levels out of a plurality of reconstruction level sets by means of a state transition process (Coban, col. 13, lines 1-31, “In particular, when using two scalar quantizers, a first quantizer Q0 may map transform coefficient levels (numbers below the points, e.g., absolute values) to even integer multiples of quantization step size Δ. The second quantizer Q1 may map the transform coefficient levels to odd integer multiples of the quantization step size Δ or to zero. Transform coefficients may be quantized in coding order, and video encoder 200 and/or video decoder 300 may select the quantizer based on a state machine that is driven by its previous state and the parity of the level (e.g., absolute value) of the previously coded coefficient. The state machine may be represented by a data structure. An example of the state machine is described in FIG. 4. FIG. 4 is a conceptual diagram illustrating an example state transition scheme for two scalar quantizers used to perform quantization…. As shown, the data structure of FIG. 4 transitions from state 0 to state 0, state 1 to state 2, state 2 to state 1, and state 3 to state 3 when a parity equals 0. In this example, the data structure of FIG. 4 transitions from state 0 to state 2, state 1 to state 0, state 2 to state 3, and state 3 to state 1 when the parity equals 1. As shown, the data structure of FIG. 4 may include a state machine.” Coban teaches selecting between two quantizers Q0 and Q1, which provide respective sets of reconstruction levels, by means of a state-transition process implemented by the disclosed state machine. When Coban’s state-transition quantizer selection is applied to sequentially processed neural-network parameters, Coban teaches selecting a reconstruction-level set for a current neural-network parameter by means of the state-transition process.) by determining, for the current neural network parameter, the set of reconstruction levels out of the plurality of reconstruction level sets depending on a state associated with the current neural network parameter, (Coban, col. 13, lines 14-24, “FIG. 4 is a conceptual diagram illustrating an example state transition scheme for two scalar quantizers used to perform quantization. In this example, coefficients in state 0 and state 1 use the Q0 (even integer multiples of stepsize) quantizer, and coefficients in states 2 and 3 use the Q1 (odd integer multiples of stepsize) quantizer. That is, video encoder 200 and/or video decoder 300 may determine a quantizer Q0 when the second state comprises state 0 or state 1. In some examples, video encoder 200 and/or video decoder 300 may determine a quantizer Q1 when the second state comprises state 2 or state 3.” Coban’s Q0 includes reconstruction levels corresponding to even integer multiples of the quantization step size, while Q1 includes reconstruction levels corresponding to odd integer multiples of the quantization step size or zero. Coban therefore teaches that states 0 and 1 determine the Q0 reconstruction-level set and states 2 and 3 determine the Q1 reconstruction-level set. Thus, when applied to Han’s neural-network parameters, the state associated with the current neural-network parameter determines which reconstruction-level set is used for that current neural-network parameter.) and updating the state for a subsequent neural network parameter depending on the quantization index decoded from the data stream for the immediately preceding neural network parameter, (Coban, col. 13, lines 7-11 and 25-31, “Transform coefficients may be quantized in coding order, and video encoder 200 and/or video decoder 300 may select the quantizer based on a state machine that is driven by its previous state and the parity of the level (e.g., absolute value) of the previously coded coefficient…. As shown, the data structure of FIG. 4 transitions from state 0 to state 0, state 1 to state 2, state 2 to state 1, and state 3 to state 3 when a parity equals 0. In this example, the data structure of FIG. 4 transitions from state 0 to state 2, state 1 to state 0, state 2 to state 3, and state 3 to state 1 when the parity equals 1.”; Coban, col. 33, lines 1-18, “Video encoder 200 and/or video decoder 300 update the data structure to a second state according to the first state and a parity of a partial set of syntax elements representing a partial set of a plurality of coefficient levels for the previous transform coefficient…. In this example, video encoder 200 and/or video decoder 300 may determine the second state using, with the parity as an input, the state transition scheme of FIG. 4.” Coban teaches processing coefficients in coding order and updating the state used for the current or subsequent coefficient based upon the previous state and parity information obtained from the previously coded or decoded coefficient level. The decoded coefficient level is represented by its corresponding quantization index, and its parity is therefore information obtained from that quantization index. Accordingly, when Coban’s sequential state update is applied to Han’s sequentially decoded neural-network parameters, Coban teaches updating the state for the subsequent neural-network parameter depending on the quantization index decoded for the immediately preceding neural-network parameter.) encoding the quantization index for the current neural network parameter into the data stream using arithmetic coding using a probability model which is selected depending on the state associated with the current neural network parameter. (Coban, col. 13, line 32 through col. 14, line 23, “In this scheme, coefficients may be quantized in coefficient coding order. In the JVET-J0014 scheme, the entire coefficient level must be coded before moving to the coding of the next coefficient. The reason for this is that the significance (greater than zero or ‘gt0’) map (the absolute value of coefficient level greater than 0) and greater than 1 (gt1) map context used in context based arithmetic coding depends on the state that is driven by the previous state and the parity of the coefficient level…. In the TCQ scheme of this disclosure, video encoder 200 and video decoder 300 may use separate context models for significance and greater than one coefficient contexts based on the current state (even or odd reconstruction level quantizer) used. Basically, video encoder 200 and video decoder 300 may maintain a separate context model for even and odd quantizers. In one example, if the current state is 0 or 1, then context set 0 is used, whereas if the current state is 2 or 3, then context set 1 is used. Switching between the context sets can be achieved by changing the parity of the level of the previous coefficients that minimizes the RD cost. Switching between the context sets may also change the quantizer, hence the reconstruction levels.” Coban expressly teaches the encoding side of the disclosed process using video encoder 200 and teaches that the context model used in context-based arithmetic coding depends on the current state. Coban’s context model constitutes the probability model used by the arithmetic coder. Because context set 0 is selected for states 0 and 1 and context set 1 is selected for states 2 and 3, the arithmetic-coding probability model is selected according to the same current state that selects the reconstruction-level quantizer Q0 or Q1. Accordingly, Coban teaches encoding the current quantization index using an arithmetic-coding probability model selected depending on the state associated with the current parameter.) The motivation to combine Han and Kasner with Coban is substantially similar to that applied for claim 1 above. Claims 106 and 112 recite substantially similar limitations to claim 1 and as such a similar analysis applies. Claims 107, 110, and 113 recite substantially similar limitations to claim 46 and as such a similar analysis applies. Claims 24-25, 28-29, 34-37, 39-40, 42-43, and 45 are rejected under 35 U.S.C. 103 as being unpatentable over Han in view of Kasner in further view of Coban and Sze et al. (US 2013/0272389 A1), hereafter referred to as Sze. Claim 24: Han, Kasner, and Coban teaches the limitations of claim 23. Sze, in the same field of machine learning further teaches the following which the above prior art fails to teach: Apparatus of claim 23, configured to decode the quantization index for the current neural network parameter from the data stream using binary arithmetic coding by using the probability model which depends on the state for the current neural network parameter for at least one bin of a binarization of the quantization index. (Sze, paragraph 6, “In CABAC, bins can be either context coded or bypass coded. Bypass coded bins do not require context selection which allows these bins to be processed at a much high throughput than context coded bins.”, CABAC’s binary arithmetic coding of “bins” (the binarized bits of the index) using a state-dependent context model for at least one bin directly maps to the claims bin-wise arithmetic decoding.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Han, Coban, and Kasner by incorporating the teachings of Sze (i.e. Context-Aware Binary Arithmetic Coding methods). Kasner further recognizes that different sets of indices (corresponding to different supersets) may be entropy coded with different coding models (e.g., different arithmetic models or different Huffman tables), supporting selection of the probability model based on the state/superset associated with the current parameter. A motivation of which is to provide an entropy coding technique (arithmetic coding with context/probability models) that would further reduce the overall bitstream size for the quantization indices.(Sze, paragraph 6, “CABAC is an inherently lossless compression technique notable for providing considerably better compression than most other encoding algorithms used in video encoding at the cost of increased complexity.”, by adopting CABAC for entropy-coding Han’s quantization indices, a POSITA would directly reduce the overall bitstream size of a neural networks parameters.) Claim 25: Han, Kasner, and Coban teaches the limitations of claim 23. Sze, in the same field of machine learning further teaches the following which the above prior art fails to teach: Apparatus of claim 23, wherein the at least one bin comprises a significance bin indicative of the quantization index of the current neural network parameter being equal to zero or not. (Sze, paragraph 6, “The theory and operation of CABAC coding for H.264/AVC is defined in the International Telecommunication Union, Telecommunication Standardization Sector (ITU-T) standard “Advanced video coding for generic audiovisual services” H.264, revision March 2005 or later, which is incorporated by reference herein. General principles are explained in “Context-Based Adaptive Binary Arithmetic Coding in the H.264/AVC Video Compression Standard,” Detlev Marpe, July 2003, which is incorporated by reference herein.”, Sze incorporates the H.264/AVC CABAC scheme by reference, where for transform coefficients a significant_flag (or significant_coeff_flag) is the first bin in the binarization signaling whether the coefficient is zero or non-zero. In that context, the bin representing significant_flag is the claimed “significance bin indicative of the quantization index … being equal to zero or not”—a bin value of 0 indicates a zero coefficient, and a bin value of 1 indicates a non-zero coefficient. The text in Sze pointing to CABAC’s standard operation is thus being interpreted as importing this significance-bin behavior.) Claim 28: Han, Kasner, and Coban teaches the limitations of claim 23. Sze, in the same field of machine learning further teaches the following which the above prior art fails to teach: Apparatus of claim 22, configured so that the dependency of the probability model involves a selection of a context out of a set of contexts for the neural network parameters using the dependency, each context having a predetermined probability model associated therewith. (Sze, paragraph 6, “ In brief, CABAC has multiple probability modes for different contexts. It first converts all non-binary symbols to binary symbols referred to as bins. Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate.”, CABAC’s explicit selection of one context (probability model) out of a set for each bin is analogous to the claim’s context-selection mechanic.) Claim 29: Han, Kasner, and Coban teaches the limitations of claim 23. Sze, in the same field of machine learning further teaches the following which the above prior art fails to teach: Apparatus of claim 28, configured to update the predetermined probability model associated with each of the contexts based on the quantization index arithmetically coded using the respective context. (Sze, paragraph 30, “The CABAC decoding process is the inverse of the encoding process and has similar feedback loops. Referring now to FIG. 1B, a CABAC decoder includes a bin decoder 112, a context modeler 110, and a de-binarizer 114. The context modeler 110 selects a context model for the next context bin to be decoded. As in the encoder, the context models are updated throughout the decoding process to track the probability estimations. That is, a bin is decoded based on the current state of the context model selected by the context modeler 110, and the context model is then updated to reflect the state transition and the MPS after the bin is decoded.”, after each bin (i.e. a portion of the quantization index) is arithmetically decoded under a chosen context model (referred to as the ‘probability model’ in paragraph 6 previously), the decoder updates that context model based on the decoded symbol (the MPS, Most Probable Symbol), which is the bin value.) Claim 34: Han, Kasner, and Coban teaches the limitations of claim 23. Sze, in the same field of machine learning further teaches the following which the above prior art fails to teach: Apparatus of claim 22, wherein the probability model additionally depends on the quantization index of previously decoded neural network parameters. (Sze, paragraph 59, “The intra-prediction estimation component 424 (IPE) performs intra-prediction estimation in which tests on CUs in an LCU based on multiple intra-prediction modes, prediction unit sizes, and transform unit sizes are performed using reconstructed data from previously encoded neighboring CUs stored in a buffer (not shown) to choose the best CU partitioning, prediction unit/transform unit partitioning, and intra-prediction modes based on coding cost, e.g., a rate distortion coding cost.”, in the Sze patent the context modeler (the probability model) relies on outputs of the intra-prediction estimation component. The intra-prediction estimation component supplies the previously decoded data that the context modeler uses as its context inputs. Thus, previously decoded neural network parameters (when combined with Han) is what the probability model (the context model of Sze) depends from.) Claim 35: Kasner, Han, Coban, and Sze teaches the limitations of claim 34. Sze further teaches: Apparatus of claim 34, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, a subset of probability models out of a plurality of probability models (Sze, paragraph 36, “Referring now to the CABAC encoder of FIG. 2A, the binarizer 200 converts syntax elements into strings of one or more binary symbols. The binarizer 200 directs each bin to either the context coding 206 or the bypass coding 208 of the bin encoder 204 based on a bin type determined by the context modeler 202.”, Sze, paragraph 6, “Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate.”, In Sze’s CABAC encoder (Fig. 2A; paragraph, 36), the binarizer 200 “directs each bin to either the context coding 206 or the bypass coding 208 of the bin encoder 204 based on a bin type determined by the context modeler 202.” Sze also explains that, “for each bin, the coder selects which probability model to use” (Paragraph, 6). Under a broadest reasonable interpretation, the “plurality of probability models” are the various context probability models used on the context-coded branch together with the (effectively fixed) model used on the bypass branch. When the binarizer/context modeler decides whether a bin is sent to the context coding 206 path or the bypass 208 path, that decision preselects which subset of the available probability models will be used for that bin (the context-coded subset vs. the bypass subset), corresponding to “preselect[ing] … a subset of probability models out of a plurality of probability models.”) and select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters. (Sze, paragraph 6, “Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate”, the selection of probability model (the context model as describe further in the patent) is based on the values of nearby elements. By combining with Han’s neural network quantization a POSITA would have applied the nearby elements paradigm to its neural network parameters.) Claim 36: Han, Kasner, Coban, and Sze teaches the limitations of claim 35. Kasner further teaches: Apparatus of claim 35, configured to preselect, depending on the state or the set of reconstruction levels selected for the current neural network parameter, the subset of probability models out of the plurality of probability models in a manner so that a subset preselected for a first state or reconstruction levels set is disjoint to a subset preselected for any other state or reconstruction levels set. (Kasner, page 1677, col. 2, paragraph 2, “Fig. 1 shows the eight state trellis used in this work… Fig. 1 shows that at any given trellis state, the next codeword must come from one of two supersets PNG media_image5.png 18 185 media_image5.png Greyscale . Given an initial state, a sequence of indices specifying which codeword was chosen from the appropriate superset at each step, is sufficient to allow the decoder to reproduce the sequence of codewords chosen [3].”, As discussed for claim 18 above, Sze teaches having a plurality of probability models (context models and bypass mode) and preselecting a subset of those probability models for a given bin (e.g., by directing the bin to the context-coded branch or to the bypass branch of the CABAC engine). Kasner, in turn, teaches that at any given trellis state “the next codeword must come from one of two supersets” S₀ or S₁ (the reconstruction level sets from which to select), and that a sequence of indices identifying which codeword from the appropriate superset was chosen allows the decoder to reproduce the sequence of codewords (Kasner, p.1677, col.2). Thus, Kasner’s trellis state determines which of two disjoint reconstruction-level supersets (S₀ vs. S₁) is active at each step. It is interpreted by the examiner that the choice of S₀ vs. S₁ serves as an additional input to Sze’s context modeler so that, when S₀ (a first reconstruction level set selected) is active, one subset of Sze’s probability models is preselected, and when S₁ is active (a second reconstruction level set selected), a different, non-overlapping subset is preselected. In this combined system, the subset of probability models preselected for a first state / reconstruction-level set (S₀) is disjoint from the subset preselected for another state / reconstruction-level set (S₁), as required by the claim.) Claim 37: Kasner, Han, Coban, and Sze teaches the limitations of claim 35. Sze further teaches: Apparatus of claim 35, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to. (Sze, paragraph 6, “Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate”, Sze (via CABAC) states that, for each bin, the probability model (context) is selected using information from nearby elements—for transform coefficients, this includes previously decoded significant_flags, levels, and positions in the same block. When this is applied to neural-network parameters (via Han/Wang), the “nearby elements” correspond to previously decoded quantized parameters in a neighboring portion of the network (e.g., adjacent weights or nodes). Thus, CABAC’s rule of choosing a context based on the already decoded neighboring coefficients is what is being interpreted as “selecting the probability model … depending on the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring the portion of the current parameter.” ) Claim 39: Han, Kasner, Coban, and Sze teaches the limitations of claim 37. Han further teaches: Apparatus of claim 37, configured to locate the previously decoded neural network parameters so that the previously decoded neural network parameters relate to the same neural network layer as the current neural network parameter. (Han, page 4, section 3.1, “We use k-means clustering to identify the shared weights for each layer of a trained network, so that all the weights that fall into the same cluster will share the same weight.”, Han processes weights per layer, applying quantization and clustering independently by layer. It is interpreted by the examiner that Han’s layer by layer processing by locating previous parameters for a layer is combined with Sze’s contextual modeling and decoding process to fully teach upon this claim.) Claim 40: Han, Kasner, Coban, and Sze teaches the limitations of claim 37. Han further teaches: Apparatus of claim 37, configured to locate one or more of the previously decoded neural network parameters in a manner so that the one or more previously decoded neural network parameters relate to neuron interconnections which emerge from, or lead towards, a neuron to which a neuron interconnection relates which the current neural network parameter refers to, or a further neuron neighboring said neuron. (Han, page 3, section 3, paragraph 2, “Weight sharing is illustrated in Figure 3. Suppose we have a layer that has 4 input neurons and 4 output neurons, the weight is a 4 × 4 matrix. On the top left is the 4 × 4 weight matrix, and on the bottom left is the 4 × 4 gradient matrix. The weights are quantized to 4 bins (denoted with 4 colors), all the weights in the same bin share the same value, thus for each weight, we then need to store only a small index into a table of shared weights. During update, all the gradients are grouped by the color and summed together, multiplied by the learning rate and subtracted from the shared centroids from last iteration. For pruned AlexNet, we are able to quantize to 8-bits (256 shared weights) for each CONV layers, and 5-bits (32 shared weights) for each FC layer without any loss of accuracy.”, Han processes neural connections at the level of individual neurons in fully connected layers. Neural network parameters are tied to specific neuron interconnections.) Claim 42: Kasner, Han, Coban, and Sze teaches the limitations of claim 1 Sze further teaches: Apparatus of claim 1, configured to decode the quantization index for the current neural network parameter from the data stream using binary arithmetic coding by using the probability model which depends on previously decoded neural network parameters for one or more leading bins of a binarization of the quantization index and by using an equi-probable bypass mode suffix bins of the binarization of the quantization index which follow the one or more leading bins. (Sze, paragraph 36, “Referring now to the CABAC encoder of FIG. 2A, the binarizer 200 converts syntax elements into strings of one or more binary symbols. The binarizer 200 directs each bin to either the context coding 206 or the bypass coding 208 of the bin encoder 204 based on a bin type determined by the context modeler 202. The binarizer also provides a bin index (binIdx) for each bin to the context modeler 202.”, Sze teaches using context models for some bins and using bypass coding (equi-probable) for others. A POSITA would have applied this to Han’s quantization indices by binarizing them and then decoding with Sze’s adaptive Context Aware Binary Arithmetic Coding instead of the Huffman coding applied in Han. This uses probability models (based on prior indices) for leading bins, and bypass mode (equi-probable) for suffix bins.) Claim 43: Kasner, Han, Coban, and Sze teaches the limitations of claim 42. Sze further teaches: Apparatus of claim 42, wherein the suffix bins of the binarization of the quantization index represent bins of a binarization code of a suffix binarization for binarizing values of the quantization index an absolute value of which exceeds a maximum absolute value representable by the one or more leading bins, wherein the apparatus is configured to selected the suffix binarization depending on the quantization index of previously decoded neural network parameters. (Sze, paragraph 84, “Referring again to FIG. 6, in this method, the variable i is a bin counter and the variable N is the absolute value of a delta quantization parameter (delta qp) syntax element. Initially, the value of the bin counter i is set to 0. In this method, for values of N greater than or equal to cMax, cMax bins with a value of 1 are context coded into the compressed bit stream followed by bypass coded bins corresponding to the EGk codeword for N-cMax.”, Sze teaches a binarization scheme where absolute value of a delta quantization parameter is first encoded using a fixed number of leading context-coded bins. If the value exceeds that representable range, additional bypass-coded suffix bins are used to encode the excess portion using an Ex-Golomb code. Thus, the suffix binarization is selected based on whether the quantization index exceeds a threshold.) Claim 45: Kasner, Han, Coban, and Sze teaches the limitations of claim 44. Sze further teaches: Apparatus of claim 44, configured to decode the quantization index for the current neural network parameter from the data stream using arithmetic coding using a probability model which depends on corresponding neural network parameter corresponding to the current neural network parameter. (Sze, paragraph 6, “Then, for each bin, the coder selects which probability model to use, and uses information from nearby elements to optimize the probability estimate. Arithmetic coding is then applied to compress the data.”, Sze, paragraph 41, “The bins generated by the context coding 224 and bypass coding 222 are provided the multiplexer 226. The multiplexor 226 selects the output of the context coding 224 or the bypass coding 222 to be provided to the de-binarizer 230 according to the bin type provided by the context modeler 228. The de-binarizer 230 receives decoded bins for a syntax element from the bin decoder 220 and operates to reverse the binarization of the encoder to reconstruct the syntax elements.”, the decoding process in Sze includes the de-binarizer, which is fed by bins generated using probability models chosen by the context modeler. Since the selection depends on the value being decoded it is dependent on the corresponding parameter, and when combined with Han would be adapted to neural network parameters.) Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Kasner in view of Han and in further view of Coban, and Schwarz et al., (Schwarz, H., Nguyen, T., Marpe, D., & Wiegand, T. (2019, March). Hybrid video coding with trellis-coded quantization. In 2019 Data Compression Conference (DCC) (pp. 182-191). IEEE.), hereafter referred to as Schwarz. Claim 8: Han, Kasner, and Coban teaches the limitations of claim 1. Coban further teaches Apparatus of claim 7, wherein the number of reconstruction level sets of the plurality of reconstruction level sets is two and the apparatus is configured to derive the intermediate value for each neural network parameter by, if the selected reconstruction level set for the respective neural network parameter is a first set, multiply the quantization index for the respective neural network parameter by two to acquire the intermediate value for the respective neural network parameter; (Coban, col. 13, line 1, “In particular , when using two scalar quantizers , a first quantizer Q0 may map transform coefficient levels ( numbers below the points , e.g. , absolute values ) to even integer multiples of quantization step size Δ.”, mapping each index to an even integer multiple of Δ is exactly the same as computing an intermediate value of the index multiplied by 2.) and if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is equal to zero, set the intermediate value for the respective neural network parameter equal to zero; (Coban, col. 13, line 4, “The second quantizer Q1 may map the transform coefficient levels to odd integer multiples of the quantization step size Δ or to zero .”, the case when the decoded index is 0 is covered by Q1 outputting 0 as the intermediate value.) Schwarz, in the same field of trellis coded quantization, teaches the following limitations which Kasner, Han, and Coban fail to teach: and if the selected reconstruction level set for a respective neural network parameter is a second set and the quantization index for the respective neural network parameter is greater than zero, multiply the quantization index for the respective neural network parameter by two and subtract one from the result of the multiplication to acquire the intermediate value for the respective neural network parameter; (Schwarz, page 184, paragraph 3, “The structure of the two scalar quantizers Q0 and Q1 used in our approach is shown in Fig. 1. The reconstruction values for Q0 are given by the even multiples of the quantization step size Δ; the reconstruction values for Q1 are given by the odd multiples of Δ and, in addition, the value of zero… For both quantizers, the reconstruction values are indicated by integer quantization indexes q, where the quantization index equal to zero corresponds to the reconstruction value equal to zero.”, Page 185, paragraph 1, “Given the N quantization indexes qk of a block, with k indicating the coding order, the associated reconstructed transform coefficients t k can be obtained by the following simple algorithm: PNG media_image7.png 90 254 media_image7.png Greyscale ”, In Schwarz, the TCQ design defines two scalar quantizers Q0 and Q1, where Q0 uses even multiples of Δ and Q1 uses odd multiples of Δ (plus zero), with each reconstruction value associated with an integer quantization index q. For each coefficient, the decoder computes the reconstructed value as t_k = (2·q_k − (s_k >> 1)·sgn(q_k))·Δ. For states corresponding to the second quantizer (Q1), (s_k >> 1) = 1, so the formula specializes to t_k = (2·q_k − sgn(q_k))·Δ. For a positive quantization index q_k > 0, sgn(q_k) = 1, and thus t_k = (2·q_k − 1)·Δ. Under the claim’s terminology, the quantization index is q_k, the “intermediate value” is the integer (2·q_k − 1), and multiplying that intermediate value by the predetermined quantization step size Δ yields the odd reconstruction level. This directly corresponds to “multiply the quantization index … by two and subtract one … to acquire the intermediate value” when the selected reconstruction-level set is the second set and the quantization index is greater than zero.) and if the selected reconstruction level set for a current neural network parameter is a second set and the quantization index for the respective neural network parameter is less than zero, multiply the quantization index for the respective neural network parameter by two and add one to the result of the multiplication to acquire the intermediate value for the respective neural network parameter. (Schwarz, page 184, paragraph 3, “The structure of the two scalar quantizers Q0 and Q1 used in our approach is shown in Fig. 1. The reconstruction values for Q0 are given by the even multiples of the quantization step size Δ; the reconstruction values for Q1 are given by the odd multiples of Δ and, in addition, the value of zero… For both quantizers, the reconstruction values are indicated by integer quantization indexes q, where the quantization index equal to zero corresponds to the reconstruction value equal to zero.”, Page 185, paragraph 1, “Given the N quantization indexes qk of a block, with k indicating the coding order, the associated reconstructed transform coefficients t k can be obtained by the following simple algorithm: PNG media_image7.png 90 254 media_image7.png Greyscale ”, As above, Schwarzs’ TCQ scheme uses two quantizers Q0 and Q1; Q1’s reconstruction values are odd integer multiples of Δ plus zero, with each value addressed by an integer quantization index q. The decoder reconstructs each coefficient via t_k = (2·q_k − (s_k >> 1)·sgn(q_k))·Δ. For the second quantizer Q1, (s_k >> 1) = 1, giving t_k = (2·q_k − sgn(q_k))·Δ. When the quantization index is negative (q_k < 0), sgn(q_k) = −1, so the expression becomes t_k = (2·q_k + 1)·Δ. The factor (2·q_k + 1) is a negative odd integer (… −5, −3, −1, …), which, when multiplied by Δ, yields the negative odd-multiple reconstruction levels required for Q1. Under the claim’s terminology, the quantization index is q_k, the “intermediate value” is the integer (2·q_k + 1), and multiplying that intermediate value by the predetermined quantization step size Δ produces the final reconstruction level. This matches “multiply the quantization index … by two and add one … to acquire the intermediate value” for the second reconstruction-level set when the quantization index is less than zero.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to incorporate the trellis-coded quantization scheme of Schwarz into the existing Kasner, Han and Coban combination in order to improve coding efficiency and rate–distortion performance when quantizing and reconstructing coefficients. Schwarz describes replacing conventional scalar quantization by trellis-coded quantization and reports that “the coding efficiency of transform coding can be improved by replacing scalar quantization with trellis-coded quantization (TCQ) and using advanced entropy coding techniques for coding the quantization indexes” (Abstract) and that their implementation in the VVC test model “yielded average bit-rate savings of 4.9% for intra-only coding and 3.3% for typical random access configurations” (Abstract) relative to a scalar-quantization baseline. A person of ordinary skill in the art implementing Han’s neural-network quantization with Kasner/Coban-style trellis coding would therefore have been motivated to adopt Schwarz explicit TCQ reconstruction rule—using intermediate values of the form 2·q_k − sgn(q_k) (i.e., 2i−1 and 2i+1 for positive and negative indices)—to obtain the same kind of bit-rate savings and improved coding efficiency for neural-network parameters. Claims 38 are rejected under 35 U.S.C. 103 as being unpatentable over Kasner in view of Han and in further view of Coban, Sze, and Marpe et al., (Marpe, D., Schwarz, H., & Wiegand, T. (2003). Context-based adaptive binary arithmetic coding in the H. 264/AVC video compression standard. IEEE Transactions on circuits and systems for video technology, 13(7), 620-636.), hereafter referred to as Marpe. Claim 38: Kasner, Han, Coban and Sze teaches the limitations of claim 35. Marpe, in the same field of Context-Aware Binary Arithmetic Coding further teaches: Apparatus of claim 35, configured to select the probability model for the current neural network parameter out of the subset of probability models depending on a characteristic of the quantization index of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, (Marpe, page 629, col. 1, paragraph 2, “Characteristic Features: For the coding of residual data within the H.264/AVC standard specifically designed syntax elements are used in CABAC entropy coding mode. These elements and their related coding scheme are characterized by the following distinct features… Context models for coding of nonzero transform coefficients are chosen based on the number of previously transmitted nonzero levels within the reverse scanning path.”, H.264’s CABAC selects a context model for each bin by examining previously coded motion-vector differences (MVD’s), which is interpreted as analogous to choosing a probability model based on previously decoded weight indices from neighboring parameters when combined with Han.) the characteristic comprising on or more of the signs of non-zero quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, (Marpe, page 629, col. 2, paragraph 2, “Significance Map: If the coded_block_flag indicates that a block has significant coefficients, a binary-valued significance map is encoded. For each coefficient in scanning order, a one-bit symbol significant_coeff_flag is transmitted.”, Marpe’s CABAC description encodes, for each coefficient, both (i) a significant_coeff_flag indicating zero vs. non-zero, and (ii) a coeff_sign_flag representing the sign of each non-zero coefficient. Contexts for these flags are conditioned on neighboring coefficients in the same block. When mapped to neural networks, each transform coefficient corresponds to a parameter; the neighboring coefficients in the block correspond to a neighboring portion of the neural network, and the coeff_sign_flag bits for those neighbors are the “signs of non-zero quantization indices of previously decoded neural network parameters.” Thus, using those sign flags in context templates is what is being interpreted as using the sign characteristics of neighboring non-zero indices to drive probability-model selection.) the number of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero (Marpe, page 625, col. 2, paragraph 1, “In contrast to all other types of context models, both types depend on the context categories of different block types, as specified below… modeling functions are specified that involve the evaluation of the accumulated number of encoded (decoded) levels with a specific value prior to the current level bin to encode (decode).”, Marpe, page 630, col. 2, last paragraph, “When a level with an absolute value greater than 1 has been encoded, i.e., when NumLgt1 is greater than 0, a context index increment of 4 is used for all remaining levels of the regarded block.”, Marpe specifies context models that depend on counters like NumLgt1(i) and NumT1(i), which are accumulated numbers of previously encoded non-zero coefficients with certain absolute values before the current bin. These counters effectively count how many prior coefficients in the same block (neighboring region) are non-zero. When mapped onto neural-network parameters, that block of coefficients is the neighboring portion of the network, and the count NumLgt1/NumT1 is what is being interpreted as the “number of quantization indices of previously decoded neural network parameters … in the neighboring portion … which are non-zero.”) a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to a difference between (Marpe, page 630, col. 2, last paragraph, “When a level with an absolute value greater than 1 has been encoded, i.e., when NumLgt1 is greater than 0, a context index increment of 4 is used for all remaining levels of the regarded block.”, Marpe defines counters such as NumLgt1 that accumulate how many previously encoded coefficients in the block have an absolute level greater than 1, and uses conditions like “when NumLgt1 is greater than 0, a context index increment of 4 is used for all remaining levels of the regarded block” (Marpe, p.630, col.2). Each coefficient with |level|>1 contributes at least 2 to the total sum of absolute values across the block, so the condition NumLgt1>0 implies that the sum of absolute values of previously encoded coefficients in that block exceeds a positive minimum. Using NumLgt1 as a context-selection variable is therefore being interpreted as using a characteristic derived from the sum of absolute values of previously decoded neighboring quantization indices when selecting a probability model, corresponding to the claimed “sum of the absolute values … of previously decoded … neighboring parameters.” Using NumLgt1(i) as a context-selection condition is therefore being interpreted, under a broad reading, as using a function of the sum of absolute values of previously decoded neighboring quantization indices when selecting a probability model, corresponding to the claimed “sum of the absolute values … of previously decoded … neighboring parameters.”) a sum of the absolute values of quantization indices of previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and the number of quantization indices of the previously decoded neural network parameters which relate to a portion of the neural network neighboring a portion which the current neural network parameter relates to, and which are non-zero. (Marpe, page 630, col. 2, last paragraph, “Then, the context for the first bin of coeff_abs_level_minus1 is determined by the current value NumT1, where the following additional rules apply. If more than three past coded coefficients have an absolute value of 1, the context index increment of three is always chosen. When a level with an absolute value greater than 1 has been encoded, i.e., when NumLgt1 is greater than 0, a context index increment of 4 is used for all remaining levels of the regarded block.”, Marpe’s decision logic jointly examines NumT1 (the count of prior levels equal to 1) and NumLgt1 (the count/sum of prior levels >1) to choose between two context-index increments (3 vs. 4). In effect, it is checking whether the difference between the total magnitude (sum) and the simple count exceeds a threshold, mapping directly to use a difference between “sum of absolute values of quantization indices” and “number of quantization indices” for model selection.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Han, Kasner, Coban, and Sze by incorporating the teachings of Marpe (i.e. Context-Aware Binary Arithmetic Coding in H.264/AVC). A motivation of which is to provide a quantization technique that would further reduce the overall bitstream size. (Marpe, page 632, col. 2, paragraph 3, “The CABAC entropy coding method is part of the Main profile of H.264/AVC [1] and may find its way into video streaming, broadcast, or storage applications within this profile. Experimental results have shown the superior performance of CABAC in comparison to the baseline entropy coding method of VLC/CAVLC. For typical test sequences in broadcast applications, averaged bit-rate savings of 9% to 14% corresponding to a range of acceptable video quality of about 30–38 dB were obtained.”, Marpe’s CABAC method provides explicit superior performance in terms of amount of compression yielding acceptable quality loss.) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Reagan, B., Gupta, U., Adolf, B., Mitzenmacher, M., Rush, A., Wei, G. Y., & Brooks, D. (2018, July). Weightless: Lossy weight encoding for deep neural network compression. In International Conference on Machine Learning (pp. 4324-4333). PMLR. Oktay, D., Ballé, J., Singh, S., & Shrivastava, A. (2019). Scalable model compression by entropy penalized reparameterization. arXiv preprint arXiv:1906.06624. US20190387259A1 - Trellis coded quantization coefficient coding Wiedemann, S., Kirchhoffer, H., Matlage, S., Haase, P., Marban, A., Marinc, T., ... & Samek, W. (2019). DeepCABAC: Context-adaptive binary arithmetic coding for deep neural network compression. arXiv preprint arXiv:1905.08318. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HYUNGJUN B YI whose telephone number is (703)756-4799. The examiner can normally be reached M-F 9-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 270-0419. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /H.B.Y./Examiner, Art Unit 2146 /DANIEL T PELLETT/Primary Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

Jun 17, 2022
Application Filed
Aug 11, 2025
Non-Final Rejection mailed — §103
Nov 11, 2025
Response Filed
Mar 09, 2026
Non-Final Rejection mailed — §103
Jun 09, 2026
Response Filed
Aug 10, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12651178
MACHINE LEARNING TECHNIQUES FOR ASSOCIATING NETWORK ADDRESSES WITH INFORMATION OBJECT ACCESS LOCATIONS
5y 4m to grant Granted Jun 09, 2026
Patent 12619888
END-TO-END SYSTEMS AND METHODS FOR CONSTRUCT SCORING
1y 7m to grant Granted May 05, 2026
Patent 12536429
INTELLIGENTLY MODIFYING DIGITAL CALENDARS UTILIZING A GRAPH NEURAL NETWORK AND REINFORCEMENT LEARNING
4y 7m to grant Granted Jan 27, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
32%
Grant Probability
58%
With Interview (+25.8%)
4y 4m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 28 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month