Prosecution Insights
Last updated: August 17, 2026
Application No. 18/823,504

METHOD, APPARATUS, AND MEDIUM FOR VISUAL DATA PROCESSING

Non-Final OA §102§103
Filed
Sep 03, 2024
Priority
Mar 03, 2022 — CN PCT/CN2022/079015 +1 more
Examiner
NAH, JONGBONG
Art Unit
Tech Center
Assignee
Bytedance Inc.
OA Round
1 (Non-Final)
76%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
90 granted / 118 resolved
+16.3% vs TC avg
Strong +16% interview lift
Without
With
+16.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
24 currently pending
Career history
135
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
64.3%
+24.3% vs TC avg
§102
22.3%
-17.7% vs TC avg
§112
1.6%
-38.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 118 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 09/03/2024, 12/08/2025, and 05/05/2026 is/are compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Office Action Summary Claim(s) 1-11 and 14-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Minnen et al (US 2023/0206512 A1). Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minnen et al (US 2023/0206512 A1) in view of Tjandrasuwita et al (US 2006/0078211 A1). Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minnen et al (US 2023/0206512 A1) in view of Tjandrasuwita et al (US 2006/0078211 A1), further in view of Ikonin et al (US 2023/0262243 A1). Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-11 and 14-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Minnen et al (US 2023/0206512 A1). Regarding claim(s) 1 and 18-20, Minnen teaches an apparatus for visual data processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor (Paragraph [0087]: “a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data”), cause the processor to perform acts comprising: obtaining, for a conversion between visual data and a bitstream of the visual data with a neural network (NN)-based model (Paragraph [0029]: “The first encoder neural network 110 is configured to process the input data 105 (x) to generate a latent representation 111 (y) of the input data 105”; and Paragraph [0030]: “the compression system 100 quantizes the latent representation 111 […] to generate […] as quantized latent representation of data”), an intermediate representation of the visual data, the intermediate representation being different from a quantized latent representation of the visual data (Paragraph [0048]: “The decoded representation 311 and the latent residual prediction (LRP) 326 is further combined (e.g., summed or concatenated) to generate a first augmented slice 131 (y1)”; Paragraph [0056]: “The decoded representation 411 and the latent residual prediction (LRP) 426 is further combined (e.g., summed or concatenated) to generate a subsequent augmented slice 136 (y2)”; and Paragraph [0060]: “During decompression, the first augmented slice and the subsequent augmented slices are combined together to generate an augmented representation of the latent representation of the data that is provided as input to a decoder neural network 155”) and being generated based on at least one of the following: at least one parameter, at least a part of the quantized latent representation, a prediction of the at least a part of the quantized latent representation, or a difference between the prediction and the at least a part of the quantized latent representation (Figure 1; Figure 3; Figure 4; Paragraph [0042]: “[…] receive as input the hyperprior parameters (mu and sigma) […] and the first slice of quantized latent representation of data 118 (y1) and generate a compressed representation of the first slice 144, and a first augmented slice 131”; Paragraph [0047]: “[…] configured to process the first slice of quantized latent representation of data 118 (y1) […]”; Paragraph [0050]: “[…] receive as input […] the subsequent slice of quantized latent representation of data which in this case is the second slice 119 (y2) […]”; Paragraph [0047] – Paragraph [0048]: “[…] hyperprior parameters mu 126 (μ) to generate the latent residual prediction (LRP) 326 of the first slice of quantized latent representation of data […] The decoded representation 311 and the latent residual prediction (LRP) 326 is further combined (e.g., summed or concatenated) to generate a first augmented slice 131 (y1)”; and Paragraph [0055] – Paragraph [0056]: “[…] to generate the latent residual prediction (LRP) 426 of the second slice of quantized latent representation of data […] The decoded representation 411 and the latent residual prediction (LRP) 426 is further combined (e.g., summed or concatenated) to generate a subsequent augmented slice 136 (y2)”); and performing, for the conversion, a synthesis transform on the intermediate representation (Paragraph [0060]: “the first augmented slice and the subsequent augmented slices are combined together to generate an augmented representation of the latent representation of the data that is provided as input to a decoder neural network 155”; and Paragraph [0061]: “the decoder neural network 155 may be configured to receive as input, augmented representation of the latent representation of the data 150 and generate a decoded output 160 […]”). Regarding claim(s) 2, Minnen teaches the method of claim 1, wherein obtaining the intermediate representation comprises: obtaining the at least a part of the quantized latent representation (Paragraph [0042]: “[…] receive as input the hyperprior parameters (mu and sigma) […] and the first slice of quantized latent representation of data 118 (y1) and generate a compressed representation of the first slice 144, and a first augmented slice 131”; and Paragraph [0050]: “[…] receive as input […] the subsequent slice of quantized latent representation of data which in this case is the second slice 119 (y2) […] and […] the first augmented slice 131 (y1)”); and updating the at least a part of the quantized latent representation based on at least one of the at least one parameter, the prediction or the difference, to obtain at least a part of the intermediate representation (Paragraph [0047] – Paragraph [0048]: “[…] hyperprior parameters mu 126 (μ) to generate the latent residual prediction (LRP) 326 of the first slice of quantized latent representation of data […] The decoded representation 311 and the latent residual prediction (LRP) 326 is further combined (e.g., summed or concatenated) to generate a first augmented slice 131 (y1)”; and Paragraph [0055] – Paragraph [0056]: “[…] to generate the latent residual prediction (LRP) 426 of the second slice of quantized latent representation of data […] The decoded representation 411 and the latent residual prediction (LRP) 426 is further combined (e.g., summed or concatenated) to generate a subsequent augmented slice 136 (y2)”), Examiner’s Note: Applicant’s specification expressly identifies generating the intermediate representation as an equivalent of updating the quantized latent representation in Paragraph [0266]. The cited reference generates an augmented slice from a slice of the quantized latent representation together with hyperprior parameters and latent residual prediction. The resulting augmented slice corresponds to an updated version of the slice of the quantized latent representation and constitutes at least a part of the claimed intermediate representation. Regarding claim(s) 3, Minnen teaches the method of claim 2, wherein updating the at least a part of the quantized latent representation comprises at least one of: scaling the at least a part of the quantized latent representation with a first parameter of the at least one parameter (Wherein clause is claimed as an alternative “at least one of”); updating the at least a part of the quantized latent representation based on a product of the prediction and a second parameter of the at least one parameter and/or a product of the difference and a third parameter of the at least one parameter (Wherein clause is claimed as an alternative “or”); or updating the at least a part of the quantized latent representation by at least adding a fourth parameter of the at least one parameter (Paragraph [0040]: “the hyperprior parameters are probability distribution parameters such as mu 126 (μ) and sigma 127 (σ) representing a Gaussian distribution that parameterizes the entropy model characterizing […]”; and Paragraph [0060]: “[…] the compression system 100 adds mu 126 (μ) to the augmented representation of the latent representation 150”), or wherein the prediction is a mean (Paragraph [0040]: “the hyperprior parameters are probability distribution parameters such as mu 126 (μ) and sigma 127 (σ) representing a Gaussian distribution that parameterizes the entropy model characterizing […]”; and Paragraph [0045]: “the convolutional neural network block 320 […] is configured to receive as input, the hyperprior parameter mu 126 (μ) […]”), or the difference is comprised in a quantized residual latent representation of the visual data (Paragraph [0047]: “[…] This process leads to a residual error (r = y − Q[y]) in the latent space […] The latent residual prediction model 325 […] is configured to process the first slice of quantized latent representation of data 118 (y1) and the hyperprior parameters mu 126 (μ) to generate the latent residual prediction (LRP) […]”), or wherein obtaining the at least a part of the quantized latent representation comprises: performing an entropy decoding process on the bitstream to obtain the at least a part of the quantized latent representation (Paragraph [0025]: “The decompression system can decompress the data by recovering the conditional entropy model from the compressed data, and using the conditional entropy model to decompress (i.e., entropy decode) the compressed code symbols”; Paragraph [0048]: “the compressed representation of the first slice 144 is decoded using an entropy decoder 310 to generate a decoded representation 311 of the compressed representation of the first slice based on the hyperprior parameters mu 126 (μ) and sigma 127 (σ)”; and Paragraph [0056]: “the compressed representation of the first slice 146 is decoded using an entropy decoder 410 to generate a decoded representation 411 of the compressed representation of the first slice based on the hyperprior parameters mu 126 (μ) and sigma 127 (σ)”). Regarding claim(s) 4, Minnen teaches the method of claim 2, wherein obtaining the at least a part of the quantized latent representation comprises: obtaining the difference by performing an entropy decoding process on the bitstream (Paragraph [0048]: “the compressed representation of the first slice 144 is decoded using an entropy decoder 310 to generate a decoded representation 311 of the compressed representation of the first slice based on the hyperprior parameters mu 126 (μ) and sigma 127 (σ)”; and Paragraph [0049]: “During decompression and […] the latent residual prediction model 325 […] is configured to process the decoded representation 311 instead of the first slice of quantized latent representation of data 118 (y1) to generate the latent residual prediction (LRP) 326 […]”); generating the prediction by using a first model in the NN-based model (Figure 3; Paragraph [0043]: “The first slice processing network 130 includes (1) an entropy encoder 305, (2) an entropy decoder 310 and (3) two convolutional neural network blocks 315 and 320 that includes convolutional neural network layers, and (4) a latent residual prediction model 325”; Paragraph [0047]: “[…] The latent residual prediction model 325 […] is configured to process the first slice of quantized latent representation of data 118 (y1) and the hyperprior parameters mu 126 (μ) to generate the latent residual prediction (LRP) […]”; and Paragraph [0049]); and generating the at least a part of the quantized latent representation based on the prediction and the difference (Paragraph [0048]: “The decoded representation 311 and the latent residual prediction (LRP) 326 is further combined (e.g., summed or concatenated) to generate a first augmented slice 131 (y1)”; and Paragraph [0056]: “The decoded representation 411 and the latent residual prediction (LRP) 426 is further combined (e.g., summed or concatenated) to generate a subsequent augmented slice 136 (y2)”). Regarding claim(s) 5, Minnen teaches the method of claim 4, wherein generating the prediction comprises: generating a prediction of a sample in the at least a part of the quantized latent representation based on at least one reconstructed sample of the quantized latent representation by using the first model (Paragraph [0049]: “During decompression and […] the latent residual prediction model 325 […] is configured to process the decoded representation 311 instead of the first slice of quantized latent representation of data 118 (y1) to generate the latent residual prediction (LRP) 326 of the first slice of quantized latent representation of data”), or wherein the first model is a prediction model (Figure 3; Paragraph [0043]: “The first slice processing network 130 includes (1) an entropy encoder 305, (2) an entropy decoder 310 and (3) two convolutional neural network blocks 315 and 320 that includes convolutional neural network layers, and (4) a latent residual prediction model 325”; Paragraph [0047]: […] The latent residual prediction model 325 […] is configured to process the first slice of quantized latent representation of data 118 (y1) and the hyperprior parameters mu 126 (μ) to generate the latent residual prediction (LRP) […]”; and Paragraph [0049]), or wherein the first model is autoregressive (Figure 1; Paragraph [0047]: “The latent residual prediction model 325 within the first slice processing network 130 is configured to process the first slice of quantized latent representation of data 118 (y1) […]”; and Paragraph [0050]: “the subsequent slice processing network 135 is configured to receive […] the first augmented slice 131 (y1) and generate as output a compressed representation of the second slice 146, and a subsequent augmented slice 136”), or wherein the first model comprises a context subnetwork or a context model subnetwork. Regarding claim(s) 6, Minnen teaches the method of claim 2, wherein the at least a part of the quantized latent representation comprises all samples of the quantized latent representation, and the intermediate representation corresponds to a result of the updating (Paragraph [0031]: “The ordered collection of code symbols 116 […] is processed into a plurality of slices of quantized latent representations of data such that each slice of quantized latent representation of data is different from the other slices of quantized latent representation of data”; Paragraph [0032]: “[…] The quantized latent representation of data 116 (ŷ) is processed into two slices: (1) a first slice of quantized latent representation of data 118 (y1), and (2) a second slice of quantized latent representation of data 119 (y2)”; and Paragraph [0060]: “During decompression, the first augmented slice and the subsequent augmented slices are combined together to generate an augmented representation of the latent representation of the data […] In this case, the first augmented slice 131 and the subsequent augmented slices 136 are combined together to generate an augmented representation of the latent representation 150”), or wherein the method further comprises: updating a further part of the quantized latent representation based on at least one further parameter different from the at least one parameter, the further part being different from the at least a part of the quantized latent representation (Paragraph [0050]: “the subsequent slice processing network 135 is configured to receive as input the hyperprior parameters (mu and sigma) […] the subsequent slice of quantized latent representation of data which in this case is the second slice 119 (y2) […] the first augmented slice 131 (y1) and generate as output a compressed representation of the second slice 146, and a subsequent augmented slice 136”; and Paragraph [0056]: “The decoded representation 411 and the latent residual prediction (LRP) 426 is further combined (e.g., summed or concatenated) to generate a subsequent augmented slice 136 (y2)”). Regarding claim(s) 7, Minnen teaches the method of claim 1, wherein obtaining the intermediate representation comprises: generating at least a part of the intermediate representation based on the prediction and the difference (Paragraph [0047]: “[…] The latent residual prediction model 325 […] is configured to process the first slice of quantized latent representation of data 118 (y1) and the hyperprior parameters mu 126 (μ) to generate the latent residual prediction (LRP) 326 of the first slice of quantized latent representation of data”; Paragraph [0048]: “The decoded representation 311 and the latent residual prediction (LRP) 326 is further combined (e.g., summed or concatenated) to generate a first augmented slice 131 (y1)”; and Paragraph [0056]: “The decoded representation 411 and the latent residual prediction (LRP) 426 is further combined (e.g., summed or concatenated) to generate a subsequent augmented slice 136 (y2)”). Regarding claim(s) 8, Minnen teaches the method of claim 7, wherein generating the at least a part of the intermediate representation comprises: generating the at least a part of the intermediate representation by adding up: a product of the prediction and a first parameter of the at least one parameter, and a product of the difference and a second parameter of the at least one parameter, or generating the at least a part of the intermediate representation by adding up: the product of the prediction and the first parameter, the product of the difference and the second parameter, and a third parameter of the at least one parameter, or wherein at least one of the prediction or the difference is generated by using a first model (Figure 3; Paragraph [0043]: “The first slice processing network 130 includes (1) an entropy encoder 305, (2) an entropy decoder 310 and (3) two convolutional neural network blocks 315 and 320 that includes convolutional neural network layers, and (4) a latent residual prediction model 325”; Paragraph [0047]: “[…] The latent residual prediction model 325 […] is configured to process the first slice of quantized latent representation of data 118 (y1) and the hyperprior parameters mu 126 (μ) to generate the latent residual prediction (LRP) […]”; and Paragraph [0049]: “During decompression […] the latent residual prediction model 325 […] is configured to process the decoded representation 311 instead of the first slice of quantized latent representation of data 118 (y1) to generate the latent residual prediction (LRP) 326 […]”). Regarding claim(s) 9, Minnen teaches the method of claim 4, wherein the first model comprises a neural network-based subnetwork, or an input of the first model comprises the bitstream, or wherein the first model comprises at least one of a first subnetwork for generating the prediction or a second subnetwork for generating a statistical value (Figure 3; Paragraph [0043]: “The first slice processing network 130 includes (1) an entropy encoder 305, (2) an entropy decoder 310 and (3) two convolutional neural network blocks 315 and 320 that includes convolutional neural network layers, and (4) a latent residual prediction model 325”; Paragraph [0044]: “the convolutional neural network block 315 […] is configured to receive as input, the hyperprior parameter sigma 127 (σ) […]”; and Paragraph [0045]: “the convolutional neural network block 320 […] is configured to receive as input, the hyperprior parameter mu 126 (μ)”). Regarding claim(s) 10, Minnen teaches the method of claim 9, wherein the first subnetwork is a hyper decoder subnetwork, the second subnetwork is a hyper scale decoder subnetwork, or the statistical value is a variance (Paragraph [0037]: “The hyperprior processing network 125 includes (1) a second quantizer 205, (2) an entropy encoder 210, (3) a predetermined entropy model 215, (4) an entropy decoder 220 and (5) two convolutional neural network blocks 225 and 230 that includes convolutional neural network layers”; Paragraph [0040]: “The decoded representations of the hyperprior parameters is further provided as input to convolutional neural network blocks to generate the hyperprior parameters […] the hyperprior parameters are probability distribution parameters such as mu 126 (μ) and sigma 127 (σ) representing a Gaussian distribution that parameterizes the entropy model characterizing one or more code symbol probability distributions”; Paragraph [0044]: “the convolutional neural network block 315 within the first slice processing network 130 is configured to receive as input, the hyperprior parameter sigma 127 (σ) that is generated by the hyperprior processing network”). Regarding claim(s) 11, Minnen teaches the method of claim 1, further comprising: determining the at least a part of the quantized latent representation from the quantized latent representation (Paragraph [0031]: “The ordered collection of code symbols 116 […] is processed into a plurality of slices of quantized latent representations of data such that each slice of quantized latent representation of data is different from the other slices of quantized latent representation of data”; and Paragraph [0032]: “[…] The quantized latent representation of data 116 (ŷ) is processed into two slices: (1) a first slice of quantized latent representation of data 118 (y1), and (2) a second slice of quantized latent representation of data 119 (y2)”). Regarding claim(s) 14, Minnen teaches the method of claim 1, wherein the at least one parameter or an indication of the at least one parameter is comprised in the bitstream, or wherein the at least one parameter is a scalar value different from zero or a vector, or wherein the at least one parameter is determined based on a quality metric, or wherein the at least a part of the quantized latent representation comprises one or more samples of the quantized latent representation (Paragraph [0031]: “The ordered collection of code symbols 116 […] is processed into a plurality of slices of quantized latent representations of data such that each slice of quantized latent representation of data is different from the other slices of quantized latent representation of data”; and Paragraph [0032]: “[…] The quantized latent representation of data 116 (ŷ) is processed into two slices: (1) a first slice of quantized latent representation of data 118 (y1), and (2) a second slice of quantized latent representation of data 119 (y2)”). Regarding claim(s) 15, Minnen teaches the method of claim 1, wherein at least one of the following is indicated in the bitstream: information on whether to apply the method, or information on how to apply the method (Wherein clause is claimed as an alternative “or”), or wherein at least one of the following is dependent on a color format and/or a color component of the visual data: information on whether to apply the method, or information on how to apply the method (Wherein clause is claimed as an alternative “or”), or wherein a value included in the bitstream is coded at one of the following: a sequence level, a picture level, a slice level, or a block level (Paragraph [0005]: “processing the quantized latent representation of data into a plurality of slices of quantized latent representations of data […] for each slice subsequent to the first slice […] generating […] a compressed representation of the respective slice […] and wherein a combination of the compressed representation of the first slice and each compressed representation of each respective slice form a compressed representation of the data”; Paragraph [0046]: “The entropy encoder 305 within the first slice processing network 130 compresses the first slice of the quantized latent representation of data 118 (y1) […]”; and Paragraph [0054]: “The entropy encoder 405 within the subsequent slice processing network 135 compresses the second slice of the quantized latent representation of data 119 (y2”), or wherein a value included in the bitstream is binarized before being coded (Wherein clause is claimed as an alternative “or”), or wherein a value included in the bitstream is coded with at least one arithmetic coding context (Paragraph [0039]: “The entropy encoder 210 can implement any appropriate entropy encoding technique, e.g., an arithmetic coding technique, a range coding technique, or a Huffman coding technique. The compressed code symbols 134 may be represented in any of a variety of ways, e.g., as a bit string”), or wherein the visual data comprise a picture of a video or an image (Paragraph [0021]: “[…] a data compression system and a data decompression system. The compression system is configured to process input data (e.g., image data, audio data, video data, text data, or any other appropriate sort of data) to generate a compressed representation of the input data”). Regarding claim(s) 16, Minnen teaches the method of claim 1, wherein the quantized latent representation is generated based on applying a first neural network in the NN-based model to the visual data (Paragraph [0005]: “processing data using a first encoder neural network to generate a latent representation of the data; processing the latent representation of data, comprising: processing the latent representation of data by a first quantizer to generate a quantized latent representation of data”; Paragraph [0029]: “The first encoder neural network 110 is configured to process the input data 105 (x) to generate a latent representation 111 (y) of the input data 105”; and Paragraph [0030]: “the compression system 100 quantizes the latent representation 111 […] to generate […] as quantized latent representation of data”). Regarding claim(s) 17, Minnen teaches the method of claim 1, wherein the conversion includes encoding the visual data into the bitstream, or wherein the conversion includes decoding the visual data from the bitstream (Paragraph [0021]: “The compression system is configured to process input data […] to generate a compressed representation of the input data […] The decompression system can process the compressed data to generate a (approximate or exact) reconstruction of the input data”; Paragraph [0024]: “The compression system compresses the code symbols by entropy encoding […] The compression system then generates the compressed representation of the input data […]”; Paragraph [0025]: “The decompression system can decompress the data […] The decompression system can then reconstruct the original input data”; and Paragraph [0028]: “The compression system 100 processes the input data 105 to generate compressed data 140 representing the input data 105 […]”). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minnen et al (US 2023/0206512 A1) in view of Tjandrasuwita et al (US 2006/0078211 A1). Regarding claim(s) 12, Minnen teaches the method of claim 11, wherein determining the at least a part of the quantized latent representation from the quantized latent representation comprises: determining whether the at least a part of the quantized latent representation comprises a sample of the quantized latent representation (Paragraph [0066]: “[…] generate an ordered collection of code symbols 116 (ŷ) also referred to as quantized latent representation of data […]”; Paragraph [0068]: “The quantized latent representation of data is processed into a plurality of slices of quantized latent representations of data such that each slice of quantized latent representation of data is different from the other slices of quantized latent representation of data”; and Paragraph [0070]: “[…] the quantized latent representation of data 116(ŷ) is processed into two slices […]”) based on at least one of the following: (Paragraph [0071]: “generate as output hyperprior parameters mu 126 (μ) and sigma 127 (σ) that represents the probability distribution of the entropy model 121 (z) and a compressed representation of the hyperprior parameters 142”; and Paragraph [0074]: “receive as input the hyperprior parameters (mu and sigma) representing the probability distribution of the entropy model generated using the hyperprior processing network 125 and the first slice of quantized latent representation of data 118 (y1)”) comprises: determining whether the at least a part of the quantized latent representation comprises samples in a block of the quantized latent representation based on at least one of the following: (Paragraph [0068]: “The quantized latent representation of data is processed into a plurality of slices of quantized latent representations of data such that each slice of quantized latent representation of data is different from the other slices of quantized latent representation of data”; Paragraph [0070]: “[…] the quantized latent representation of data 116(ŷ) is processed into two slices […]”; Paragraph [0071]: “generate as output hyperprior parameters mu 126 (μ) and sigma 127 (σ) that represents the probability distribution of the entropy model 121 (z) and a compressed representation of the hyperprior parameters 142”; and Paragraph [0074]: “receive as input the hyperprior parameters (mu and sigma) representing the probability distribution of the entropy model generated using the hyperprior processing network 125 and the first slice of quantized latent representation of data 118 (y1)”). Minnen fails to teach a comparison between a first threshold and a statistical value corresponding to the sample, a comparison between a second threshold and a value determined based on the statistical value, or a comparison between a third threshold and an index of the sample and a comparison between a fourth threshold and a statistical value corresponding each of the samples, a comparison between a fifth threshold and a metric determined based on statistical values corresponding the samples, or a comparison between a sixth threshold and an index of each of the samples. However, Tjandrasuwita teaches a comparison between a first threshold and a statistical value corresponding to the sample, a comparison between a second threshold and a value determined based on the statistical value, or a comparison between a third threshold and an index of the sample (Paragraph [0061]: “during run-length encoding of a block, should the number of bits encoded reach the threshold in the budget (the "first threshold") […] those transformed and quantized values that are "smaller" than a second threshold value are not encoded […] the second threshold corresponds to the respective magnitudes of the transformed and quantized values […] Alternatively, the second threshold value can correspond to the number of bits needed to encode the transformed and quantized values”; and Paragraph [0062]: “once the first (budget) threshold is reached, the value of the second threshold can be increased as the number of encoded bits approaches the budget limit, so that progressively larger transformed and quantized values within the block are not encoded, leaving budget for the largest of the values remaining to be encoded”) and a comparison between a fourth threshold and a statistical value corresponding each of the samples, a comparison between a fifth threshold and a metric determined based on statistical values corresponding the samples, or a comparison between a sixth threshold and an index of each of the samples (Figure 7; Paragraph [0074]: “ the values in the block are run-length encoded. In step 704, a value in the block is accessed and the number of bits associated with that value is counted”; Paragraph [0077]: “a determination is made as to whether a first threshold value within the budget has been reached or exceeded […] If the first threshold has not been reached or exceeded, then flowchart 700 proceeds to step 712. Otherwise, flowchart 700 proceeds to step 710”; Paragraph [0078]: “a determination is made as to whether a second threshold value is satisfied […] the second threshold value corresponds to the size (e.g., magnitude or number of bits) associated with the value accessed in step 704. If the value accessed in step 704 does not satisfy the second threshold, then flowchart 700 returns proceeds to step 714. Otherwise, flowchart 700 proceeds to step 712”; and Paragraph [0080]: “a determination is made as to whether the value at hand is the last value in the block. If not, then flowchart 700 returns to step 704, where the encoding process is started for a new value in the block. Otherwise, flowchart 700 proceeds to step 716.”). Minnen teaches a learned image compression framework in which quantized latent samples are entropy coded using statistical parameters (e.g., μ and σ) associated with the quantized latent representation to determine coding probabilities for entropy coding. Tjandrasuwita further teaches selectively determining whether individual transformed and quantized values are encoded by comparing threshold criteria during the encoding process, such that transformed and quantized values that do not satisfy the threshold criteria are not encoded, thereby efficiently compressing image data into a target file or bitstream size without significantly reducing image fidelity. Therefore, it would have been obvious to one of ordinary skill in the art to combine the threshold based selective encoding technique of Tjandrasuwita into the learned image compression framework of Minnen before the effective filing date of the claimed invention. The motivation for this combination of references would have been to modify the entropy coding process of Minnen by incorporating Tjandrasuwita’s threshold-based selective encoding technique in order to determine whether at least a part of the quantized latent representation comprises a sample based on a comparison between threshold values and information associated with the sample, thereby selectively encoding samples during entropy coding while efficiently compressing image data into a target file or bitstream size without significantly reducing image fidelity. This motivation for the combination of Minnen and Tjandrasuwita is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minnen et al (US 2023/0206512 A1) in view of Tjandrasuwita et al (US 2006/0078211 A1), further in view of Ikonin et al (US 2023/0262243 A1). Regarding claim(s) 13, Minnen as modified by Tjandrasuwita teaches the method of claim 12, but do not specifically teach wherein the metric is an average, a minimum or a maximum, or wherein an index of a sample indicates one of the following: a channel number of the sample, a feature map identifier of the sample, or a spatial coordinate of the sample, or wherein at least one of the following thresholds or an indication of at least one of the following is indicated in the bitstream: the first threshold, the second threshold, the third threshold, the fourth threshold, the fifth threshold, or the sixth threshold. However, Ikonin teaches wherein the metric is an average, a minimum or a maximum (Wherein clause is claimed as an alternative “or”), or wherein an index of a sample indicates one of the following: a channel number of the sample, a feature map identifier of the sample, or a spatial coordinate of the sample (Paragraph [0166]: “In the table below an exemplary implementation of the bitstream syntax is provided: for (i = 0; i < channels_num; i++) ... for (y = 0; y < latent_space_height[i]; y++) ... for(x = 0; x < latent_space_width[i]; x++) ... y_cap[i][y][x] = decode_latent_value(i, x, y)”; Paragraph [0167]: “Variable channels_num defines number of channels over which is iterated […] Variables latent_space_height and latent_space_width can be derived based on high-level syntax information about picture width and height and architecture of the generative model”; Paragraph [0168]: “One of these values indicating presence […] indicating absence of the channel (feature map) data. It is further noted that in FIG. 10, the channel indicator is signaled which indicates that an entire channel is present/absent”; and Paragraph [0253]: “The sample positions are positions (index) of the transformed coefficients […]”), or wherein at least one of the following thresholds or an indication of at least one of the following is indicated in the bitstream: the first threshold, the second threshold, the third threshold, the fourth threshold, the fifth threshold, or the sixth threshold (Wherein clause is claimed as an alternative “or”). Therefore, it would have been obvious to one of ordinary skill in the art to combine Minnen, Tjandrasuwita, and Ikonin before the effective filing date of the claimed invention. The motivation for this combination of references would have been to incorporate Ikonin’s explicit indexing of latent samples by channel and spatial position into Minnen’s learned image compression framework, thereby facilitating efficient identification, access, and processing of individual latent samples during encoding and decoding while maintaining compatibility with Minnen’s latent representation architecture. This motivation for the combination of Minnen, Tjandrasuwita, and Ikonin is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Relevant Prior Art Directed to State of Art Besenbruch et al (US 2022/0279183 A1) are relevant prior art not applied in the rejection(s) above. Besenbruch discloses a computer implemented method of training a first neural network and a second neural network, the neural networks being for use in lossy image or video compression, transmission and decoding, the method including the steps of: (i) receiving an input training image; (ii) encoding the input training image using the first neural network, to produce a latent representation; (iii) quantizing the latent representation to produce a quantized latent; (iv) using the second neural network to produce an output image from the quantized latent, wherein the output image is an approximation of the input image; (v) evaluating a loss function based on differences between the output image and the input training image; (vi) evaluating a gradient of the loss function; (vii) back-propagating the gradient of the loss function through the second neural network and through the first neural network, to update weights of the second neural network and of the first neural network; and (viii) repeating steps (i) to (vii) using a set of training images, to produce a trained first neural network and a trained second neural network, and (ix) storing the weights of the trained first neural network and of the trained second neural network; wherein the loss function is a weighted sum of a rate term and a distortion term, wherein split quantisation is used during the evaluation of the gradient of the loss function, with a combination of two quantisation proxies for the rate term and the distortion term. Zhou et al (US 2021/0297667 A1) are relevant prior art not applied in the rejection(s) above. Zhou discloses a training device for an image processing apparatus, in which an image encoder and an image decoder are trained by using a training image, the training device comprises: a memory to store a plurality of instructions; and a processor coupled to the memory and configured to: acquire a latent variable obtained by the image encoder by encoding input training image data; acquire first restored image data obtained by the image decoder by decoding the latent variable and second restored image data obtained by the image decoder by decoding a sum of the latent variable and a noise; and train the image encoder and the image decoder according to a cost function, the cost function being related to a deviation between the input training image data and the first restored image data and a deviation between the first restored image data and the second restored image data. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONGBONG NAH whose telephone number is (571) 272-1361. The examiner can normally be reached M - F: 9:00 AM - 5:30 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ONEAL MISTRY can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JONGBONG NAH/Examiner, Art Unit 2674
Read full office action

Prosecution Timeline

Sep 03, 2024
Application Filed
Aug 03, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705791
LOCALIZATION PROCESSING SERVICE
4y 2m to grant Granted Aug 11, 2026
Patent 12705905
ROAD BOUNDARY DETECTION BASED ON RADAR AND VISUAL INFORMATION
3y 9m to grant Granted Aug 11, 2026
Patent 12688721
DETECTING A CONDITION FOR A CULTURE DEVICE USING A MACHINE LEARNING MODEL
3y 11m to grant Granted Jul 21, 2026
Patent 12659419
AUGMENTED REALITY SELF-PORTRAITS
3y 11m to grant Granted Jun 16, 2026
Patent 12645937
SYSTEM, METHOD, AND COMPUTER PROGRAM FOR ITERATIVE CONTENT ADAPTIVE ONLINE TRAINING IN NEURAL IMAGE COMPRESSION
3y 8m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
76%
Grant Probability
93%
With Interview (+16.3%)
2y 10m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 118 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month