Prosecution Insights
Last updated: October 02, 2026
Application No. 18/294,785

METHOD AND DATA PROCESSING SYSTEM FOR LOSSY IMAGE OR VIDEO ENCODING, TRANSMISSION AND DECODING

Non-Final OA §103
Filed
Feb 02, 2024
Priority
Aug 03, 2021 — GB 2111188.5 +1 more
Examiner
FIGUEROA, KEVIN W
Art Unit
Tech Center
Assignee
InterDigital Inc.
OA Round
1 (Non-Final)
70%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
264 granted / 376 resolved
+10.2% vs TC avg
Strong +21% interview lift
Without
With
+20.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
22 currently pending
Career history
393
Total Applications
across all art units

Statute-Specific Performance

§101
25.4%
-14.6% vs TC avg
§103
56.1%
+16.1% vs TC avg
§102
6.2%
-33.8% vs TC avg
§112
6.4%
-33.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 376 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Specification The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 97-99, 102, 104-106 109-111, and 113 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minnen et al. US 2020/0027247 in view of Wen et al. US2020/0372684. Regarding claim 97, Minnen teaches “a method of training one or more neural networks, the one or more neural networks being for use in lossy image or video encoding, transmission and decoding, the method comprising the steps of” (abstract “method comprises: processing data using an encoder neural network to generate a latent representation of the data”): “receiving an input image at a first computer system” ([0006] “In some implementations, the data includes an image and the encoder neural network is a convolutional neural network.”); “encoding the input image using a first neural network to produce a latent representation” ([0066] “The encoder neural network 106 is configured to process the input data 102 (x) to generate a latent representation 116 (y) of the input data 102. As used throughout this document, a “latent representation” of data refers to a representation of the data as an ordered collection of numerical values, e.g., a vector or matrix of numerical values”); “entropy encoding the latent representation” ([0067] “To facilitate compression of the latent representation 116 of the input data using entropy encoding techniques, the compression system 100 quantizes the latent representation 116 of the input data using a quantizer Q 118 to generate an ordered collection of code symbols 120”); “transmitting the entropy encoded latent representation to a second computer system” (fig. 1 shows the compression system and fig. 2 shows the decompression system, two separate systems i.e. a first and second computer system, [0079] “The decompression system 200 is an example system implemented as computer programs on one or more computers in one or more locations in which the systems, components” and [0080] “The decompression system 200 processes the compressed data 104 generated by the compression system to generate a reconstruction 202 that approximates the original input data”); “entropy decoding the entropy encoded latent representation” ([0113] “the system sequentially entropy decodes the code symbols of the quantized latent representation of the data.”) and “decoding the latent representation using a second neural network to produce an output image, wherein the output image is an approximation of the input image” ([0114] “The system determines the reconstruction of the data by processing the quantized latent representation of the data using the decoder neural network (810). In one example, the decoder neural network may be a de-convolutional neural network that processes the quantized latent representation of image data to generate an (approximate or exact) reconstruction of the image data.”); “determining a quantity based on a difference between the output image and the input image” ([0019] “In some implementations, the rate-distortion performance measure includes: (i) a first rate term based on a size of the entropy encoded representation of the latent representation of the data, (ii) a second rate term based on a size of the entropy encoded representation of the latent representation of the entropy model, and (iii) a distortion term based on a difference between the data and a reconstruction of the data.”); “updating the parameters of the first neural network and the second neural network based on the determined quantity” ([0087] “The compression system and the decompression system can be jointly trained using machine learning training techniques (e.g., stochastic gradient descent) to optimize a rate-distortion objective function. More specifically, the encoder neural network, the hyper-encoder neural network, the hyper-decoder neural network, the context neural network, the entropy model neural network, and the decoder neural network can be jointly trained to optimize the rate distortion objective function.”); and “repeating the above steps using a first set of input images to produce a first trained neural network and a second trained neural network” (previous citation, neural networks are not limited to one pass and training is always iterative); While Minnen generally operates in a pixel by pixel basis, Wen more specifically teaches “wherein the entropy decoding of the entropy encoded latent representation is performed pixel by pixel” (Wen [0044] “since the entropy model is very important for image compression, as a part of the input of the entropy model, the context model may effectively improve accuracy of prediction by using information on a pixel preceding a current pixel. However, as the context model is an autoregressive network, the latent representation is coded pixel by pixel” and [0039] “the image coding apparatus 101 is used to transform the input image (i.e. pixels of the input image in the embodiment of this disclosure) into a latent representation that is able to reduce a dimensional space (that is, dimension reduction). The image decoding apparatus 103 attempts to map the latent representation back to the above pixels via an approximate inverse function.”); and “the order of the pixel by pixel decoding is additionally updated based on the determined quantity” (previous citation, Wen [0044] “since the entropy model is very important for image compression, as a part of the input of the entropy model, the context model may effectively improve accuracy of prediction by using information on a pixel preceding a current pixel. However, as the context model is an autoregressive network, the latent representation is coded pixel by pixel” wherein the order is updated based on the quantity since the image changes based on the quantity determined) It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Minnen with that of Wen since a combination of known methods would yield predictable results. As shown in Wen, it is understood in the art that some methods operate on a pixel basis when analyzing image data. Therefore these techniques would translate easily and perform predictably in the system above when combined. Note that independent claims 110 and 111 recite broader limitations of independent claim 97, including using the method of claim 97 as training. Therefore the claims are subject to the same rejection as they overlap in scope. Regarding claim 98, the Minnen and Wen references have been addressed above. Wen further teaches “wherein the order of the pixel by pixel decoding is based on the latent representation” (Wen [0044] “since the entropy model is very important for image compression, as a part of the input of the entropy model, the context model may effectively improve accuracy of prediction by using information on a pixel preceding a current pixel. However, as the context model is an autoregressive network, the latent representation is coded pixel by pixel”) Regarding claim 99, the Minnen and Wen references have been addressed above. Wen further teaches “wherein the entropy decoding of the entropy encoded latent comprises an operation based on previously decoded pixels” (Wen [0044] “since the entropy model is very important for image compression, as a part of the input of the entropy model, the context model may effectively improve accuracy of prediction by using information on a pixel preceding a current pixel. However, as the context model is an autoregressive network, the latent representation is coded pixel by pixel” i.e. autoregressive) Regarding claim 102, the Minnen and Wen references have been addressed above. Wen further teaches “wherein the determining of the order of the pixel by pixel decoding comprises dividing the latent representation into a plurality of sub-images” (Wen [0044] “since the entropy model is very important for image compression, as a part of the input of the entropy model, the context model may effectively improve accuracy of prediction by using information on a pixel preceding a current pixel. However, as the context model is an autoregressive network, the latent representation is coded pixel by pixel” a pixel is a sub-image) Regarding claim 104, the Minnen and Wen references have been addressed above. Minnen further teaches “wherein the determining of the order of the pixel by pixel decoding comprises ranking a plurality of pixels of the latent representation based on the magnitude of a quantity associated with each pixel” ([0041] “For example, if the input data is an image, the compression system may directly entropy encode the pixel intensity/color values of the image. When the compression system directly entropy encodes the components of the input data (i.e., without using an encoder neural network), the decompression system similarly directly entropy decodes the components of the input data (i.e., without using a decoder neural network).” i.e. the decoding depends on the magnitude/intensity of the pixel, which inherently has a ranked order i.e. brightest to least) Regarding claim 105, the Minnen and Wen references have been addressed above. Minnen further teaches “wherein the quantity associated with each pixel is the location or scale parameter associated with that pixel” (previous citation, brightness/intensity is a scale parameter i.e. a scale of how bright or dark it is) Regarding claim 106, the Minnen and Wen references have been addressed above. Minnen further teaches “wherein the quantity associated with each pixel is additionally updated based on the evaluated difference” (previous citations, the image/context changes with each iteration) Regarding claim 109, the Minnen and Wen references have been addressed above. Minnen further teaches “further comprising the steps of: encoding the latent representation using a fourth trained neural network to produce a hyper-latent representation” ([0046] “To determine the conditional entropy model, the system can process a quantized latent representation of the data using a “hyper-encoder” neural network to generate a “hyper-prior” that implicitly characterizes the conditional entropy model. The hyper-prior is subsequently compressed and included as side-information in the compressed representation of the input data. Generally, a more complex hyper-prior can specify a more accurate conditional entropy model that enables the input data to be compressed at a higher rate”); “transmitting the hyper-latent to the second computer system; and decoding the hyper-latent using a fifth trained neural network, wherein the order of the pixel by pixel decoding is based on the output of the fifth trained neural network” ([0068] “The compression system 100 uses the hyper-encoder neural network 108, the hyper-decoder neural network 110, and the entropy model neural network 112 to generate a conditional entropy model for entropy encoding the code symbols 120 representing the input data, as will be described in more detail next”) Regarding claim 113, the Minnen and Wen references have been addressed above. They both further teaches “a data processing system configured to perform the method of claim 97” (both references contain data processing systems /neural networks to perform the steps as claimed) Claim(s) 100 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minnen in view of Wen, further in view of Vosoughi et al. US 2019/0114805 [herein Vos]. Regarding claim 100, the Minnen and Wen references have been addressed above. They do not explicitly teach the claim limitations. Vos however teaches “wherein the determining of the order of the pixel by pixel decoding comprises ordering a plurality of the pixels of the latent representation in a directed acyclic graph” (Vos [0016] “FIG. 1 illustrates a flowchart of a palette coding encoder and decoder according to some embodiments. An encoder 100 is utilized to implement the palette coding for color compression. In the step 102, a serialized Directed Acyclic Graph (DAG) tree with uncompressed color of a cloud point is received/acquired/generated. In the step 104, a Matrix M is generated. The Matrix M contains rgb components of pixels from the point cloud. I”) It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Minnen and Wen with that of Vos since a combination of known methods would yield predictable results. There are numerous methods to determine how to decode pixel data and Vos teaches one of the many ways of doing so, by using a graph to show adjacency. Therefore these techniques would operate in a known and predictable manner with the systems above. Claim(s) 101 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minnen in view of Wen further in view of Fracastoro et al. US 2018/0262759. Regarding claim 101, the Minnen and Wen references have been addressed above. They do not explicitly teach the claim limitations. Fracastoro however teaches “wherein the determining of the order of the pixel by pixel decoding comprises operating on the latent representation with a plurality of adjacency matrices” (Fracastoro [0082] “In FIG. 4(a), the reference pixel (black) is located at the center of the image, and has four adjacencies (dark grey pixels) on top, left, bottom and right, whereas in an embodiment of the invention the remaining four pixels (light gray) are preferably not considered adjacencies of the reference pixel. By using this adjacency structure, i.e., a square grid where each pixel is a vertex of the graph and is connected to each of its 4-connected neighbors, it is possible to simplifying advantageously the determination of the eigenvectors, since it has been proved that, when the graph is a 4-connected grid and all the four edges have the same weight, the 2D DCT basis functions are eigenvectors of the graph Laplacian, and thus the transform matrix U can be the 2D DCT matrix.”) It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Minnen and Wen with that of Fracastoro since a combination of known methods would yield predictable results. There are numerous methods to determine how to decode pixel data and Fracastoro teaches one of the many ways of doing so, by using adjacency matrices to show neighboring pixels. Therefore these techniques would operate in a known and predictable manner with the systems above. Claim(s) 103 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minnen in view of Wen further in view of Li, Mu, et al. "Efficient and effective context-based convolutional entropy modeling for image compression." Regarding claim 103, the Minnen and Wen references have been addressed above. While Wen generally teaches binary masking, Li more explicitly teaches “wherein the plurality of sub-images are obtained by convolving the latent representation with a plurality of binary mask kernels” (Li pg. 3 left col. last ¶ “Together, the two assumptions guarantee the legitimacy of context-based entropy modeling in fully convolutional networks, which can be achieved by placing translation invariant binary masks to convolution filters”) It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Minnen and Wen with that of Li since as shown in Li, it outright states that binary masking is used for convolution filters, which is what is happening in Wen. This is a common and known technique in the art and thus would combine with the systems above. Claim(s) 107-108 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minnen in view of Wen, further in view of Chou et al. US 2021/0250616. Regarding claim 107, the Minnen and Wen references have been addressed above. The references do not explicitly teach the claim limitations. Chou however teaches “wherein the determining of the order of the pixel by pixel decoding comprises a wavelet decomposition of a plurality of pixels of the latent representation” (Chou [0032] “In some embodiments, the video encoding system may perform a wavelet transform on the pixel data prior to encoding to decompose the pixel data into frequency bands. The frequency bands are then organized into blocks that are provided to a block-based encoder for encoding/compression. As an example, a frame may be divided into 128×128 blocks, and a two-level wavelet decomposition may be applied to each 128×128 block to generate 16 32×32 blocks of frequency data representing seven frequency bands that may then be sent to an encoder (e.g., a High Efficiency Video Coding (HEVC) encoder) to be encoded.” ) It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Minnen and Wen with that of Chou. As shown in Chou, Chou describes wavelet decomposition, a known mathematical technique that operates on the data used by the systems above. Therefore by combining these techniques, the system would have alternate methods of processing the data in order to achieve its results. Regarding claim 108, the Minnen, Wen, and Chou references have been addressed above. Chou further teaches “wherein the order of the pixel by pixel decoding is based on the frequency components of the wavelet decomposition associated with the plurality of pixels” (previous citation “The frequency bands are then organized into blocks that are provided to a block-based encoder for encoding/compression”) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Chen, Honggang, et al. "DPW-SDNet: Dual pixel-wavelet domain deep CNNs for soft decoding of JPEG-compressed images." 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). IEEE, 2018. Lee et al. US 2020/0107023 Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEVIN W FIGUEROA whose telephone number is (571)272-4623. The examiner can normally be reached Monday-Friday, 10AM-6PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. KEVIN W FIGUEROA Primary Examiner Art Unit 2124 /Kevin W Figueroa/ Primary Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Feb 02, 2024
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743727
Machine Learning Portfolio Simulating and Optimizing Apparatuses, Methods and Systems
5y 2m to grant Granted Sep 22, 2026
Patent 12743437
Method And System For Implementing Machine Learning Classifications
2y 4m to grant Granted Sep 22, 2026
Patent 12731033
NON-LINEAR LATTICE LAYER FOR PARTIALLY MONOTONIC NEURAL NETWORK
4y 10m to grant Granted Sep 08, 2026
Patent 12731040
DEVICE, A COMPUTER PROGRAM AND A COMPUTER-IMPLEMENTED METHOD FOR DETERMINING NEGATIVE SAMPLES FOR TRAINING A KNOWLEDGE GRAPH EMBEDDING OF A KNOWLEDGE GRAPH
4y 2m to grant Granted Sep 08, 2026
Patent 12731041
DISTRIBUTED TRAINING PROCESS WITH BOTTOM-UP ERROR AGGREGATION
3y 11m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
70%
Grant Probability
91%
With Interview (+20.7%)
3y 11m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 376 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month