Prosecution Insights
Last updated: October 01, 2026
Application No. 17/644,372

TASK-DEPENDENT SELECTION OF DECODER-SIDE NEURAL NETWORK

Non-Final OA §103§112
Filed
Dec 15, 2021
Examiner
BOSTWICK, SIDNEY VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Nokia Corporation
OA Round
3 (Non-Final)
51%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
78 granted / 152 resolved
-3.7% vs TC avg
Strong +35% interview lift
Without
With
+35.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
40 currently pending
Career history
216
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
4.6%
-35.4% vs TC avg
§112
24.0%
-16.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 152 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 7/28/2026 has been entered. Remarks This Office Action is responsive to Applicants' Amendment filed on July 28, 2026, in which claims 1, 5, 7, 11, 15, and 19 are currently amended. Claims 1-9 and 11-21 are currently pending. Specification Applicant's amendments made to the specification are acknowledged. Examiner’s objection to the specification are hereby withdrawn, as necessitated by Applicant’s amendments made to the specification. Response to Arguments Applicant’s arguments with respect to rejection of claims 1-9 and 11-21 under 35 U.S.C. 112(b) based on amendment have been considered. With respect to Applicant’s arguments on p. 10 of the Remarks submitted 7/28/2026 that the amended limitations are operative to “removing any ambiguity” directed towards claims 1 and 15, Examiner respectfully disagrees. "Wherein the new task neural network was one of: not used, not known at the codec design stage, or not associated with one or more of the plurality of decoder side neural networks" is indefinite. While none of the alternative choices have clear bounds, "not used" is especially problematic. It would not be clear to one of ordinary skill in the art when or how the new task neural network was not used: not used ever, within a time range, in designing the codec, by the decoder, for the new task, or something else altogether. These are all considerably different readings such that the scope of the claim is ambiguous. "not associated with one or more of the plurality of decoder side neural networks" is similarly ambiguous as the scope of the one to many association is unclear: The broadest reasonable interpretation of the claim would appear to anticipate none of the plurality of decoder side neural networks having any association whatsoever with the new task neural network, which appears logically inconsistent with the claim and the instant specification. For at least these reasons Examiner asserts that the scope of the claim cannot reasonably be determined. With respect to Applicant’s arguments on p. 10 of the Remarks submitted 7/28/2026 directed towards claims 5 and 19, while Examiner agrees that the amendments strengthen the clarity, the claim still lacks proper antecedent basis in multiple locations and does not address the ambiguous relative clause (“that”). Specifically, regarding claims 1, 5, 15, and 19, "a known association with a decoder side neural network" lacks antecedent basis. Claims 1 and 15 recite "selecting a decoder side neural network" such that it's unclear if the "decoder side neural network" in "a known association with a decoder side neural network" is the same as the "decoder side neural network" in "selecting a decoder side neural network" or another separate decoder side neural network. Similarly, claims 5 and 19 recite "a decoder side neural network" and "that decoder side neural network" which lack antecedent basis. With respect to Applicant’s arguments on p. 10 of the Remarks submitted 7/28/2026 directed to claim 7, "wherein the two or more decoder side neural networks are associated with a different level of at least one of" is grammatically indefinite. The claim limitation could be interpreted as "Networks A and B are associated with a different level of computational complexity" which can naturally mean A and B, collectively, are associated with a particular level that is different from some other level. Alternatively, it could mean that Networks A and B are associated with unique respective levels. The word "different" has no explicit comparator making the intended meaning clear such that the scope of the claim cannot reasonably be determined. In the interest of further examination the claim is interpreted as the two or more decoder side neural networks are associated with respective unique levels. With respect to Applicant’s arguments on p. 10 of the Remarks submitted 7/28/2026 that the amendments to claim 11, Applicant’s arguments are persuasive. The rejection is withdrawn in view of Applicant’s amendments and remarks made to the rejection. Applicant’s arguments with respect to rejection of claims 1-9 and 11-20 under 35 U.S.C. 103 based on amendment have been considered, however, are not persuasive. With respect to Applicant’s arguments on p. 11 of the Remarks submitted 7/28/2026 that “the amendments makes it clear that the claim suggests comparing the features of an entirely new task with the features of known tasks, and then picks the decoder-side neural network linked to the most similar known task and is significantly more specific than merely comparing task-specific training parameters”, Examiner respectfully disagrees. The amendment claims “wherein, to select the decoder side neural network, the apparatus is further caused to: run at least part of a new task neural network on data derived from a bitstream received by a decoder, wherein the new task neural network was one of: not used, not known at the codec design stage, or not associated with one or more of the plurality of decoder side neural networks;”, which as noted above in the response to arguments under 35 USC §112, is highly ambiguous. Examiner asserts that given the uncertain scope of the actual claim language it would be very reasonable to interpret Meyerson who explicitly anticipates using a new task neural network for a new task ([¶0032] "Learning a different soft ordering of layers for each task amounts to discovering a set of generalizable modules that are assembled in different ways for different tasks. This perspective points to future approaches that train a collection of layers on a set of training tasks, which can then be assembled in novel ways for future unseen tasks.") and explicitly anticipates computing layer-wise differences ([¶0073] "For a representative two-task soft order experiment the layer-wise distance between scalings of the tasks increases by iteration, as shown by (b) in FIG. 5") as covering the instant claims, especially since the instant claims appear to analogize neural network layers/subnetworks to neural networks (i.e. decoder side neural network). With respect to Applicant’s arguments on p. 12 of the Remarks submitted 7/28/2026 that “Meyerson’s tensor S is a set of task specific scaling parameters learned jointly during training […] and not “features extracted from the new task neural network”, Examiner respectfully disagrees. Examiner asserts that a “task specific scaling parameter” is a feature which is explicitly learned jointly from the new task neural network gradient update ([¶0068] “S can be learned jointly with the other learnable parameters in the Wkεi,[…] via backpropagation” [¶0110] “The optimization algorithm can be based on stochastic gradient descent”). With respect to Applicant’s arguments on p. 12 of the Remarks submitted 7/28/2026 that the explicit distance between the scaling parameters (features) is “not a distance computed between the features of a new task neural network and the features of task neural networks associated with known tasks”, Examiner respectfully disagrees. Meyerson teaches ([¶0073] "For a representative two-task soft order experiment the layer-wise distance between scalings of the tasks increases by iteration, as shown by (b) in FIG. 5"). As noted above the instant claims and specification support this interpretation as the instant claims appear to analogize neural network layers/subnetworks to neural networks (i.e. decoder side neural network). Meyerson further discloses ([¶0035] “FIG. 2, (a) shows classical approaches which add a task-specific decoder to the output of the core single-task model for each task. In FIG. 2, (b) shows column based approaches which include a network column for each task and define a mechanism for sharing between columns. In FIG. 2, (c) shows supervision at custom depths which add output decoders at depths based on a task hierarchy” [¶0118] “Speciation can create new subpopulations, add new items to pre-existing subpopulations, remove pre-existing items from pre-existing subpopulations, move pre-existing items from one pre-existing subpopulation to another pre-existing subpopulation, move pre-existing items from a pre-existing subpopulation to a new subpopulation, and so on. For example, a population of items is divided into subpopulations such that items with similar topologies, i.e., topology hyperparameters, are in the same subpopulation.”). With respect to Applicant’s arguments on p. 12 of the Remarks submitted 7/28/2026 that “selection premised on a task whose identity and decoder are already known”, Examiner respectfully disagrees. As noted above, Meyerson explicitly anticipates ([¶0032] "Learning a different soft ordering of layers for each task amounts to discovering a set of generalizable modules that are assembled in different ways for different tasks. This perspective points to future approaches that train a collection of layers on a set of training tasks, which can then be assembled in novel ways for future unseen tasks."). Similarly, Applicant’s arguments on p. 12 of the Remarks submitted 7/28/2026 that Meyerson “does not select a decoder side neural network on the basis of computed feature distance to a new task” relies on an unreasonably narrow interpretation of “based on”. Examiner notes that instant claim 5 which this argument appears to be directed towards recites “smallest among a plurality of distance metrics that are computed” where “a plurality of distance metrics that are computed” is not pre-defined in the claim or having any antecedent basis that would limit the claim in a way to make Examiner’s interpretation unreasonable. For at least these reasons and those further detailed below, Examiner asserts that the interpretation of the combination of Meyerson and Hoogeboom to cover the instant claims is reasonable and should be maintained. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-9 and 11-21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 1, "to select an optimal decoder side neural network" is indefinite. "Optimal" is a relative term with no relative basis for comparison. In the interest of further examination the claim is interpreted as "to select a second decoder side neural network". Regarding claim 1, "wherein the new task neural network was one of: not used, not known at the codec design stage, or not associated with one or more of the plurality of decoder side neural networks" is indefinite. While none of the alternative choices have clear bounds, "not used" is especially problematic. It would not be clear to one of ordinary skill in the art when or how the new task neural network was not used: not used ever, within a time range, in designing the codec, by the decoder, for the new task, or something else altogether. These are all considerably different readings such that the scope of the claim is ambiguous. "not associated with one or more of the plurality of decoder side neural networks" is similarly ambiguous as the scope of the one to many association is unclear: The broadest reasonable interpretation of the claim would appear to anticipate none of the plurality of decoder side neural networks having any association whatsoever with the new task neural network, which appears logically inconsistent with the claim and the instant specification. For at least these reasons Examiner asserts that the scope of the claim cannot reasonably be determined. In the interest of further examination "not used" is interpreted as "not used for decoding". Regarding claims 1, 5, 15, and 19, "a known association with a decoder side neural network" lacks antecedent basis. Claims 1 and 15 recite "selecting a decoder side neural network" such that it's unclear if the "decoder side neural network" in "a known association with a decoder side neural network" is the same as the "decoder side neural network" in "selecting a decoder side neural network" or another separate decoder side neural network. Similarly, claims 5 and 19 recite "a decoder side neural network" and "that decoder side neural network" which lack antecedent basis. Regarding claim 7, "the two or more decoder side neural networks perform differently in terms of a rate-distortion trade-off or a trade-off between a rate and a task accuracy" is indefinite. "Differently" is a relative term without a relative basis for comparison. One of ordinary skill in the art would recognize that there are different ways to decide whether the trade-offs themselves are different and those different tests could produce different infringement results. Regarding claim 7, "wherein the two or more decoder side neural networks are associated with a different level of at least one of" is grammatically indefinite. The claim limitation could be interpreted as "Networks A and B are associated with a different level of computational complexity" which can naturally mean A and B, collectively, are associated with a particular level that is different from some other level. Alternatively, it could mean that Networks A and B are associated with unique respective levels. The word "different" has no explicit comparator making the intended meaning clear such that the scope of the claim cannot reasonably be determined. In the interest of further examination the claim is interpreted as the two or more decoder side neural networks are associated with respective unique levels. The remaining claims are rejected with respect to their dependence on the rejected claims. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-9 and 11-21 are rejected under U.S.C. §103 as being unpatentable over the combination of Meyerson (US20190130257A1) which incorporates ("Pseudo-task Augmentation: From Deep Multitask Learning to Intratask Sharing—and Back", 2018) by reference, and Hoogeboom (“Integer Discrete Flows and Lossless Compression”, 2019). Regarding claim 1, Meyerson teaches An apparatus comprising at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: ([¶0028] "one or more of the functional blocks (e.g., modules, processors, or memories) can be implemented in a single piece of hardware (e.g., a general purpose signal processor or a block of random access memory, hard disk, or the like) or multiple pieces of hardware. Similarly, the programs can be stand-alone programs, can be incorporated as subroutines in an operating system, can be functions in an installed software package, and the like") organizing a plurality of decoder side neural networks based on one or more task categories or one or more tasks; and ([¶0035] "In FIG. 2, (c) shows supervision at custom depths which add output decoders at depths based on a task hierarchy. In FIG. 2, (d) shows universal representations which adapt each layer with a small number of task-specific scaling parameters") selecting a decoder side neural network based at least on the one or more task categories or the one or more tasks.([0185] "The method includes selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks") wherein the apparatus is further caused to select an optimal decoder side neural network for a new task that was not known at a codec design stage;([¶0032] "Learning a different soft ordering of layers for each task amounts to discovering a set of generalizable modules that are assembled in different ways for different tasks. This perspective points to future approaches that train a collection of layers on a set of training tasks, which can then be assembled in novel ways for future unseen tasks.") wherein the new task neural network was one of: not used, not known at the codec design stage, or not associated with one or more of the plurality of decoder side neural networks;([¶0032] "Learning a different soft ordering of layers for each task amounts to discovering a set of generalizable modules that are assembled in different ways for different tasks. This perspective points to future approaches that train a collection of layers on a set of training tasks, which can then be assembled in novel ways for future unseen tasks.") extract features from the new task neural network;([¶0035] "Underlying each of these approaches is the assumption of parallel ordering of shared layers. Also, each one requires aligned sequences of feature extractors across tasks.") extract features from the task neural network; ([¶0035] "Underlying each of these approaches is the assumption of parallel ordering of shared layers. Also, each one requires aligned sequences of feature extractors across tasks." [¶0045] "A common interpretation of deep learning is that layers extract progressively higher level features at later depths" [¶0016] "the feature hierarchy implied by the model architecture." [¶0035] "adapt each layer with a small number of task-specific scaling parameters" [¶0066] "Soft ordering (Eq. 7) generalizes Eqs. 2 and 3 by learning a tensor S of task-specific scaling parameters" Meyerson explicitly discloses that the layers themselves are feature extractors, the model architecture being a feature hierarchy, the feature hierarchy including extracted task-specific scaling parameters (tensor S which is a latent representation) which are extracted through learning) and compute a distance metric between the features extracted from the new task neural network and the features extracted from the task neural network associated with the task;([¶0073] "For a representative two-task soft order experiment the layer-wise distance between scalings of the tasks increases by iteration, as shown by (b) in FIG. 5" Meyerson explicitly computes a distance metric between the extracted layer-wise (between neural networks), task-specific, scaling features) and select the decoder side neural network associated with the task having a predetermined distance metric or a lowest computed distance metric.([0185] "The method includes selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks" [¶0073] "For a representative two-task soft order experiment the layer-wise distance between scalings of the tasks increases by iteration, as shown by (b) in FIG. 5"). However, Meyerson does not explicitly teach wherein, to select the decoder side neural network, the apparatus is further caused to: run at least part of a new task neural network on data derived from a bitstream received by a decoder, for each task for which a known association with a decoder side neural network exists: use the decoder side neural network associated with the task to process data derived from the bitstream received by the decoder; run at least part of the task neural network on data output by the decoder side neural network associated with the task or data derived from the output of the decoder side neural network;. Hoogeboom, in the same field of endeavor, teaches wherein, to select the decoder side neural network, the apparatus is further caused to: run at least part of a new task neural network on data derived from a bitstream received by a decoder, ([p. 5 §3.3] "Lossless compression is an essential technique to limit the size of representations without destroying information. Methods for lossless compression require i) a statistical model of the source, and ii) a mapping from source symbols to bit stream [...] The mapping between symbols and bit streams may be provided by any entropy encoder" [p. 2 §1] "To obtain x, the decoder uses pZ(·) and c to reconstruct z. Subsequently, z is mapped to x using the inverse of the IDF") for each task for which a known association with a decoder side neural network exists: use the decoder side neural network associated with the task to process data derived from the bitstream received by the decoder;([p. 5 §3.3] "Lossless compression is an essential technique to limit the size of representations without destroying information. Methods for lossless compression require i) a statistical model of the source, and ii) a mapping from source symbols to bit stream [...] The mapping between symbols and bit streams may be provided by any entropy encoder" [p. 2 §1] "To obtain x, the decoder uses pZ(·) and c to reconstruct z. Subsequently, z is mapped to x using the inverse of the IDF") run at least part of the task neural network on data output by the decoder side neural network associated with the task or data derived from the output of the decoder side neural network;([p. 5] "Figure 4: Example of a 2-level flow architecture. The squeeze layer reduces the spatial dimensions by two, and increases the number of channels by four. A single integer flow layer consists of a channel permutation and an integer discrete coupling layer. Each level consists of D flow layers" See FIG. 4 which shows that inner decoder neural networks are inversely processed by providing decoder output to outer decoder layer(s).). Meyerson as well as Hoogeboom are directed towards stacked decoder layer autoencoder architectures. Therefore, Meyerson as well as Hoogeboom are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Meyerson with the teachings of Hoogeboom by simply using both models sequentially (reconstructing bitstream using IDF decoder and then providing that reconstructed image to task-specific Meyerson decoder heads). Hoogeboom provides as additional motivation for combination ([p. 1 §1] “Every day, 2500 petabytes of data are generated. Clearly, there is a need for compression to enable efficient transmission and storage of this data [...] maximizing the log-likelihood (of data) is equivalent to minimizing the expected number of bits required per message”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 2, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 1, wherein the apparatus is further caused to: associate one or more decoder side neural networks of the plurality of decoder side neural networks with the one or more task categories or the one or more tasks; and(Meyerson [¶0185] "selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks and transmitting the accumulated output encoding produced for the final clone as input to the selected decoder." [¶0186] "The method includes processing the accumulated output encoding through the selected decoder to produce classification scores for classes defined for the particular one of the classification tasks.") select the decoder side neural network based on the association between the one or more tasks and the plurality of decoder side neural networks.(Meyerson [¶0185] "selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks and transmitting the accumulated output encoding produced for the final clone as input to the selected decoder."). Regarding claim 3, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 1, wherein the apparatus is further caused to select an optimal decoder side neural network based on one or more predetermined criteria.(Meyerson [¶0169] "The system comprises a decoder selector. The decoder selector selects, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks and transmits the accumulated output encoding produced for the final clone as input to the selected decoder." [¶0170] "The selected decoder processes the accumulated output encoding and produces classification scores for classes defined for the particular one of the classification tasks."). Regarding claim 4, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 1, wherein the plurality of decoder side neural networks comprise one or more shared decoder side neural networks, and wherein the one or more shared decoder side neural networks comprise one or more shared parameters and one or more non-shared parameters, and wherein the one or more shared parameters do not depend on a task to be run on decoded data, wherein the task to be run on the decoded data is a task of the one or more tasks, (Meyerson [Abstract] "The technology disclosed identifies parallel ordering of shared layers as a common assumption underlying existing deep multitask learning (MTL) approaches" [¶0021] "FIG. 4 depicts soft ordering of shared layers." [¶0047] "at each shared depth the applied weight tensors for each task are similar and compatible for sharing" [¶0066] " Each Fj in FIG. 4 includes a shared weight layer and a nonlinearity. This architecture enables the learning of layers that are used in different ways at different depths for different tasks." Meyerson explicitly teaches that the selected decoder is a soft ordering shared decoder with shared and unshared parameters where the shared parameters are not task dependent.) and wherein the one or more non-shared parameters depend at least on the task to be run on the decoded data, (Meyerson [¶0066] "Soft ordering (Eq. 7) generalizes Eqs. 2 and 3 by learning a tensor S of task-specific scaling parameters. S is learned jointly with the Fj, to allow flexible sharing across tasks and depths. Each Fj in FIG. 4 includes a shared weight layer and a nonlinearity. This architecture enables the learning of layers that are used in different ways at different depths for different tasks." tensor S interpreted as non-shared task dependent parameter.) and wherein the apparatus is further caused to select a subset of the one or more non-shared parameters that are to be used by the decoder side neural network associated with the task.(Meyerson [¶0068] "by s(i,σi(k),k)=1∀(i,k). S can be learned jointly with the other learnable parameters in the Wkεi, and Di via backpropagation. In training, all s(i,j,k) are initialized with equal values, to reduce initial bias of layer function across tasks" Meyerson explicitly selects subset s_(i,j,k) from set S where i corresponds to selected decoder Di.). Regarding claim 5, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 1, wherein the decoder side neural network that is selected is a decoder side neural network for which the distance metric between the features extracted from the new task neural network and the features extracted from the task neural network having the known association with the respective decoder side neural network is smallest among a plurality of distance metrics that are computed.(Meyerson [¶0035] "Underlying each of these approaches is the assumption of parallel ordering of shared layers. Also, each one requires aligned sequences of feature extractors across tasks." [¶0045] "A common interpretation of deep learning is that layers extract progressively higher level features at later depths" [¶0016] "the feature hierarchy implied by the model architecture." [¶0035] "adapt each layer with a small number of task-specific scaling parameters" [¶0066] "Soft ordering (Eq. 7) generalizes Eqs. 2 and 3 by learning a tensor S of task-specific scaling parameters" [¶0073] "For a representative two-task soft order experiment the layer-wise distance between scalings of the tasks increases by iteration, as shown by (b) in FIG. 5" Meyerson explicitly computes a distance metric between the extracted layer-wise, task-specific, scaling features where the distance is explicitly smallest in the first iteration). Regarding claim 6, the combination of Meyerson, and Hoogeboom teaches The apparatus claim 1, wherein the apparatus is further caused to: decode a bitstream, received from an encoder side, by using a lossless or substantially lossless codec;(Hoogeboom [p. 5 §3.3] "Lossless compression is an essential technique to limit the size of representations without destroying information. Methods for lossless compression require i) a statistical model of the source, and ii) a mapping from source symbols to bit stream [...] The mapping between symbols and bit streams may be provided by any entropy encoder") provide the decoded bitstream to one or more of the plurality of decoder side neural networks, based on the one or more tasks; run one or more decoder side neural networks associated with at least one task of the one or more tasks; and(Hoogeboom [p. 2] "Figure 1: Overview of IDF based lossless compression. An image x is transformed to a latent representation z with a tractable distribution pZ(·). An entropy encoder takes z and pZ(·) as input, and produces a bitstream c. To obtain x, the decoder uses pZ(·) and c to reconstruct z. Subsequently, z is mapped to x using the inverse of the IDF." [p. 5] "the mapping f : x 7! z is defined by the IDF. Subsequently, z is encoded under the distribution pZ(z) to a bitstream c using an entropy encoder. Note that, when using factor-out layers, pZ(z) is also defined using the IDF. Finally, in order to decode a bitstream c, an entropy encoder uses pZ(z) to obtain z. and the original image is obtained by using the map f") provide an output of the one or more decoder side neural networks or data derived from the output of the one or more decoder side neural networks to a task neural network associated with the task.(Hoogeboom [p. 5] "Figure 4: Example of a 2-level flow architecture. The squeeze layer reduces the spatial dimensions by two, and increases the number of channels by four. A single integer flow layer consists of a channel permutation and an integer discrete coupling layer. Each level consists of D flow layers" See FIG. 4 which shows that inner decoder neural networks are inversely processed by providing decoder output to outer decoder layer(s).). Regarding claim 7, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 1, wherein the apparatus is further caused to associate two or more decoder side neural networks with a task, (Meyerson [p. 1 §1] "when only a single task is available for training. The method is formalized as pseudo-task augmentation (PTA), in which a single task has multiple distinct decoders projecting the output of the shared structure to task predictions. By training the shared structure to solve the same problem in multiple ways, PTA simulates the effect of training towards distinct but closely-related tasks drawn from the same universe") and wherein the two or more decoder side neural networks perform differently in terms of a rate-distortion trade-off or a trade-off between a rate and a task accuracy, and (Meyerson [p. 3 §3.1] "since each decoder induces a distinct model for a task, what matters is not the average over decoders, but the best performing decoder for each task, i.e. [...] Equation 5 is used for model validation, and to select the best performing decoder for each task from the final joint model" Meyerson explicitly chooses the best decoder which requires the decoders to perform differently in terms of a rate-distortion (loss/error).) wherein the two or more decoder side neural networks are associated with a different level of at least one of a computation complexity, a memory complexity, or a power complexity.(Meyerson [p. 1 §1] "pseudo-task augmentation is a broadly applicable and efficient way to boost performance in deep learning systems" Boosting performance interpreted as synonymous with being associated with a computational complexity.). Regarding claim 8, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 2, wherein a task category of the one or more task categories comprises one or more tasks having characteristics that are same or substantially same.(Meyerson [¶0185] "selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks and transmitting the accumulated output encoding produced for the final clone as input to the selected decoder."). Regarding claim 9, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 1, wherein, the apparatus is further caused to: select one or more task neural networks based on one or more predetermined neural network tasks;(Meyerson [¶0185] "selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks and transmitting the accumulated output encoding produced for the final clone as input to the selected decoder." See also FIG. 1. Model 101 with selected decoder side neural network and encoder 102 is interpreted as a selected task neural network based on one or more predetermined neural network tasks.) select one or more decoder side neural networks associated with the one or more tasks; and use the selected one or more decoder side neural networks to decode or process associated input data.(Meyerson [¶0185] "selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks and transmitting the accumulated output encoding produced for the final clone as input to the selected decoder." Decoder interpreted as decoding input data by definition.). Regarding claim 11, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 1, wherein, to select the decoder side neural network, the apparatus is further caused to evaluate a performance of the new task neural network when applied on one or more sample data, (Meyerson [¶0053] "In Eq. 3 above, σi is a task-specific permutation of size D, and σi is fixed before training. If there are sets of tasks for which joint training of the model defined by Eq. 3 achieves similar or improved performance over Eq. 2, then parallel ordering is not a necessary requirement for deep MTL. Of course, in this formulation, it is required that the wk can be applied in any order. See Section 6 for examples of possible generalizations."). Regarding claim 12, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 11, wherein to evaluate the performance of the new task neural network, the apparatus is further caused to measure a task performance on a low-quality data by using an approximated ground truth from high quality data, and wherein to measure the task performance, the apparatus is further caused to: derive the approximated ground truth from an output of the new task neural network, when an input data is the high quality data;(Meyerson [p. 3 §3] “Consider again the case where there are T distinct true tasks, but now let there be D decoders for each task. Then, the model for the dth decoder of the tth task is given by [See Eqn. 3] and the overall loss for the joint model from Eq. 2 becomes [See Eqn. 4]"” [p. 6 §4.1] “This section evaluates and compares the various PTA methods on Omniglot character recognition (Lake et al., 2015). The Omniglot dataset consists of 50 alphabets of handwritten characters, each of which induces its own character recognition task”) determine one or more candidate decoder side neural networks from the plurality of decoder side neural networks;(Meyerson [0185] "The method includes selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks" [¶0073] "For a representative two-task soft order experiment the layer-wise distance between scalings of the tasks increases by iteration, as shown by (b) in FIG. 5") for each candidate decoder side neural network perform: providing the low-quality data as an input to the candidate decoder side neural network;(Meyerson [¶0105] "The training, validation, and test splits provided by Liu et al. (2015b) were used. There are ≈160K images for training, ≈20K for validation, and ≈20K for testing" validation split interpreted as low-quality data from the full dataset (high-quality data) where the validation split comprises the ground truth examples.) providing output of the candidate decoder side neural network as an input to the new task neural network; and(Meyerson See FIG. 1) comparing at least one output of one or more outputs of the new task neural network with the approximated ground truth; and(Meyerson [p. 3 §3] “Consider again the case where there are T distinct true tasks, but now let there be D decoders for each task. Then, the model for the dth decoder of the tth task is given by [See Eqn. 3] and the overall loss for the joint model from Eq. 2 becomes [See Eqn. 4]"” [p. 6 §4.1] “This section evaluates and compares the various PTA methods on Omniglot character recognition (Lake et al., 2015). The Omniglot dataset consists of 50 alphabets of handwritten characters, each of which induces its own character recognition task”) wherein comparing the at least one output of the one or more outputs of the new task neural network with the approximated ground truth comprises determining an accuracy value associated with the candidate decoder side neural network; (Meyerson [¶0057] "The following experiments investigate how accurately a model can jointly fit two tasks of n samples" [¶0059] "in the linear case, permuted ordering of shared layers does not lose accuracy compared to the single-task case") and select a candidate decoder side neural network of the one or more candidate decoder side neural networks, based on the accuracy value associated with the candidate decoder side neural network providing a predetermined accuracy value as the decoder side neural network.(Meyerson [0185] "The method includes selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks" [¶0073] "For a representative two-task soft order experiment the layer-wise distance between scalings of the tasks increases by iteration, as shown by (b) in FIG. 5"). Regarding claim 13, the combination of Meyerson, and Hoogeboom teaches The apparatus of claim 11, wherein to evaluate the performance of the new task neural network, the apparatus is further caused to measure a distance between features extracted from high quality data and reconstructed data, and wherein to measure the distance, the apparatus is further caused to: derive the approximated ground truth from an output of a feature-extraction neural network when an input is the high-quality data;(Meyerson [p. 3 §3] “Consider again the case where there are T distinct true tasks, but now let there be D decoders for each task. Then, the model for the dth decoder of the tth task is given by [See Eqn. 3] and the overall loss for the joint model from Eq. 2 becomes [See Eqn. 4]"” [p. 6 §4.1] “This section evaluates and compares the various PTA methods on Omniglot character recognition (Lake et al., 2015). The Omniglot dataset consists of 50 alphabets of handwritten characters, each of which induces its own character recognition task”) determine one or more candidate decoder side neural network from the plurality of decoder side neural networks;(Meyerson [0185] "The method includes selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks" [¶0073] "For a representative two-task soft order experiment the layer-wise distance between scalings of the tasks increases by iteration, as shown by (b) in FIG. 5") for each candidate decoder side neural network perform: providing the low-quality data as an input to the each candidate decoder side neural network; (Meyerson See FIG. 1) providing output of the candidate decoder side neural network as an input to the feature-extraction neural network;(Meyerson [p. 3 §3] “Consider again the case where there are T distinct true tasks, but now let there be D decoders for each task. Then, the model for the dth decoder of the tth task is given by [See Eqn. 3] and the overall loss for the joint model from Eq. 2 becomes [See Eqn. 4]"” [p. 6 §4.1] “This section evaluates and compares the various PTA methods on Omniglot character recognition (Lake et al., 2015). The Omniglot dataset consists of 50 alphabets of handwritten characters, each of which induces its own character recognition task”) extracting features based at least on the output of the each candidate decoder side neural network; and(Meyerson [¶0035] "Underlying each of these approaches is the assumption of parallel ordering of shared layers. Also, each one requires aligned sequences of feature extractors across tasks.") comparing the features extracted from the output of the each candidate decoder side neural network with the approximated ground-truth; and(Meyerson [p. 3 §3] “Consider again the case where there are T distinct true tasks, but now let there be D decoders for each task. Then, the model for the dth decoder of the tth task is given by [See Eqn. 3] and the overall loss for the joint model from Eq. 2 becomes [See Eqn. 4]"” [p. 6 §4.1] “This section evaluates and compares the various PTA methods on Omniglot character recognition (Lake et al., 2015). The Omniglot dataset consists of 50 alphabets of handwritten characters, each of which induces its own character recognition task”) select a candidate decoder side neural network providing a predetermined distance value or a lowest distance value as the decoder side neural network.(Meyerson [0185] "The method includes selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks" [¶0073] "For a representative two-task soft order experiment the layer-wise distance between scalings of the tasks increases by iteration, as shown by (b) in FIG. 5" [p. 3 §3.1] “Equation 5 is used for model validation, and to select the best performing decoder for each task from the final joint model. This decoder is then applied to future data”). Regarding claim 14, the combination of Meyerson, and Hoogeboom teaches The apparatus claim 1, wherein the one or more tasks comprises one or more task-NNs.(Meyerson [¶0185] "selecting, from among numerous decoders, a decoder that is specific to the particular one of the classification tasks and transmitting the accumulated output encoding produced for the final clone as input to the selected decoder."). Regarding claims 15-20, claims 15-20 are directed towards the method performed by the apparatus of claims 1-5, respectively. Therefore, the rejections applied to claims 1-5 also apply to claims 15-20. Regarding claim 21, the combination of Meyerson and Hoogeboom teaches The apparatus of claim 12, wherein the candidate decoder side neural network that is selected provides a predetermined accuracy value or an accuracy value that is highest among a plurality of accuracy values that are determined.(Meyerson [¶0057] "The following experiments investigate how accurately a model can jointly fit two tasks of n samples" [¶0059] "in the linear case, permuted ordering of shared layers does not lose accuracy compared to the single-task case"). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Han (“Learning Tree Structure in Multi-Task Learning”, 2015) which teaches ([p. 398] “we devise sequential constraints each of which makes the distance between the component parameters in the component matrices for a pair of tasks decrease over layers and as a consequence”). Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Dec 15, 2021
Application Filed
Oct 24, 2025
Non-Final Rejection mailed — §103, §112
Mar 20, 2026
Response Filed
Apr 29, 2026
Final Rejection mailed — §103, §112
Jul 28, 2026
Response after Non-Final Action
Aug 04, 2026
Request for Continued Examination
Aug 06, 2026
Response after Non-Final Action
Sep 01, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699874
Leveraging Redundancy in Attention with Reuse Transformers
3y 10m to grant Granted Aug 04, 2026
Patent 12675673
NEURAL NETWORK PROCESSING DEVICE, METHOD, AND COMPUTER-READABLE RECORDING MEDIUM
3y 7m to grant Granted Jul 07, 2026
Patent 12645914
INSTRUCTION PRUNING FOR NEURAL NETWORKS
3y 6m to grant Granted Jun 02, 2026
Patent 12626139
SECRET SOFTMAX FUNCTION CALCULATION SYSTEM, SECRET SOFTMAX FUNCTION CALCULATION APPARATUS, SECRET SOFTMAX FUNCTION CALCULATION METHOD, SECRET NEURAL NETWORK CALCULATION SYSTEM, SECRET NEURAL NETWORK LEARNING SYSTEM, AND PROGRAM
4y 3m to grant Granted May 12, 2026
Patent 12619815
Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
1y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
51%
Grant Probability
86%
With Interview (+35.1%)
4y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 152 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month