Prosecution Insights
Last updated: October 04, 2026
Application No. 17/756,461

METHOD AND SYSTEM FOR DETERMINING TASK COMPATIBILITY IN NEURAL NETWORKS

Final Rejection §103§112
Filed
May 25, 2022
Priority
Nov 25, 2019 — EU 19211218.3 +1 more
Examiner
PHAM, JESSICA THUY
Art Unit
2121
Tech Center
2100 — Computer Architecture & Software
Assignee
Continental Automotive GmbH
OA Round
4 (Final)
18%
Grant Probability
At Risk
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants only 18% of cases
18%
Career Allowance Rate
2 granted / 11 resolved
-36.8% vs TC avg
Strong +90% interview lift
Without
With
+90.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
21 currently pending
Career history
46
Total Applications
across all art units

Statute-Specific Performance

§101
28.5%
-11.5% vs TC avg
§103
38.5%
-1.5% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
20.4%
-19.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 11 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment/Status of Claims Claims 1, 2, 6, 7, and 9-14 were amended. Claim 8 was cancelled. Claim 16 is new. Claims 1-2 and 4-16 are pending and examined herein. Claims 1, 9, 10, and 14 are objected to. Claims 1-2, 4-7, and 9-16 are rejected under 35 U.S.C. 112(b). Claims 1-2, 4-7, and 9-16 are rejected under 35 U.S.C. 103. Response to Arguments Applicant’s arguments, see pages 13, filed 6/4/2026, with respect to the rejection(s) of claim(s) 1-2, 4-7, and 9-16 under 35 U.S.C. 112 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. Note that Applicant’s amendments introduced new issues, and new rejections under 35 U.S.C. 112(a) and 35 U.S.C. 112(b) are made. Applicant’s arguments, see pages 13, filed 6/4/2026, with respect to the rejection(s) of claim(s) 1-2 and 4-15 under 35 U.S.C. 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Vandenhende et al., “Branched Multi-Task Networks: Deciding What Layers to Share,” November 2, 2019,” hereinafter “Vandenhende,” Jianshu Li et al., “Task Relation Networks,” January 7, 2019, IEEE, 2019 IEEE Winter Conference on Applications of Computer Vision, pp. 932-940, hereinafter “Li”, and Ben Poole et al., “On Variational Bounds of Mutual Information,” May 16, 2019, arXiv, Proceedings of the 36th International Conference on Machine Learning, hereinafter “Poole”. Claim Objections Claims 1, 9, 10, and 14 are objected to because of the following informalities: Claims 1 and 14 state “the information share measures including a first measure indicating how much of the image information contains second encoded image information about the first encoded image information, and a second measure indicating how much of the image information contains first encoded image information about the second encoded image information.” For clarity, this limitation should be rewritten. For purposes of examination, this limitation will be interpreted as “the information share measures including a first measure indicating an amount of information about the first encoded information that is contained in the second encoded information and a second measure indicating a second amount of information about the second encoded information that is contained in the first encoded information." Claims 1 and 14 state “estimating, by the auxiliary neural network, information share measures for each of the first encoded information and the second encoded information generated by the trained first and second sub-networks according to the approximation of the probability function.” Though it is clear from context that the estimation of the information share measures is according to the approximation of the probability function, the syntax of this sentence is ambiguous. It reads that the first encoded information and the second encoded information are generated by the trained first and second sub-networks according to the approximation of the probability function, which is not the intended meaning of the limitation. For clarity, this limitation should be rewritten. Claims 1 and 14 recite the term “the image information” in the paragraph beginning with “receiv[ing]”. It is clear from the context that this should be “the input image information”. Claim 9 recites "estimating an information share measure." This should be “estimating the information share measures.” Claim 10 recites "the information share measures a first individual task is compared against other individual tasks." This should likely be “wherein the information share measure for a first individual task is compared against other individual tasks.” Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-2, 4-7, and 9-16 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claims 1 and 14 recites the limitation "the image information" in the paragraph beginning with “estimat[ing]”. There is insufficient antecedent basis for this limitation in the claim. It is unclear what image information this refers to, the input image information or either the first or second encoded image information. Claims 1 and 14 recite the limitation "select[ing] from each of the trained first and second sub-networks multiple parameters that establish a parameterizable distribution describing a probability function of a first set of information occurring as a result of a second information." However, the specification does not support this limitation. [0070] states "Broadly, the auxiliary neural network AUXNN is used to determine parameters of a parametrized probability function Q(AIB), which defines the probability event A occurring given that event B has occurred. Based on the parametrized probability function Q(AIB), which parameters can be evaluated by training the auxiliary neural network AUXNN, it is possible to approximate an upper bound of information content missing from second encoded image information B compared to first encoded image information A, that is conditional entropy H(AIB)." Therefore, the parameters for the parametrized probability function are not selected from the trained first and second sub-networks, rather determined by the parameters of the auxiliary neural network. Also, "a probability function of a first set of information occurring as a result of a second information" is not the same as "the probability event A occurring given that event B has occurred". This limitation is new matter. For purposes of examination, this limitation will be interpreted as “establishing a parameterizable distribution describing a conditional probability function”. Claims 1 and 14 recite the limitation “training the auxiliary network on the parameterizable distribution to obtain an approximation of the probability function.” [0070] states "Based on the parametrized probability function Q(AIB), which parameters can be evaluated by training the auxiliary neural network AUXNN, it is possible to approximate an upper bound of information content missing from second encoded image information B compared to first encoded image information A, that is conditional entropy H(AIB)." This does not mean that the parameterizable distribution is used to train the auxiliary network, rather that the auxiliary network is used to determine the parameterizable distribution, which is an approximation of the probability function. For purposes of examination, this limitation will be interpreted as “training the auxiliary network to obtain an approximation of the probability function.” Dependent claims 2, 4-7, 9-13, and 15-16 fail to resolve the issue and are rejected with the same rationale. Claim 2 recites the limitations “estimating the first measure includes approximating an upper bound of information missing in the second encoded image information compared to an amount of the input image information included in the first encoded image information" and “estimating the first measure includes approximating an lower bound of information missing in the second encoded image information compared to an amount of the input image information included in the first encoded image information.” The specification does not provide support for the first measure, which indicates how much information the second encoded image information contains about the first image information, being an approximation of an upper bound. The specification does not provide support for the second measure, which indicates how much information the first encoded image information contains about the second image information, being an approximation of a lower bound. For purposes of examination, these limitations will be interpreted as “approximating an upper bound of information missing in the second encoded image information compared to an amount of the input image information included in the first encoded image information" and “approximating an lower bound of information missing in the second encoded image information compared to an amount of the input image information included in the first encoded image information,” respectively. Claims 7 and 16 recite "wherein estimating an information share measure based on the auxiliary network comprises selecting the parameterizable distribution among multiple parameterizable distributions, the selected parameterizable distribution providing the probability function as a parameterizable probability distribution function used for determining how much information content exists in the second encoded image which is not covered in the first encoded image." There is insufficient support for this limitation in the specification. While [0072] states that "In a further step, a parametrizable distribution is selected which comprise includes multiple parameters to be chosen and which is suitable to describe a probability distribution," the specification does not provide support for the multiple parameterizable distributions. For purposes of examination, this limitation will be interpreted as "wherein estimating an information share measure based on the auxiliary network comprises selecting the parameterizable distribution, the selected parameterizable distribution providing the probability function as a parameterizable probability distribution function used for determining how much information content exists in the second encoded image which is not covered in the first encoded image." The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-2, 4-7, and 9-16 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 1 and 14 recite the limitation "select[ing] from each of the trained first and second sub-networks multiple parameters that establish a parameterizable distribution describing a probability function of a first set of information occurring as a result of a second information." This is inconsistent with the specification, which states in [0070] that "Broadly, the auxiliary neural network AUXNN is used to determine parameters of a parametrized probability function Q(AIB), which defines the probability event A occurring given that event B has occurred. Based on the parametrized probability function Q(AIB), which parameters can be evaluated by training the auxiliary neural network AUXNN, it is possible to approximate an upper bound of information content missing from second encoded image information B compared to first encoded image information A, that is conditional entropy H(AIB)." Therefore, it is uncertain where the parameters of the probability function come from and what the parametrized probability function describes. The claims are rendered indefinite. For purposes of examination, this limitation will be interpreted as “establishing a parameterizable distribution describing a conditional probability function”. Claims 1 and 14 recite the limitation “training the auxiliary network on the parameterizable distribution to obtain an approximation of the probability function.” This is inconsistent with the specification, in which [0070] states "Based on the parametrized probability function Q(AIB), which parameters can be evaluated by training the auxiliary neural network AUXNN, it is possible to approximate an upper bound of information content missing from second encoded image information B compared to first encoded image information A, that is conditional entropy H(AIB)." This does not mean that the parameterizable distribution is used to train the auxiliary network, rather that the auxiliary network is used to determine the parameterizable distribution, which is an approximation of the probability function. As it is unclear whether the parameterizable distribution is used to train the auxiliary network, the claim is rendered indefinite. For purposes of examination, this limitation will be interpreted as “training the auxiliary network to obtain an approximation of the probability function.” Dependent claims 2, 4-7, 9-13, and 15-16 fail to resolve the issue and are rejected with the same rationale. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 1-2, 5, 9-11, and 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vandenhende et al., “Branched Multi-Task Networks: Deciding What Layers to Share”, November 2, 2019”, hereinafter “Vandenhende”, Jianshu Li et al., “Task Relation Networks”, January 7, 2019, IEEE, 2019 IEEE Winter Conference on Applications of Computer Vision, pp. 932-940, hereinafter “Li”, and Ben Poole et al., “On Variational Bounds of Mutual Information”, May 16, 2019, arXiv, Proceedings of the 36th International Conference on Machine Learning, hereinafter “Poole”. Regarding claim 1, Vandenhende teaches A computer-implemented method for determining clusters of tasks, the clusters at least partially including multiple tasks to be executed by a computer processor configured with at least a joint encoder portion of a neural network, the method comprising (Section 1, page 2 states "The proposed method aims to find an effective task grouping for the sharable layers f l of the encoder, i.e. grouping related tasks together in the same branches of the tree." The grouping is interpreted as clustering, and the sharable layers of the encoder are interpreted as the joint encoder portion of a neural network. Algorithm 1 on page 4 shows the method, which one of ordinary skill in the art would realize is implemented on a computer which necessarily uses a processor to perform the method.) training a neural network for performing processing tasks, the neural network having at least a first sub-network and a second sub-network; (Page 3 states "Consider a backbone architecture: an encoder, consisting of a sequence of shared layers or blocks f l , followed by a decoder with a few task-specific layers. We assume an appropriate structure for layer sharing to take the shape of a tree. In particular, the first layers are shared by all tasks, while later layers gradually split off as they show more task-specific behavior." Section 3.1, page 4 states “As a first step, we train a single-task model for each task t i ∈ T . The single-task models use an identical encoder E -made of all sharable layers f l -followed by a task-specific decoder D t i ." As the models are neural networks that are composed of layers from a branched neural network architecture (including the identical encoder and the task-specific decoders), the single-task models are sub-networks under the broadest reasonable interpretation. Figure 1(a) shows that there are at least four tasks, which means there is a first and second task and a corresponding first and second neural network. Training the sub-networks is interpreted as training a neural network.) forming an estimation neural network, the estimation neural network comprising the trained first sub-network, the trained second sub-network, and an auxiliary [function]; (Section 3.1 states “To calculate these task affinities, we have to compare the representation dissimilarity matrices (RDM) of the single-task networks – trained in the previous step – at the specified D locations.” Section 3.1 further states “Specifically, RDM d ,   i ,   j is found by calculating the dissimilarity score between the features at location d for image i and j .” Section 3.1 further states “For a specific location d in the network, the computed RDMs are symmetrical, with a diagonal of zeros. For every such location, we measure the similarity between the upper or lower triangular part of the RDMs belonging to the different single-task networks. We use the Spearman’s correlation coefficient r s to measure similarity. When repeated for every pair of tasks, at a specific location d , the result is a symmetrical matrix of size N × N , with a diagonal of ones. Concatenating over the D locations in the sharable encoder, we end up with the desired task affinity tensor of size D × N ×   N .” The function that obtains the task affinity is interpreted as the auxiliary function. It receives the features from images processed through the trained neural networks, and therefore receives information from the trained first and second neural networks.) receiving input image information at the estimation neural network, the trained first sub-network generating first encoded image information from the input Image information and trained second sub-network generating second encoded image information from the input image information; (Section 3.1 states “The single-task models use an identical encoder E – made of all sharable layers f l – followed by a task-specific decoder D t i .” Section 3.1 further states “To do this, a held-out subset of K images is required. The latter images serve to compare the dissimilarity of their feature representations in the single-task networks for every pair of images. Specifically, for every task t i , we characterize these learned feature representations at the selected locations by filling a tensor of size D × K × K .” As the images are used to input to the trained single-task networks, and the single-task networks comprise an encoder and a decoder, it is inherent that each neural network will generate encoded image information.) estimating, by the auxiliary [function], information share measures for each of the first encoded information and the second encoded information generated by the trained first and second sub-networks . . . , the information share measures including a first measure indicating how much of the image information contains second encoded image information about the first encoded image information, and a second measure indicating how much of the image information contains first encoded image information about the second encoded image information; (Section 1 states “To this end, we base the layer sharing on measurable levels of task affinity or task relatedness: two tasks are strongly related, if their single task models rely on a similar set of features.” Features are obtained from the encoder-decoder single-task models, and, as established above, the single-task models provide encoded image information. Section 1 further states “Given a dataset and a number of tasks, our approach uses RSA [representation similarity analysis] to assess the task affinity at arbitrary locations in a neural network.” As one of ordinary skill in the art would understand, similarity between representations/features (interpreted as encoded information) measures how much information is shared between the encoded information that is being compared. Section 3.1 states “Specifically, RDM d ,   i ,   j is found by calculating the dissimilarity score between the features at location d for image i and j .” At location d , images i and j are encoded. Therefore, as discussed above, the task affinity, calculated using the RDM, is interpreted as the information share measure, and the encoded images are compared using the information share measure. Section 3.1, page 5 states “When repeated for every pair of tasks, at a specific location d , the result is a symmetrical matrix of size N × N , with a diagonal of ones.” As the matrix is symmetrical, one of ordinary skill in the art would reason that the tasks are compared both ways, meaning that there are two information share measures.) [having a measure for the] first and second encoded image information which measures a difficulty for the first processing task to be executed together with the second processing task in a shared multitask architecture (Page 2 states "Given a dataset and a number of tasks, our approach uses RSA to assess the task affinity at arbitrary locations in a neural network. The task affinity scores are then used to construct a branched multi task network in a fully automated manner. In particular, our task clustering algorithm groups similar tasks together in common branches, and separates dissimilar tasks by assigning them to different branches, thereby reducing the negative transfer between tasks." The negative transfer is interpreted as the difficulty for the first and second processing task to be executed together in the shared multitask architecture. As RSA reduces this, it is a measure of the difficulty for task grouping.) comparing the first measure and the second measure to a threshold value indicating a limit for an information overlap specifying which of the processing should be executed in a joint encoder potion of the neural network; (Section 3.2, page 5 states “The task dissimilarity score of a tree is defined as C c l u s t e r = ∑ l C c l u s t e r l , where C c l u s t e r l is found by averaging the maximum distance between the dissimilarity scores of the elements in every cluster.” Section 3.2 further states “The branched multi-task network is built with the intention to separate dissimilarity tasks by assigning them to separate branches. To this end, we define the dissimilarity score between two tasks t i and t j at location d as 1 - A d ,   i ,   j , with A the task affinity tensor.” By using the maximum distance between tasks in a cluster, there exists a functional threshold of dissimilarity, beyond which tasks will increase the cluster cost and be excluded from the same group. Tasks below this threshold are clustered together. As dissimilarity involves the task affinity, which measures information overlap between representations, the threshold on dissimilarity also indicates information overlap. The calculation of the dissimilarity score is interpreted as the comparison.) determining clusters of the processing tasks to be executed in a joint encoder portion of the neural network based on the comparison, and (The claim mapping of step g) explains how the clusters are determined based on the comparison. Additionally, Section 3.2 states “By taking into account the clustering cost at all depths, the procedure can find a task grouping that is considered optimal in a global sense.” The task grouping is interpreted as the clusters of tasks.) allocating computational resources to perform the clusters of processing tasks in the joint encoder portion of the neural network based on the determination. (Page 5 states "Given a computational budget C , we need to derive how the layers (or blocks) in the sharable f l encoder should be shared among the tasks in T ." Page 5 further states "Since the number of tasks is finite, we can enumerate all possible trees that fall within the given computational budget C ." Page 6 states "Depending on the available computational budget C , our method generates a specific task grouping." As the available computational budget is applied to the task grouping, the resources (computational budget) are allocated based on the determination of the task grouping.) Vandenhende does not appear to explicitly teach selecting from each of the trained first and second sub-networks multiple parameters that establish a parameterizable distribution describing a probability function of a first set of information occurring as a result of a second information; [forming an estimation neural network, the estimation neural network comprising the trained first sub-network, the trained second sub-network, and an auxiliary] neural network, wherein the auxiliary network is arranged to receive a trained first sub-network output and a trained second sub-network output; training the auxiliary network on the parametrizable distribution to obtain an approximation of the probability function; [estimating], by the auxiliary neural network, [an information share measure] according to the approximation of the probability function However, Li—directed to analogous art—teaches forming an estimation neural network, the estimation neural network comprising the trained first sub-network, the trained second sub-network, and an auxiliary neural network, wherein the auxiliary network is arranged to receive a trained first sub-network output and a trained second sub-network output; (Fig. 1 shows the estimation neural network. The TRL layers are interpreted as the auxiliary neural network. Page 935 states "We apply TRL between the corresponding layers in the task-specific subnets. Specifically, we can place a TRL between the output layers in the task-specific subnets to form an output TRL. Taking as input the features from output layers, the output TRL generates STR.") It would have been obvious to a person having ordinary skill in the art before the effective filing date of this application to combine the teachings of Vandenhende and Li because, as stated by Li on page 933, "Traditionally, multi-task learning schemes can be categorized into multi-task feature learning and multi-task relation learning. The first category mainly focuses on learning features shared for different tasks, such that it enables implicit data augmentation and mitigates representation bias [20]. The second category usually uses task covariance [2, 24] to model the relationship between tasks to achieve mutual performance boosts. Different from them, our proposed TRN explicitly models the relations between the features, so that more favorable features can be learned for improving upon multiple tasks." The combination of Vandenhende and Li does not appear to explicitly teach selecting from each of the trained first and second sub-networks multiple parameters that establish a parameterizable distribution describing a probability function of a first set of information occurring as a result of a second information; However, Poole—directed to analogous art—teaches selecting from each of the trained first and second sub-networks multiple parameters that establish a parameterizable distribution describing a probability function of a first set of information occurring as a result of a second information; (Page 8 states "We use the convolutional encoder architecture from Burgess et al. (2018); Locatello et al. (2018) for p(y|x)." Therefore, the parameters of the convolutional encoder are used to estimate p(y|x), which is the conditional probability of y under the condition x.) training the auxiliary network on the parametrizable distribution to obtain an approximation of the probability function; (Page 8 states "We use the convolutional encoder architecture from Burgess et al. (2018); Locatello et al. (2018) for p(y|x)." Therefore, the parameters of the convolutional encoder are used to estimate p(y|x), which is the conditional probability of y under the condition x. One of ordinary skill in the art would reason that the convolutional encoder is trained to estimate the conditional probability, as it is involved in the optimization of eq. 15.) [estimating], by the auxiliary neural network, [an information share measure] according to the approximation of the probability function (Pages 7-8 state "To estimate and maximize the information contained in the representation Y about the input X, we use the I J S lower bound, with a structured critic that leverages the known stochastic encoder p(y|x) but learns an unnormalized variational approximation q(y) to the prior.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of this application to combine the teachings of Vandenhende and Li with the teachings of Poole because, as Poole states on page 1, "Estimating the relationship between pairs of variables is a fundamental problem in science and engineering. Quantifying the degree of the relationship requires a metric that captures a notion of dependency." Additionally, according to Yixuan Li et al., "Convergent Learning: Do Different Neural Networks Learn the Same Representations?", February 28, 2016, ICLR 2016, "Because correlation is a relatively simple mathematical metric that may miss some forms of statistical dependence, we also performed one-to-one alignments of neurons by measuring the mutual information between them. Mutual information measures how much knowledge one gains about one variable by knowing the value of another." Regarding claim 2, the rejection of claim 1 is incorporated herein. Vandenhende teaches second encoded image information (See below.) the input image information included in the first encoded image information (Section 3.1 states “The single-task models use an identical encoder E – made of all sharable layers f l – followed by a task-specific decoder D t i .” Section 3.1 further states “To do this, a held-out subset of K images is required. The latter images serve to compare the dissimilarity of their feature representations in the single-task networks for every pair of images. Specifically, for every task t i , we characterize these learned feature representations at the selected locations by filling a tensor of size D × K × K .” As the images are used to input to the trained single-task networks, and the single-task networks comprise an encoder and a decoder, it is inherent that each neural network will provide encoded image information.) The combination of Vandenhende and Li does not appear to explicitly teach estimating an information share measure based on the auxiliary neural network includes approximating an upper bound of the input information missing in [data] compared to an amount of [other data]. estimating an information share measure based on the auxiliary neural network includes approximating a lower bound of information included in [data] compared to the amount of [other data]. However, Poole—directed to analogous art—teaches estimating an information share measure based on the auxiliary neural network includes approximating an upper bound of information missing in [data] compared to the amount of [other data]. ("Upper bounding MI is challenging, but is possible when the conditional distribution p ( y | x ) is known (e.g. in deep representation learning where y   is the stochastic representation). We can build a tractable variational upper bound by introducing a variational approximation q ( y ) to the intractable marginal p y = ∫ d x   p ( x ) p ( y | x ) . By multiplying and dividing the integrand in MI by q ( y ) and dropping a negative KL term, we get a tractable variational upper bound". Deep representation learning includes a neural network. Mutual information is again interpreted as the information share measure. In this method, x and y are the data being compared.) estimating the first measure based on the auxiliary neural network includes approximating a lower bound of information included in [data] compared to the amount of [other data]. (Page 7 states "To estimate and maximize the information contained in the representation Y about the input X, we use the I J S lower bound." In this method, x and y are the data being compared.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of this application to combine the teachings of Vandenhende and Li with the teachings of Poole for the reasons given above in regards to claim 1. Regarding claim 5, the rejection of claim 1 is incorporated herein. The combination of Vandenhende and Li does not appear to explicitly teach wherein estimating an information share measure is performed at least a random probability function However, Poole—directed to analogous art—teaches wherein estimating an information share measure is performed based on a random probability function (Section 5 states "Upper bounding MI is challenging, but is possible when the conditional distribution p ( y | x ) is known (e.g. in deep representation learning where y   is the stochastic representation). We can build a tractable variational upper bound by introducing a variational approximation q ( y ) to the intractable marginal p y = ∫ d x   p ( x ) p ( y | x ) . By multiplying and dividing the integrand in MI by q ( y ) and dropping a negative KL term, we get a tractable variational upper bound". Mutual information is again interpreted as the information share measure. In this method, x and y are the data being compared. q ( y ) is interpreted as the random probability function.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of this application to combine the teachings of Vandenhende and Li with the teachings of Poole for the reasons given above in regards to claim 1. Regarding claim 9, the rejection of claim 1 is incorporated herein. Vandenhende teaches wherein the input image information includes multi-dimensional image information, and wherein estimating an information share measure comprises calculating information share measures for multiple different data points of the multi-dimensional image information and calculating a mean information share measure by averaging at least the first and second measures. (Page 4 shows input images, which include multiple dimensions of data, including pixel color, both in the dimensions of height and width.Section 3.2, page 5 states “The task dissimilarity score of a tree is defined as C c l u s t e r = ∑ l C c l u s t e r l , where C c l u s t e r l is found by averaging the maximum distance between the dissimilarity scores of the elements in every cluster.” Section 3.2, page 5 further states “To this end, we define the dissimilarity score between two tasks t i and t j at location d as 1 - A d , i , j , with A the task affinity tensor.” Therefore, by averaging the distance between the dissimilarity scores, an average of task affinity, interpreted as the information share measure, is found.) Regarding claim 10, the rejection of claim 1 is incorporated herein. Vandenhende teaches wherein information share measures for different tuples of tasks of the processing tasks are calculated for individual tasks in the different tuples of tasks, the information share measures a first individual task is compared against other individual tasks. (Section 3.1, page 5 states “When repeated for every pair of tasks, at a specific location d , the result is a symmetrical matrix of size N × N , with a diagonal of ones.” As the matrix is symmetrical, one of ordinary skill in the art could reason that the tasks are compared both ways. Section 1 states “Given a dataset and a number of tasks, our approach uses RSA [representation similarity analysis] to assess the task affinity at arbitrary locations in a neural network.” As one of ordinary skill in the art would understand, similarity between representations/features (interpreted as encoded information) measures how much information is shared between the encoded information that is being compared. Section 1 further states “To this end, we base the layer sharing on measurable levels of task affinity or task relatedness: two tasks are strongly related, if their single task models rely on a similar set of features.” This means that the information in the representations are indicative of the tasks used to encode the representation. Page 5 states "To this end, we define the dissimilarity score between two tasks t i and t j at location d as 1 - A d , i , j with A the task affinity tensor." Therefore, the similarity measures between each tasks are compared to each other.) Regarding claim 11, the rejection of claim 10 is incorporated herein. Vandenhende teaches wherein determining clusters of processing tasks comprises grouping individual tasks together to be executed in the joint encoder portion of the neural network if the first and second information share measure is below the threshold value. (Section 3.2, page 5 states “The task dissimilarity score of a tree is defined as C c l u s t e r = ∑ l C c l u s t e r l , where C c l u s t e r l is found by averaging the maximum distance between the dissimilarity scores of the elements in every cluster.” Section 3.2 further states “The branched multi-task network is built with the intention to separate dissimilarity tasks by assigning them to separate branches. To this end, we define the dissimilarity score between two tasks t i and t j at location d as 1 - A d ,   i ,   j , with A the task affinity tensor.” By using the maximum distance between tasks in a cluster, there exists a functional threshold of dissimilarity, beyond which tasks will increase the cluster cost and be excluded from the same group. Tasks below this threshold are clustered together. As dissimilarity involves the task affinity, which measures information overlap between representations, the threshold on dissimilarity also indicates information overlap. As stated above in regards to claim 10, the first and second information share measures are identical, as the task affinity tensor is symmetrical, meaning that comparing one to a threshold would be the same as comparing both to a threshold.) Regarding claim 14, Vandenhende teaches A system for determining clusters of tasks, the clusters at least partially including processing tasks to be executed in a joint encoder portion of a neural network, the system comprusing: (Section 1, page 2 states "The proposed method aims to find an effective task grouping for the sharable layers f l of the encoder, i.e. grouping related tasks together in the same branches of the tree." The grouping is interpreted as clustering, and the sharable layers of the encoder are interpreted as the joint encoder portion of a neural network. Algorithm 1 on page 4 shows the method, which one of ordinary skill in the art would realize is implemented on a computer. Page 5 states that “In this section, we quantitatively and qualitatively evaluate the proposed method on a number of diverse multi-tasking datasets, that range from real to semi-real data, from few to many tasks, from dense prediction to classification tasks, and so on.” This means that a system was used to execute the method.) a processor configured to execute instructions which cause the processor to be configured to: (As the method is implemented on a computer, the computer necessarily has a processor that executes the instructions to perform the method.) The remainder of claim 14 recites substantially similar subject matter to claim 1 and is rejected with the same rationale, mutatis mutandis. Regarding claim 15, the rejection of claim 5 is incorporated within. The combination of Vandenhende and Li do not appear to teach wherein the random probability function is determined by performing a variational mutual information maximization approach. However, Poole—directed to analogous art—teaches wherein the random probability function is determined by performing a variational mutual information maximization approach. (Page 8 states “We use the convolutional encoder architecture from Burgess et al. (2018); Locatello et al. (2018) for p ( y | x ) , and a two hidden layer fully-connected neural network to parameterize the unnormalized variational marginal q y used by I J S . Empirically, we find that this variational regularized info-max objective is able to learn x and y position, and scale, but not rotation”. q y is again interpreted as the random probability function) It would have been obvious to a person having ordinary skill in the art before the effective filing date of this application to combine the teachings of Vandenhende and Li with the teachings of Poole for the reasons given above in regards to claim 1. Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vandenhende, Li, and Poole as applied to claim 1 above, further in view of R Devon Hjelm et al., “Learning Deep Representations by Mutual Information Estimation and Maximization”, February 22, 2019, arXiv, 2019, hereinafter, “Hjelm.” Regarding claim 4, the rejection of claim 1 is incorporated herein. Vandenhende teaches the trained first neural network [and] the trained second neural network (Section 3.1, page 4 states “As a first step, we train a single-task model for each task t i ∈ T .” Figure 1(a) shows that there are at least four tasks, which means there is a first and second task and a corresponding first and second neural network.) The combination of Vandenhende, Li, and Poole does not appear to explicitly teach wherein estimating an information share measure includes training the auxiliary neural network by adapting the weights of the auxiliary neural network and keeping weights of [a trained neural network] constant. However, Hjelm—directed to analogous art—teaches wherein estimating an information share measure includes training the auxiliary neural network by adapting the weights of the auxiliary neural network and keeping weights of [a trained neural network] constant. (Section 4.1 states, "To summarize, we use the following metrics for evaluating representations. For each of these, the encoder is held fixed unless noted otherwise: " and further states, as one of the options, “Mutual information neural estimate (MINE), I ρ ^ ( X ,   E ψ x ) , between the input, X , and the output representation, E ψ ( x ) , by training a discriminator with parameters ρ to maximize the DV estimator of the KL-divergence.” MINE is interpreted as the auxiliary neural network that estimates mutual information, interpreted as the information share measure, and the encoder is interpreted as the trained neural network. As one of ordinary skill in the art would understand, holding the encoder fixed means keeping the weights of the encoder constant. Training a discriminator with parameters, as one of ordinary skill in the art would understand, is adapting the weights of the discriminator.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Vandenhende, Li, and Poole with the teachings of Hjelm because, as stated by Hjelm, "Evaluation of representations is case-driven and relies on various proxies. Linear separability is commonly used as a proxy for disentanglement and mutual information (MI) between representations and class labels. Unfortunately, this will not show whether the representation has high MI with the class labels when the representation is not disentangled." In the instant application’s claimed invention, the representations obtained by the trained neural networks are not disentangled, and would therefore require another method of evaluation as taught by Helm. Therefore, one of ordinary skill in the art would be motivated to add this element. Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vandenhende, Li, and Poole as applied to claim 1 above, further in view of Deng (“Multi-Task Learning with Multi-View Attention for Answer Selection and Knowledge Base Question Answering”, January 2019). Regarding claim 6, the rejection of claim 1 is incorporated herein. Vandenhende teaches wherein training the first sub-network comprises training an encoder of the first sub-network to perform a first processing task of the processing tasks, and (Section 3.1 states “The single-task models use an identical encoder E – made of all sharable layers f l – followed by a task-specific decoder D t i .” Section 3.1 further states “To do this, a held-out subset of K images is required. The latter images serve to compare the dissimilarity of their feature representations in the single-task networks for every pair of images. Specifically, for every task t i , we characterize these learned feature representations at the selected locations by filling a tensor of size D × K × K .” As the images are used to input to the trained single-task networks, and the single-task networks comprise an encoder and a decoder, it is inherent that each neural network will generate encoded image information. Page 4 states “We train a single-task model for every task t in T .) The combination of Vandenhende, Rusu, Belghazi, and Boudiaf does not appear to explicitly teach training the second sub-network comprises training an encoder of the second sub network for performing a second processing task of the processing tasks. However, Deng—directed to analogous art—teaches training a second neural network comprises training an encoder of the second neural network for performing the second processing task. (Page 6319, “Task-specific Encoder Layer”, states “Therefore, each task is equipped with a task-specific Siamese encoder for both question and answers, and each task-specific encoder contains a word encoder and a knowledge encoder to learn the integral sentence representations, as shown in Figure 2.” Page 6320, “Multi-Task Learning” states “The overall multi-task learning model is trained” which means that the task-specific encoders are also trained. As there is more than one task, a second encoder is trained for the second task.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Vandenhende, Rusu, Belghazi, and Boudiaf with the teachings of Deng because, as stated on page 6319, “Task-specific Encoder Layer”, "Different QA tasks are supposed to be diverse in data distributions and low-level representations. Therefore, each task is equipped with a task-specific siamese encoder for both questions and answers, and each task-specific encoder contains a word encoder and a knowledge encoder to learn the integral sentence representations, as shown in Figure 2." One would be motivated to combine because the tasks, as in Deng, may be diverse in data distributions and low-level representations. Claim(s) 7 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vandenhende, Li, and Poole as applied to claim 1 above, further in view of Chen (“InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets”, 2016). Regarding claim 7, the rejection of claim 1 is incorporated herein. Vandenhende teaches the first encoded image [and] the second encoded image (Section 3.1 states “The single-task models use an identical encoder E – made of all sharable layers f l – followed by a task-specific decoder D t i .” Section 3.1 further states “To do this, a held-out subset of K images is required. The latter images serve to compare the dissimilarity of their feature representations in the single-task networks for every pair of images. Specifically, for every task t i , we characterize these learned feature representations at the selected locations by filling a tensor of size D × K × K .” As the images are used to input to the trained single-task networks, and the single-task networks comprise an encoder and a decoder, it is inherent that each neural network will provide encoded image information.) The combination of Vandenhende, Li, and Poole does not appear to explicitly teach wherein estimating an information share measure based on the auxiliary neural network comprises selecting a parametrizable distribution among multiple parameterizable distributions, the selected parameterizable distribution providing the probability function used for determining how much information content exists in [the data] is not covered in [the other data]. However, Chen—directed to analogous art—teaches wherein estimating an information share measure based on the auxiliary neural network comprises selecting a parametrizable distribution among multiple parameterizable distributions, the selected parameterizable distribution providing the probability function used for determining how much information content exists in [the data] is not covered in [the other data]. (Page 3, section 5 states "In practice, the mutual information term   I ( c ; G ( z , c ) ) is hard to maximize directly as it requires access to the posterior P ( c | x ) . Fortunately we can obtain a lower bound of it by defining an auxiliary distribution Q ( c | x ) to approximate P ( c | x ) : I c ; G z , c = H c - H c G z , c In the equation, one can see that the part of the equation in the underbrace is equal to the conditional entropy H c G z , c . Wikipedia (“Conditional Entropy”) states that the conditional entropy is written as H Y X , in this case, the variable Y is instead c and X is G z , c . Further, Wikipedia, in the Venn diagram on page 1, where X is on the left in red and Y is on the right in blue, states that the conditional entropy is the part of the Venn diagram that is blue. This means that conditional entropy is the part of Y that is not covered by X. Therefore, this is inherent in Chen. Chen further states on page 4, section 5, that “Eq. (4) shows that the lower bound becomes tight as the auxiliary distribution Q approaches the true posterior distribution: E x [ D K L P ⋅ x     Q ⋅ x ) ] → 0 .” The auxiliary distribution Q therefore determines the conditional entropy when it is maximized. Section 6 states “In practice, we parameterize the auxiliary distribution Q as a neural network. In most experiments Q and D share all convolutional layers and there is one final fully connected layer to output parameters for the conditional distribution Q ( c | x ) , which means InfoGAN only adds a negligible computation cost to GAN.” The conditional distribution Q ( c | x ) is interpreted as the parametrizable probability distribution function. The conditional entropy is by definition is how much information content exists in [Y] is not covered in [X].) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Vandenhende, Li, and Poole with the teachings of Chen because as Chen states on page 4, We note in addition that the entropy of latent codes   H ( c ) can be optimized over as well since for common distributions it has a simple analytical form. However, in this paper we opt for simplicity by fixing the latent code distribution and we will treat H ( c ) as a constant. So far we have bypassed the problem of having to compute the posterior P ( c | x )   explicitly via this lower bound but we still need to be able to sample from the posterior in the inner expectation. Next we state a simple lemma, with its proof deferred to Appendix, that removes the need to sample from the posterior.” Regarding claim 16, Vandenhende teaches the first encoded image [and] the second encoded image (Section 3.1 states “The single-task models use an identical encoder E – made of all sharable layers f l – followed by a task-specific decoder D t i .” Section 3.1 further states “To do this, a held-out subset of K images is required. The latter images serve to compare the dissimilarity of their feature representations in the single-task networks for every pair of images. Specifically, for every task t i , we characterize these learned feature representations at the selected locations by filling a tensor of size D × K × K .” As the images are used to input to the trained single-task networks, and the single-task networks comprise an encoder and a decoder, it is inherent that each neural network will provide encoded image information.) The combination of Vandenhende, Li, and Poole does not appear to explicitly teach wherein estimating an information share measure based on the auxiliary neural network comprises selecting a parametrizable distribution among multiple parameterizable distributions, the selected parameterizable distribution providing the probability function used for determining how much information content exists in [the data] is not covered in [the other data]. However, Chen—directed to analogous art—teaches wherein estimating an information share measure based on the auxiliary neural network comprises selecting a parametrizable distribution among multiple parameterizable distributions, the selected parameterizable distribution providing the probability function used for determining how much information content exists in [the data] is not covered in [the other data]. (Page 3, section 5 states "In practice, the mutual information term   I ( c ; G ( z , c ) ) is hard to maximize directly as it requires access to the posterior P ( c | x ) . Fortunately we can obtain a lower bound of it by defining an auxiliary distribution Q ( c | x ) to approximate P ( c | x ) : I c ; G z , c = H c - H c G z , c In the equation, one can see that the part of the equation in the underbrace is equal to the conditional entropy H c G z , c . Wikipedia (“Conditional Entropy”) states that the conditional entropy is written as H Y X , in this case, the variable Y is instead c and X is G z , c . Further, Wikipedia, in the Venn diagram on page 1, where X is on the left in red and Y is on the right in blue, states that the conditional entropy is the part of the Venn diagram that is blue. This means that conditional entropy is the part of Y that is not covered by X. Therefore, this is inherent in Chen. Chen further states on page 4, section 5, that “Eq. (4) shows that the lower bound becomes tight as the auxiliary distribution Q approaches the true posterior distribution: E x [ D K L P ⋅ x     Q ⋅ x ) ] → 0 .” The auxiliary distribution Q therefore determines the conditional entropy when it is maximized. Section 6 states “In practice, we parameterize the auxiliary distribution Q as a neural network. In most experiments Q and D share all convolutional layers and there is one final fully connected layer to output parameters for the conditional distribution Q ( c | x ) , which means InfoGAN only adds a negligible computation cost to GAN.” The conditional distribution Q ( c | x ) is interpreted as the parametrizable probability distribution function.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Vandenhende, Li, and Poole with the teachings of Chen because as Chen states on page 4, We note in addition that the entropy of latent codes   H ( c ) can be optimized over as well since for common distributions it has a simple analytical form. However, in this paper we opt for simplicity by fixing the latent code distribution and we will treat H ( c ) as a constant. So far we have bypassed the problem of having to compute the posterior P ( c | x )   explicitly via this lower bound but we still need to be able to sample from the posterior in the inner expectation. Next we state a simple lemma, with its proof deferred to Appendix, that removes the need to sample from the posterior.” Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vandenhende, Li, and Poole as applied to claim 1 above, further in view of Zheng (“Data-driven Task Allocation for Multi-task Transfer Learning on the Edge”, July 2019). Regarding claim 12, the rejection of claim 1 is incorporated herein. Vandenhende teaches joint encoder portion of the neural network (Section 1, page 2 states "The proposed method aims to find an effective task grouping for the sharable layers f l of the encoder, i.e. grouping related tasks together in the same branches of the tree." The sharable layers of the encoder are interpreted as the joint encoder portion of a neural network.) The combination of Vandenhende, Li, and Poole does not appear to explicitly teach wherein the computing resources are allocated to at least one [section] based on a total number of processing tasks being handled by the at least one [section]. However, Zheng —directed to analogous art—teaches wherein computing resources are allocated to each [section] based on the number of tasks being handled by the respective [section]. (Page 1043 states “Since each task is assigned to exactly one processor, we have the following constraint: ∑ p ∈ P u j , p = 1 ,     ∀ j ∈ J . Additionally, the execution time and resource of all tasks assigned to the processor p should satisfy the following constraints: ∑ j ∈ J t j ⋅ u j , p ≤ T ,   ∀ p ∈ P ,     ∑ j ∈ J v j ⋅ u j , p ≤ V p ,   ∀ p ∈ P , where t j denotes the execution time of task j ; T denotes the time limit; v j denotes the resource required for task j ; V p denotes the resource capacity of processor p .” Each task is allocated to exactly one processor, meaning that the resources are based on the number of tasks.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Vandenhende, Li, and Poole with the teachings of Zheng because, as Zheng states on page 1041, “The benefits of multiple tasks come in mainly two ways. First, similar tasks can transfer their knowledge between each other during the training process, which reduces the negative effect of data scarcity, especially on the edge. Second, in the real-world scenario, it is common to make the final decision by aggregating the output of multiple tasks. Maintaining the high performance of all these tasks contribute to the final aggregated decision performance. Again in the example of a self-driving car, the final driving operation of the car is conducted based on the result of multiple data-driven tasks, e.g., the neighboring car, traffic-sign, and pedestrian detection.” Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vandenhende, Li, and Poole as applied to claim 1 above, further in view of Hattori (“Synthesizing a Scene-Specific Pedestrian Detector and Pose Estimator for Static Video Surveillance”, 2018), Regarding claim 13, the rejection of claim 1 is incorporated herein. The combination of Vandenhende, Li, and Poole does not appear to explicitly teach wherein the processing tasks include at least one tuple of processing tasks, and the at least one tuple of processing tasks include at least one of: depth estimation, detection of pedestrians, detection of traffic signs, pose detection of pedestrians, detection of drivable area. However, Hattori —directed to analogous art—teaches wherein the processing tasks include at least one tuple of processing tasks, and the at least one tuple of processing tasks include at least one of: depth estimation, detection of pedestrians, detection of traffic signs, pose detection of pedestrians, detection of drivable area. (Each processing task is interpreted as a tuple of processing tasks. Page 1031, Fig. 2 shows “Overview of our fully convolutional neural network architecture for multi-task learning method: with physically grounded and geometrically accurate renders of pedestrians for every grid location, our region specific pedestrian detection and pose estimation networks are trained on this synthetic data. At test time, our model takes a single image and outputs pedestrian detections, segmentation mask and body pose estimates.” It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Vandenhende, Li, and Poole with the task of Hattori because, as stated by Hattori in the abstract, "We demonstrate that when real human annotated data is scarce or non-existent, our data generation strategy can provide an excellent solution for an array of tasks for human activity analysis including detection, pose estimation and segmentation. Experimental results show that our approach (1) outperforms classical models and hybrid synthetic-real models, (2) outperforms various combinations of off-the-shelf state-of-the-art pedestrian detectors and pose estimators that are trained on real data, and (3) surprisingly, our method using purely synthetic data is able to outperform models trained on real scene-specific data when data is limited." Additionally, as stated by Hattori in the abstract, "We consider scenarios where we have zero instances of real pedestrian data (e.g., a newly installed surveillance system in a novel location in which no labeled real data or unsupervised real data exists yet) and a pedestrian detector must be developed prior to any observations of pedestrians. Given a single image and auxiliary scene information in the form of camera parameters and geometric layout of the scene, our approach infers and generates a large variety of geometrically and photometrically accurate potential images of synthetic pedestrians along with purely accurate ground-truth labels through the use of computer graphics rendering engine." Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JESSICA THUY PHAM whose telephone number is (571)272-2605. The examiner can normally be reached Monday - Friday, 9 A.M. - 5:00 P.M.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached at (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.T.P./Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

Show 5 earlier events
Jan 28, 2026
Request for Continued Examination
Feb 05, 2026
Response after Non-Final Action
Mar 04, 2026
Non-Final Rejection mailed — §103, §112
Mar 26, 2026
Interview Requested
Apr 02, 2026
Examiner Interview Summary
Apr 02, 2026
Applicant Interview (Telephonic)
Jun 04, 2026
Response Filed
Sep 22, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711363
SYSTEM FOR DYNAMIC AUTHENTICATION AND PROCESSING OF ELECTRONIC ACTIVITIES BASED ON PARALLEL NEURAL NETWORK PROCESSING
4y 10m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
18%
Grant Probability
99%
With Interview (+90.0%)
4y 0m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 11 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month