Prosecution Insights
Last updated: October 02, 2026
Application No. 18/468,873

MULTI-TASK GATING FOR MACHINE LEARNING SYSTEMS

Non-Final OA §101§102§103§112
Filed
Sep 18, 2023
Examiner
KIM, SEHWAN
Art Unit
2129
Tech Center
2100 — Computer Architecture & Software
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
61%
Grant Probability
Moderate
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
95 granted / 156 resolved
+5.9% vs TC avg
Strong +67% interview lift
Without
With
+67.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
32 currently pending
Career history
188
Total Applications
across all art units

Statute-Specific Performance

§101
20.3%
-19.7% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
7.7%
-32.3% vs TC avg
§112
23.3%
-16.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 156 resolved cases

Office Action

§101 §102 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Examiner’s Note Providing supporting paragraph(s) for each limitation of amended/new claim(s) in Remarks is strongly requested for clear and definite claim interpretations by Examiner (e.g., to avoid rejections under 35 U.S.C § 112(a) “Lack of written description”) Applicant can schedule interviews (via Automated Interview Request (AIR)) at any stage of the prosecution (e.g., Non-Final, Final, and After-Final) to discuss any issues related to, for example, rejections under 35 U.S.C § 101 and § 102/103, for moving toward allowance. If a limitation has bold brackets (i.e., [·]) around claim languages, the bracketed claim languages indicate that they have not been taught yet by the current prior art reference but they will be taught by another prior art reference afterwards. If a limitation has one or more bold underlines, the one or more bold underlined claim languages indicate that they are taught by the current prior art reference, while the one or more non-underlined claim languages indicate that they have been taught already by one or more previous art references. Priority Acknowledgment is made of applicant's claim for the present application filed on 09/18/2023. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim(s) 7-10, 26-29 is/are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim(s) 7 recite(s) the limitation “the layer of a second task-specific branch” (line 2). There is insufficient antecedent basis for this limitation in the claim. It is not clear what it is referring to. It appears it may need to read “a layer of a second task-specific branch”, or something else. For the purposes of examination, “a layer of a second task-specific branch” is used. In addition, claim(s) 26 is/are rejected for the same reason. Claim(s) 7, 26 each recite(s) limitations that raise issues of indefiniteness as set forth above, and their dependent claims are rejected at least based on their direct and/or indirect dependency from the claims listed above. Appropriate explanation and/or amendment is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 17-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 17 Step 1: “Is the claim to a process, machine, manufacture, or composition of matter?” The claim is directed to a method. Therefore, yes. Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?” perform the shared function on shared features of the input data for the first task using at least one of the one or more shared channels of the shared branch to generate a shared feature map; (i.e., mental process) perform the first task-specific function on first task-specific features of the input data using at least one of the one or more first task-specific channels associated with the first task-specific function to generate a first task-specific feature map; and (i.e., mental process) generate an output for the first task-specific branch based on performing the shared function on the shared features of the input data and performing the first task-specific function on the first task-specific features of the input data (i.e., mental process) The claim is directed to an abstract idea. Therefore, yes. Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements: An apparatus for performing at least one task, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: (well-understood, routine, and conventional generic computer and/or model, see MPEP 2106.05(f)) receive input data for a first task in a layer in a neural network (insignificant extra-solution activity of receiving data, see MPEP 2106.05(g)), wherein the layer is associated with a shared function of a shared branch and a first task-specific function of a first task-specific branch, and wherein the shared function is associated with one or more shared channels of the shared branch, and wherein the first task-specific function is associated with one or more first task-specific channels of the first task-specific branch; (a particular type or source of model/data, Field of Use and Technological Environment, see MPEP 2106.05(h)) Therefore, no. Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?” The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Specifically, the claimed inventions simply append well-understood, routine and conventional activities previously known to the industry, both when viewed independently and as an ordered combination, specified at a high level of generality, to the judicial exception, (e.g., a claim to an abstract idea requiring no more than a generic computer to perform generic computer functions that are well-understood, routine and conventional activities previously known to the industry). Therefore, no. Regarding claim 18 Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?” perform the shared function on shared features of the input data for the second task using at least one of the one or more shared channels of the shared feature map to generate a second shared feature map; (i.e., mental process) perform the second task-specific function on second task-specific features of the input data for the second task using at least one of the one or more second task-specific channels to generate a second task-specific feature map; and (i.e., mental process) generate an output for the first task-specific branch based on performing the shared function on the shared features of the input data and performing the first task-specific function on the first task-specific features of the input data. (i.e., mental process) The claim is directed to an abstract idea. Therefore, yes. Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements: wherein the at least one processor is further configured to: (well-understood, routine, and conventional generic computer and/or model, see MPEP 2106.05(f)) receive input data for a second task in the layer in the neural network (insignificant extra-solution activity of receiving data, see MPEP 2106.05(g)), wherein the layer is further associated with a second task-specific function, and wherein the second task-specific function is associated with one or more second task-specific channels; (a particular type or source of model/data, Field of Use and Technological Environment, see MPEP 2106.05(h)) Therefore, no. Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?” The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, no. Regarding claim 19 Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?” perform the third task-specific function on third task-specific features of the output using at least one of the plurality of third task-specific channels to generate a first subsequent feature map; (i.e., mental process) perform the second shared function on shared features of the output using at least one of the plurality of shared channels to generate a second subsequent feature map; and (i.e., mental process) generate a subsequent output for the first task-specific branch based on performing the second shared function on the shared features of the output and performing the third task-specific function on the third task-specific features of the output. (i.e., mental process) The claim is directed to an abstract idea. Therefore, yes. Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements: wherein the at least one processor is further configured to: (well-understood, routine, and conventional generic computer and/or model, see MPEP 2106.05(f)) receive the output in a subsequent layer of the neural network (insignificant extra-solution activity of receiving data, see MPEP 2106.05(g)), wherein the subsequent layer includes a third task-specific function of the first task-specific branch and a second shared function of the shared branch, and wherein the third task-specific function includes a plurality of third task-specific channels, and wherein the second shared function includes a plurality of shared channels; (a particular type or source of model/data, Field of Use and Technological Environment, see MPEP 2106.05(h)) Therefore, no. Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?” The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, no. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-4, 6-23, 25-30 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Bragman et al. (Stochastic Filter Groups for Multi-Task CNNs: Learning Specialist and Generalist Convolution Kernels) Regarding claim 1 Bragman teaches An apparatus for training a neural network to perform at least one task, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: (Bragman [sec(s) A] “All networks were trained with ADAM optimiser [36] with an initial learning rate of 10−3 and β = [0.9,0.999]. We used values of λ1 = 10−6 and λ2 = 10−5 for the weight and entropy regularisation factors in Equation (5) in Section 3.2. All stochastic filter group (SFG) modules were initialised with grouping probabilities p=[0.2, 0.6, 0.2] for every convolution kernel. Positivity of the grouping probabilities p is enforced by passing the output through a soft plus function f(x) = ln(1 + ex) as in [37]. The scheduler τ = max(0.10,exp(−rt)) recommended in [31] was used to anneal the Gumbel-Softmax temperature τ where r is the annealing rate and t is the current training iteration. We used r = 10−5 for our models. Hyper-parameters for the annealing rate and the entropy regularisation weight were obtained by analysis of the net work performance on a secondary randomly split on the UTK dataset (70/15/15). They were then applied to all trained models (large and small dataset for UTKFace and medical imaging dataset). … We used Tensorflow and implemented our models within the NiftyNet framework [41]. Models were trained on NVIDIA Titan Xp, P6000 and V100. All networks were trained in the Stochastic Filter Group paradigm.”;) -- obtain training data for a first task in a layer in a neural network, wherein the layer is associated with a first gating mechanism configured to determine whether to process shared features of the training data for the first task using a shared function of a shared branch or first task-specific features of the training data for the first task using a first task-specific function of a first task-specific branch, wherein the shared function is associated with one or more shared channels of the shared branch, and wherein the first task-specific function is associated with one or more first task-specific channels of the first task-specific branch; (Bragman [fig(s) 5] [sec(s) 3] “At l = 0, input image x is simply convolved with the first set of filter groups to yield F(1)i = h(1) x∗G(1)i ,i∈{1,2,s}.” [sec(s) 4] “UTKFace dataset: We tested our method on UTKFace [18], which consists of 23,703 cropped faced images in the wild with labels for age and gender. We created a dataset with a 70/15/15% split. We created a secondary separate dataset containing only 10% of images from the initial set, so as to simulate a data-starved scenario.” [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “For simplicity, we describe SFGs for the case of multitask learning with two tasks, but can be trivially extended to a larger number of tasks. At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3.”;) perform, based on a determination from the first gating mechanism, the shared function on the shared features of the training data for the first task using at least one of the one or more shared channels to generate a shared feature map; (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) perform, based on the determination from the first gating mechanism, the first task-specific function on the first task-specific features of the training data for the first task using at least one of the one or more first task-specific channels to generate a first task-specific feature map; (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) generate an output for the first task-specific branch based on performing the shared function on the shared features of the training data and performing the first task-specific function on the first task-specific features of the training data; and (Bragman [fig(s) 5] “This process repeats in the remaining SFG modules in the architecture until the last layer where the outputs of the final SFG module are combined into task-specific predictions y^1 and y^2.” [sec(s) 3.1] “Feature Routing: as shown in Fig. 4 (i), the features F(l)1, F(l)s, F(l)2 are routed to the filter groups G(l+1)1, G(l+1)s, G(l+1)2 in the subsequent (l+1)th layer in such a way to respect the task-specificity and sharedness of filter groups in the lth layer. Specifically, we perform the following routing for l > 0: PNG media_image1.png 228 702 media_image1.png Greyscale where each h(l+1) defines the choice of non-linear function, * denotes convolution operation and | denotes a merging operation of arrays (e.g. concatenation). … The merging modules, denoted as black circles, combine the task-specific and shared features appropriately, i.e. [F(l)i | F(l)s]; i = 1; 2 and pass them to the filter groups in the next layer. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) update at least one parameter of the first gating mechanism based on the output. (Bragman [fig(s) 4] “The task-specific groups G1, G2 are only updated based on the associated losses, while the shared group Gs is updated based on both” [sec(s) 3.2] “As the posterior distribution over the convolution kernels in SFG modules p(WjX;Y(1);Y(2)) is intractable, we approximate it with a simpler distribution qФ(W) … The set of variational parameters for each kernel in each layer is thus given by PNG media_image2.png 84 1121 media_image2.png Greyscale . … PNG media_image3.png 302 1106 media_image3.png Greyscale where λ1 > 0; λ2 > 0 are regularization coefficients. We note that the discrete sampling operation during filter group assignment (eq. (2)) creates discontinuities, giving the first term in the objective function (eq. 5) zero gradient with respect to the grouping probabilities {p(l);k}. We therefore, as employed in [16] for the binary case, approximate each of the categorical variables Cat(p(l),k) by the Gumbel-Softmax distribution, GSM(p(l),k, τ) [30, 31], a continuous relaxation which allows for sampling, differentiable with respect to the parameters p(l),k through a reparametrisation trick.”;) Regarding claim 2 The combination of Bragman teaches claim 1. Bragman teaches wherein each shared channel of the one or more channels respectively corresponds to each first task-specific channel of the one or more first task-specific channels. (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices.”;) Regarding claim 3 The combination of Bragman teaches claim 1. Bragman teaches the first gating mechanism includes a gate for each set of corresponding shared channels and first task-specific channels. (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices.” [sec(s) 3.2] “ PNG media_image4.png 128 828 media_image4.png Greyscale where z(l),k is the one-hot encoding of a sample from the categorical distribution over filter group assignments, and M(l),k denotes the parameters of the pre-grouping convolution kernel. The set of variational parameters for each kernel in each layer is thus given by PNG media_image2.png 84 1121 media_image2.png Greyscale .”;) Regarding claim 4 The combination of Bragman teaches claim 1. Bragman teaches the at least one parameter of the first gating mechanism includes a plurality of weights. (Bragman [sec(s) 3.2] “The set of variational parameters for each kernel in each layer is thus given by PNG media_image2.png 84 1121 media_image2.png Greyscale . … PNG media_image3.png 302 1106 media_image3.png Greyscale where λ1 > 0; λ2 > 0 are regularization coefficients. We note that the discrete sampling operation during filter group assignment (eq. (2)) creates discontinuities, giving the first term in the objective function (eq. 5) zero gradient with respect to the grouping probabilities {p(l);k}. We therefore, as employed in [16] for the binary case, approximate each of the categorical variables Cat(p(l),k) by the Gumbel-Softmax distribution, GSM(p(l),k, τ) [30, 31], a continuous relaxation which allows for sampling, differentiable with respect to the parameters p(l),k through a reparametrisation trick.” [sec(s) A] “All stochastic filter group (SFG) modules were initialised with grouping probabilities p=[0:2, 0:6, 0:2] for every convolution kernel. Positivity of the grouping probabilities p is enforced by passing the output through a softplus function.”;) Regarding claim 6 The combination of Bragman teaches claim 1. Bragman teaches determine a loss for first task-specific branch based on the output of the first task-specific branch; and (Bragman [fig(s) 4] “The task-specific groups G1, G2 are only updated based on the associated losses, while the shared group Gs is updated based on both.” [sec(s) 3.2] “As the posterior distribution over the convolution kernels in SFG modules p(WjX;Y(1);Y(2)) is intractable, we approximate it with a simpler distribution qФ(W) … The set of variational parameters for each kernel in each layer is thus given by PNG media_image2.png 84 1121 media_image2.png Greyscale . … PNG media_image3.png 302 1106 media_image3.png Greyscale where λ1 > 0; λ2 > 0 are regularization coefficients. We note that the discrete sampling operation during filter group assignment (eq. (2)) creates discontinuities, giving the first term in the objective function (eq. 5) zero gradient with respect to the grouping probabilities {p(l);k}. We therefore, as employed in [16] for the binary case, approximate each of the categorical variables Cat(p(l),k) by the Gumbel-Softmax distribution, GSM(p(l),k, τ) [30, 31], a continuous relaxation which allows for sampling, differentiable with respect to the parameters p(l),k through a reparametrisation trick.”;) update, using backpropagation, the at least one parameter of the first gating mechanism based on the loss. (Bragman [fig(s) 4] “The task-specific groups G1, G2 are only updated based on the associated losses, while the shared group Gs is updated based on both.” [sec(s) 3.2] “We therefore, as employed in [16] for the binary case, approximate each of the categorical variables Cat(p(l),k) by the Gumbel-Softmax distribution, GSM(p(l),k, τ) [30, 31], a continuous relaxation which allows for sampling, differentiable with respect to the parameters p(l),k through a reparametrisation trick.” [sec(s) 3] “We employ variational inference to learn the distributions over the possible grouping of kernels and network parameters that determines the connectivity between layers and the shared and task-specific features. This naturally results in a learning algorithm that optimally allocate representation capacity across multi-tasks via gradient-based stochastic optimization, e.g. stochastic gradient descent.”;) Regarding claim 7 The combination of Bragman teaches claim 1. Bragman teaches obtain training data for a second task in the layer of a second task-specific branch in the neural network, wherein the layer is further associated with a second gating mechanism configured to determine whether to process shared features of the training data for the second task using the shared function of the shared branch or second task-specific features of the training data for the second task using a second task-specific function, and wherein the second task-specific function is associated with one or more second task-specific channels of the second task-specific branch; (Bragman [sec(s) 3] “At l = 0, input image x is simply convolved with the first set of filter groups to yield F(1)i = h(1) x∗G(1)i ,i∈{1,2,s}.” [sec(s) 4] “UTKFace dataset: We tested our method on UTKFace [18], which consists of 23,703 cropped faced images in the wild with labels for age and gender. We created a dataset with a 70/15/15% split. We created a secondary separate dataset containing only 10% of images from the initial set, so as to simulate a data-starved scenario.” [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “For simplicity, we describe SFGs for the case of multitask learning with two tasks, but can be trivially extended to a larger number of tasks. At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3.”;) perform, based on a determination from the second gating mechanism, the shared function on the shared features of the training data for the second task using at least one of the one or more shared channels of the shared feature map to generate a second shared feature map; (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) perform, based on a determination from the second gating mechanism, the second task-specific function on the second task-specific features of the training data for the second task using at least one of the one or more second task-specific channels to generate a second task-specific feature map; (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale ”;) generate a second output for the second task-specific branch based on performing the shared function on the shared features of the training data and performing the second task-specific function on the second task-specific features; and (Bragman [fig(s) 5] “This process repeats in the remaining SFG modules in the architecture until the last layer where the outputs of the final SFG module are combined into task-specific predictions y^1 and y^2.” [sec(s) 3.1] “Feature Routing: as shown in Fig. 4 (i), the features F(l)1, F(l)s, F(l)2 are routed to the filter groups G(l+1)1, G(l+1)s, G(l+1)2 in the subsequent (l+1)th layer in such a way to respect the task-specificity and sharedness of filter groups in the lth layer. Specifically, we perform the following routing for l > 0: PNG media_image1.png 228 702 media_image1.png Greyscale where each h(l+1) defines the choice of non-linear function, * denotes convolution operation and | denotes a merging operation of arrays (e.g. concatenation). … The merging modules, denoted as black circles, combine the task-specific and shared features appropriately, i.e. [F(l)i | F(l)s]; i = 1; 2 and pass them to the filter groups in the next layer.”;) update at least one parameter of the second gating mechanism based on the second output. (Bragman [fig(s) 4] “The task-specific groups G1, G2 are only updated based on the associated losses, while the shared group Gs is updated based on both” [sec(s) 3.2] “As the posterior distribution over the convolution kernels in SFG modules p(WjX;Y(1);Y(2)) is intractable, we approximate it with a simpler distribution qФ(W) … The set of variational parameters for each kernel in each layer is thus given by PNG media_image2.png 84 1121 media_image2.png Greyscale . … PNG media_image3.png 302 1106 media_image3.png Greyscale where λ1 > 0; λ2 > 0 are regularization coefficients. We note that the discrete sampling operation during filter group assignment (eq. (2)) creates discontinuities, giving the first term in the objective function (eq. 5) zero gradient with respect to the grouping probabilities {p(l);k}. We therefore, as employed in [16] for the binary case, approximate each of the categorical variables Cat(p(l),k) by the Gumbel-Softmax distribution, GSM(p(l),k, τ) [30, 31], a continuous relaxation which allows for sampling, differentiable with respect to the parameters p(l),k through a reparametrisation trick.”;) Regarding claim 8 The combination of Bragman teaches claim 7. Bragman teaches determine a second loss for the second task-specific branch based on the output of the second task-specific branch; and (Bragman [fig(s) 4] “The task-specific groups G1, G2 are only updated based on the associated losses, while the shared group Gs is updated based on both.” [sec(s) 3.2] “As the posterior distribution over the convolution kernels in SFG modules p(WjX;Y(1);Y(2)) is intractable, we approximate it with a simpler distribution qФ(W) … The set of variational parameters for each kernel in each layer is thus given by PNG media_image2.png 84 1121 media_image2.png Greyscale . … PNG media_image3.png 302 1106 media_image3.png Greyscale where λ1 > 0; λ2 > 0 are regularization coefficients. We note that the discrete sampling operation during filter group assignment (eq. (2)) creates discontinuities, giving the first term in the objective function (eq. 5) zero gradient with respect to the grouping probabilities {p(l);k}. We therefore, as employed in [16] for the binary case, approximate each of the categorical variables Cat(p(l),k) by the Gumbel-Softmax distribution, GSM(p(l),k, τ) [30, 31], a continuous relaxation which allows for sampling, differentiable with respect to the parameters p(l),k through a reparametrisation trick.”;) update, using backpropagation, the at least one parameter of the second gating mechanism based on the second loss. (Bragman [fig(s) 4] “The task-specific groups G1, G2 are only updated based on the associated losses, while the shared group Gs is updated based on both.” [sec(s) 3.2] “We therefore, as employed in [16] for the binary case, approximate each of the categorical variables Cat(p(l),k) by the Gumbel-Softmax distribution, GSM(p(l),k, τ) [30, 31], a continuous relaxation which allows for sampling, differentiable with respect to the parameters p(l),k through a reparametrisation trick.” [sec(s) 3] “We employ variational inference to learn the distributions over the possible grouping of kernels and network parameters that determines the connectivity between layers and the shared and task-specific features. This naturally results in a learning algorithm that optimally allocate representation capacity across multi-tasks via gradient-based stochastic optimization, e.g. stochastic gradient descent.”;) Regarding claim 9 The combination of Bragman teaches claim 7. Bragman teaches wherein a first subset of the one or more shared channels is allocated to a first subset of the training data for the first task and a second subset of the one or more shared channels is allocated to the first subset of the training data for the second task. (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “For simplicity, we describe SFGs for the case of multitask learning with two tasks, but can be trivially extended to a larger number of tasks. At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3.”;) Regarding claim 10 The combination of Bragman teaches claim 7. Bragman teaches wherein a classification of the first task is different from a classification of the second task. (Bragman [sec(s) 4] “We tested stochastic filter groups (SFG) on two multitask learning (MTL) problems: 1) age regression and gender classification from face images on UTKFace dataset [18] and 2) semantic image regression (synthesis) and segmentation on a medical imaging dataset. Full details of the training and datasets are provided in Sec. A in the supplementary materials.”;) Regarding claim 11 The combination of Bragman teaches claim 1. Bragman teaches based on the first gating mechanism, select between at least one channel of the one or more first task-specific channels and at least one corresponding channel of the one or more shared channels; and (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “For simplicity, we describe SFGs for the case of multitask learning with two tasks, but can be trivially extended to a larger number of tasks. At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3.”;) prune unselected channels based on the selection by the first gating mechanism. (Bragman [sec(s) 3.1] “At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices.”;) Regarding claim 12 The combination of Bragman teaches claim 1. Bragman teaches a number of active first task-specific channels is different from a number of active shared channels. (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “For simplicity, we describe SFGs for the case of multitask learning with two tasks, but can be trivially extended to a larger number of tasks. At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3.”;) Regarding claim 13 The combination of Bragman teaches claim 1. Bragman teaches wherein a first channel of the one or more shared channels corresponds to a first channel of the one or more first task-specific channels. (Bragman [sec(s) 3.1] “At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices.”;) Regarding claim 14 The combination of Bragman teaches claim 13. Bragman teaches the first channel of the one or more shared channels is active and the first channel of the one or more first task-specific channels is inactive. (Bragman [sec(s) 3.1] “At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices.”;) Regarding claim 15 The combination of Bragman teaches claim 13. Bragman teaches the first channel of the one or more shared channels in inactive and the first channel of the one or more first task-specific channels is active. (Bragman [sec(s) 3.1] “At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices.”;) Regarding claim 16 The combination of Bragman teaches claim 1. Bragman teaches the first task is one of image segmentation, surface normal estimation, depth estimation, or classification. (Bragman [sec(s) 4] “We tested stochastic filter groups (SFG) on two multitask learning (MTL) problems: 1) age regression and gender classification from face images on UTKFace dataset [18] and 2) semantic image regression (synthesis) and segmentation on a medical imaging dataset. Full details of the training and datasets are provided in Sec. A in the supplementary materials.”;) Regarding claim 17 Bragman teaches An apparatus for performing at least one task, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: (Bragman [sec(s) A] “All networks were trained with ADAM optimiser [36] with an initial learning rate of 10−3 and β = [0.9,0.999]. We used values of λ1 = 10−6 and λ2 = 10−5 for the weight and entropy regularisation factors in Equation (5) in Section 3.2. All stochastic filter group (SFG) modules were initialised with grouping probabilities p=[0.2, 0.6, 0.2] for every convolution kernel. Positivity of the grouping probabilities p is enforced by passing the output through a soft plus function f(x) = ln(1 + ex) as in [37]. The scheduler τ = max(0.10,exp(−rt)) recommended in [31] was used to anneal the Gumbel-Softmax temperature τ where r is the annealing rate and t is the current training iteration. We used r = 10−5 for our models. Hyper-parameters for the annealing rate and the entropy regularisation weight were obtained by analysis of the net work performance on a secondary randomly split on the UTK dataset (70/15/15). They were then applied to all trained models (large and small dataset for UTKFace and medical imaging dataset). … We used Tensorflow and implemented our models within the NiftyNet framework [41]. Models were trained on NVIDIA Titan Xp, P6000 and V100. All networks were trained in the Stochastic Filter Group paradigm.”;) receive input data for a first task in a layer in a neural network, wherein the layer is associated with a shared function of a shared branch and a first task-specific function of a first task-specific branch, and wherein the shared function is associated with one or more shared channels of the shared branch, and wherein the first task-specific function is associated with one or more first task-specific channels of the first task-specific branch; (Bragman [fig(s) 5] [sec(s) 3] “At l = 0, input image x is simply convolved with the first set of filter groups to yield F(1)i = h(1) x∗G(1)i ,i∈{1,2,s}.” [sec(s) 4] “UTKFace dataset: We tested our method on UTKFace [18], which consists of 23,703 cropped faced images in the wild with labels for age and gender. We created a dataset with a 70/15/15% split. We created a secondary separate dataset containing only 10% of images from the initial set, so as to simulate a data-starved scenario.” [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “For simplicity, we describe SFGs for the case of multitask learning with two tasks, but can be trivially extended to a larger number of tasks. At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3.”;) perform the shared function on shared features of the input data for the first task using at least one of the one or more shared channels of the shared branch to generate a shared feature map; (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) perform the first task-specific function on first task-specific features of the input data using at least one of the one or more first task-specific channels associated with the first task-specific function to generate a first task-specific feature map; and (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) generate an output for the first task-specific branch based on performing the shared function on the shared features of the input data and performing the first task-specific function on the first task-specific features of the input data. (Bragman [fig(s) 5] “This process repeats in the remaining SFG modules in the architecture until the last layer where the outputs of the final SFG module are combined into task-specific predictions y^1 and y^2.” [sec(s) 3.1] “Feature Routing: as shown in Fig. 4 (i), the features F(l)1, F(l)s, F(l)2 are routed to the filter groups G(l+1)1, G(l+1)s, G(l+1)2 in the subsequent (l+1)th layer in such a way to respect the task-specificity and sharedness of filter groups in the lth layer. Specifically, we perform the following routing for l > 0: PNG media_image1.png 228 702 media_image1.png Greyscale where each h(l+1) defines the choice of non-linear function, * denotes convolution operation and | denotes a merging operation of arrays (e.g. concatenation). … The merging modules, denoted as black circles, combine the task-specific and shared features appropriately, i.e. [F(l)i | F(l)s]; i = 1; 2 and pass them to the filter groups in the next layer. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) Regarding claim 18 The combination of Bragman teaches claim 17. Bragman teaches receive input data for a second task in the layer in the neural network, wherein the layer is further associated with a second task-specific function, and wherein the second task-specific function is associated with one or more second task-specific channels; (Bragman [fig(s) 5] [sec(s) 3] “At l = 0, input image x is simply convolved with the first set of filter groups to yield F(1)i = h(1) x∗G(1)i ,i∈{1,2,s}.” [sec(s) 4] “UTKFace dataset: We tested our method on UTKFace [18], which consists of 23,703 cropped faced images in the wild with labels for age and gender. We created a dataset with a 70/15/15% split. We created a secondary separate dataset containing only 10% of images from the initial set, so as to simulate a data-starved scenario.” [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “For simplicity, we describe SFGs for the case of multitask learning with two tasks, but can be trivially extended to a larger number of tasks. At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3.”;) perform the shared function on shared features of the input data for the second task using at least one of the one or more shared channels of the shared feature map to generate a second shared feature map; (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) perform the second task-specific function on second task-specific features of the input data for the second task using at least one of the one or more second task-specific channels to generate a second task-specific feature map; and (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) generate an output for the first task-specific branch based on performing the shared function on the shared features of the input data and performing the first task-specific function on the first task-specific features of the input data. (Bragman [fig(s) 5] “This process repeats in the remaining SFG modules in the architecture until the last layer where the outputs of the final SFG module are combined into task-specific predictions y^1 and y^2.” [sec(s) 3.1] “Feature Routing: as shown in Fig. 4 (i), the features F(l)1, F(l)s, F(l)2 are routed to the filter groups G(l+1)1, G(l+1)s, G(l+1)2 in the subsequent (l+1)th layer in such a way to respect the task-specificity and sharedness of filter groups in the lth layer. Specifically, we perform the following routing for l > 0: PNG media_image1.png 228 702 media_image1.png Greyscale where each h(l+1) defines the choice of non-linear function, * denotes convolution operation and | denotes a merging operation of arrays (e.g. concatenation). … The merging modules, denoted as black circles, combine the task-specific and shared features appropriately, i.e. [F(l)i | F(l)s]; i = 1; 2 and pass them to the filter groups in the next layer. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices. In the simplest form with no additional transformation (i.e. the grey circles in Fig. 5 are identity functions), we define the merging operation [F(l)i|F(l)s], i = 1; 2 as pixel-wise summation.”;) Regarding claim 19 The combination of Bragman teaches claim 17. Bragman teaches receive the output in a subsequent layer of the neural network, wherein the subsequent layer includes a third task-specific function of the first task-specific branch and a second shared function of the shared branch, and wherein the third task-specific function includes a plurality of third task-specific channels, and wherein the second shared function includes a plurality of shared channels; (Bragman [fig(s) 5] “This process repeats in the remaining SFG modules in the architecture until the last layer where the outputs of the final SFG module are combined into task-specific predictions y^1 and y^2.” [sec(s) 3.1] “Feature Routing: as shown in Fig. 4 (i), the features F(l)1, F(l)s, F(l)2 are routed to the filter groups G(l+1)1, G(l+1)s, G(l+1)2 in the subsequent (l+1)th layer in such a way to respect the task-specificity and sharedness of filter groups in the lth layer. Specifically, we perform the following routing for l > 0: PNG media_image1.png 228 702 media_image1.png Greyscale where each h(l+1) defines the choice of non-linear function, * denotes convolution operation and | denotes a merging operation of arrays (e.g. concatenation). … The merging modules, denoted as black circles, combine the task-specific and shared features appropriately, i.e. [F(l)i | F(l)s]; i = 1; 2 and pass them to the filter groups in the next layer. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices. In the simplest form with no additional transformation (i.e. the grey circles in Fig. 5 are identity functions), we define the merging operation [F(l)i|F(l)s], i = 1; 2 as pixel-wise summation.”;) perform the third task-specific function on third task-specific features of the output using at least one of the plurality of third task-specific channels to generate a first subsequent feature map; (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices.”;) perform the second shared function on shared features of the output using at least one of the plurality of shared channels to generate a second subsequent feature map; and (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … PNG media_image1.png 228 702 media_image1.png Greyscale … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices”;) generate a subsequent output for the first task-specific branch based on performing the second shared function on the shared features of the output and performing the third task-specific function on the third task-specific features of the output. (Bragman [fig(s) 5] “This process repeats in the remaining SFG modules in the architecture until the last layer where the outputs of the final SFG module are combined into task-specific predictions y^1 and y^2.” [sec(s) 3.1] “Feature Routing: as shown in Fig. 4 (i), the features F(l)1, F(l)s, F(l)2 are routed to the filter groups G(l+1)1, G(l+1)s, G(l+1)2 in the subsequent (l+1)th layer in such a way to respect the task-specificity and sharedness of filter groups in the lth layer. Specifically, we perform the following routing for l > 0: PNG media_image1.png 228 702 media_image1.png Greyscale where each h(l+1) defines the choice of non-linear function, * denotes convolution operation and | denotes a merging operation of arrays (e.g. concatenation). … The merging modules, denoted as black circles, combine the task-specific and shared features appropriately, i.e. [F(l)i | F(l)s]; i = 1; 2 and pass them to the filter groups in the next layer. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices. In the simplest form with no additional transformation (i.e. the grey circles in Fig. 5 are identity functions), we define the merging operation [F(l)i|F(l)s], i = 1; 2 as pixel-wise summation.”;) Regarding claim 20 The claim is a method claim corresponding to the system claim 1, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Regarding claim 21 The claim is a method claim corresponding to the system claim 2, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Regarding claim 22 The claim is a method claim corresponding to the system claim 3, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Regarding claim 23 The claim is a method claim corresponding to the system claim 4, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Regarding claim 25 The claim is a method claim corresponding to the system claim 6, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Regarding claim 26 The claim is a method claim corresponding to the system claim 7, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Regarding claim 27 The claim is a method claim corresponding to the system claim 8, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Regarding claim 28 The claim is a method claim corresponding to the system claim 9, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Regarding claim 29 The claim is a method claim corresponding to the system claim 10, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Regarding claim 30 The claim is a method claim corresponding to the system claim 11, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 5, 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bragman et al. (Stochastic Filter Groups for Multi-Task CNNs: Learning Specialist and Generalist Convolution Kernels) in view of Hu et al. (Squeeze-and-Excitation Networks) Regarding claim 5 The combination of Bragman teaches claim 4. Bragman teaches process the at least one parameter using a [sigmoid] function to generate a value; (Bragman [sec(s) 3.2] “The set of variational parameters for each kernel in each layer is thus given by PNG media_image2.png 84 1121 media_image2.png Greyscale . … PNG media_image3.png 302 1106 media_image3.png Greyscale where λ1 > 0; λ2 > 0 are regularization coefficients. We note that the discrete sampling operation during filter group assignment (eq. (2)) creates discontinuities, giving the first term in the objective function (eq. 5) zero gradient with respect to the grouping probabilities {p(l);k}. We therefore, as employed in [16] for the binary case, approximate each of the categorical variables Cat(p(l),k) by the Gumbel-Softmax distribution, GSM(p(l),k, τ) [30, 31], a continuous relaxation which allows for sampling, differentiable with respect to the parameters p(l),k through a reparametrisation trick.” [sec(s) A] “All stochastic filter group (SFG) modules were initialised with grouping probabilities p=[0:2, 0:6, 0:2] for every convolution kernel. Positivity of the grouping probabilities p is enforced by passing the output through a softplus function.”;) compare the value to a threshold value to provide a binary selection; and (Bragman [sec(s) 3.2] “ PNG media_image4.png 128 828 media_image4.png Greyscale where z(l),k is the one-hot encoding of a sample from the categorical distribution over filter group assignments, and M(l),k denotes the parameters of the pre-grouping convolution kernel. The set of variational parameters for each kernel in each layer is thus given by PNG media_image2.png 84 1121 media_image2.png Greyscale .”;) select, based on the binary selection, between the first task-specific function and the shared function. (Bragman [sec(s) 3] “We propose stochastic filter groups (SFG), a probabilistic mechanism to partition kernels in each convolution layer into “specialist” groups or a “shared” group, which are specific to or shared across different tasks, respectively.” [sec(s) 3.1] “At the lth convolution layer in a CNN architecture with Kl kernels {w(l),k}Klk=1, the associated SFG performs two operations: 1. Filter Assignment: each kernel w(l)k is stochastically assigned to either: i) the “task-1 specific group” G(l)1, ii) “shared group” G(l)s or iii) “task-2 specific group” G(l)2 with respective probabilities p(l);k = [p(l);k1 ; p(l);ks ; p(l);k2] ϵ [0; 1]3. Convolving with the respective filter groups yields distinct sets of features F(l)1, F(l)s, F(l)2. Fig. 2 illustrates this operation and Fig. 3 shows different learnable patterns. … At each SFG module, we first convolve the input features with all kernels, and generate the output features from each filter group by zeroing out the channels that root from the kernels in the other groups, resulting in F(l)1, F(l)s, F(l)2 that are sparse at non-overlapping channel indices.”;) However, the combination of Bragman does not appear to explicitly teach: process the at least one parameter using a [sigmoid] function to generate a value; Hu teaches process the at least one parameter using a sigmoid function to generate a value; (Bragman [fig(s) 2-3] [sec(s) 3] “To fulfil this objective, the function must meet two criteria: first, it must be flexible (in particular, it must be capable of learning a nonlinear interaction between channels) and second, it must learn a non-mutually-exclusive relationship since we would like to ensure that multiple channels are allowed to be emphasised opposed to one-hot activation. To meet these criteria, we opt to employ a simple gating mechanism with a sigmoid activation: s =Fex(z,W) = σ(g(z,W)) = σ(W2δ(W1z)), (3) where δ refers to the ReLU[30]function, W1 ∈ RC/r ×C and W2 ∈RC×C/r . To limit model complexity and aid generalisation, we parameterise the gating mechanism by forming a bottleneck with two fully connected (FC) layers around the non-linearity, i.e. a dimensionality-reduction layer with parameters W1 with reduction ratio r (this parameter choice is discussed in Sec. 6.4), a ReLU and then a dimensionality increasing layer with parameters W2.” [sec(s) 2] “It is typically implemented in combination with a gating function (e.g. a soft max or sigmoid) and sequential techniques [12, 41]. Recent work has shown its applicability to tasks such as image captioning [4, 48] and lip reading [7]. In these applications, it is often used on top of one or more layers rep resenting higher-level abstractions for adaptation between modalities.”;) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bragman with the sigmoid of Hu. One of ordinary skill in the art would have been motived to combine in order to improve the accuracy by a large margin at minimal increases in computational cost. (Hu [sec(s) 6] “Finally, we evaluate on two representative efficient architectures, MobileNet [13] and ShuffleNet [52] in Table 3, showing that SE blocks can consistently improve the accuracy by a large margin at minimal increases in computational cost. These experiments demonstrate that improvements induced by SE blocks can be used in combination with a wide range of architectures. Moreover, this result holds for both residual and non-residual foundations.”) Regarding claim 24 The claim is a method claim corresponding to the system claim 5, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the system claim. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEHWAN KIM whose telephone number is (571)270-7409. The examiner can normally be reached Mon - Fri 9:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached on (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SEHWAN KIM/Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Sep 18, 2023
Application Filed
Aug 21, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12619853
DECISION-MAKING DEVICE, UNMANNED SYSTEM, DECISION-MAKING METHOD, AND PROGRAM
5y 6m to grant Granted May 05, 2026
Patent 12619921
PREDICTIVE FOG DATA CENTER MIGRATION
3y 8m to grant Granted May 05, 2026
Patent 12608592
AUTOMATED ELECTRIC SUBMERSIBLE PUMP (ESP) FAILURE ANALYSIS
3y 4m to grant Granted Apr 21, 2026
Patent 12602595
SYSTEM AND METHOD OF USING A KNOWLEDGE REPRESENTATION FOR FEATURES IN A MACHINE LEARNING CLASSIFIER
9y 4m to grant Granted Apr 14, 2026
Patent 12602580
Dataset Dependent Low Rank Decomposition Of Neural Networks
6y 9m to grant Granted Apr 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
61%
Grant Probability
99%
With Interview (+67.3%)
4y 0m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 156 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month