Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on July 27, 2026 has been entered.
Response to Amendment
In the previous Office Action issued February 25, 2026 (hereinafter “the previous Office Action”), claims 1-20 were pending.
This action is in response to the amendment and remarks filed July 27, 2026. In the amendment, claims 1, 5-6, 11, and 15-16 were amended, no claims were canceled, and no claims were added. Thus, claims 1-20 are pending.
The rejections of claims 1-20 under 35 U.S.C. § 112(b), set forth in the previous Office Action, have been withdrawn in view of Applicant’s amendments and remarks.
The rejections of claims 1-20 under 35 U.S.C. § 101, set forth in the previous Office Action, have been withdrawn in view of Applicant’s amendments and remarks.
Drawings
FIGS. 5, 6, 7, 8A, 8B, and 9 of the drawings, filed February 3, 2022, are objected to because they contain text oriented in a different than the view [see CFR 1.84(p)(1)]. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-7 and 11-17 are rejected under 35 U.S.C. 103 as being unpatentable over Ma et al. (“Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts”), hereinafter Ma, in view of Tang et al. (“Progressive Layered Extraction (PLE): A Novel Multi-Task Learning (MTL) Model for Personalized Recommendations”), hereinafter Tang.
Regarding Claim 1:
Ma discloses:
A system for training a heterogeneous multi-task learning network, the system comprising: at least one processor; and a memory comprising instructions which, when executed by the processor, configure the processor to:
Ma, p. 1935, “All of them are trained together via standard backpropagation”
p. 1933, “The model is implemented using TensorFlow [1] and trained using Adam optimizer [25] with the default setting”
Figure 1(c), p. 1931, teaches the MMOE model is implanted using TensorFlow and run the backpropagation corresponding to backpropagation use the processor to run the algorithm, MMOE corresponding to heterogeneous multi-task learning network.
provide a learning network architecture comprising:
PNG
media_image1.png
407
989
media_image1.png
Greyscale
On p. 1931, Ma discloses their model architecture in FIG. 1 [provide a learning network architecture].
a plurality of expert models assigned to a plurality of tasks in the multi-task learning network, the plurality of expert models including:
Ma, p. 1934, “The new model is called Multi-gate Mixture-of-Experts (MMoE) model, where the key idea is to substitute the shared bottom network
f
in Eq 1 with the MoE layer in Eq 5. More importantly, we add a separate gating network
g
k
for each task
k
…Each gating network can learn to ‘select’ a subset of experts to use conditioned on the input example.”
In view of FIG. 1 above, (c) depicts and Ma discloses a network where experts [a plurality of expert models] are connected to multiple tasks through a gating network [assigned to a plurality of tasks in the multi-task learning network].
at least one shared expert model contributing to the first task and at least one other task of the plurality of tasks
Ma, p. 1934, “Each gating network can learn to ‘select’ a subset of experts to use conditioned on the input example…The MMoE is able to model the task relationships in a sophisticated way by deciding how the separations resulted by different gates overlap with each other. If the tasks are less related, then sharing experts will be penalized and the gating networks of these tasks will learn to utilize different experts instead.”
p. 1931, “each task has an individual ‘tower’ of network on top of the bottom representations.”
On p. 1934, Ma discloses sharing experts [at least one shared expert model]. In view of FIG. 1(c) above, Ma depicts expert models such as sharing experts that are connected to multiple towers, also referred to as tasks [at least one shared expert model contributing to the first task and at least one other task of the plurality of tasks].
a gating network defined by a plurality of gate weights, the gate weights controlling a contribution of the plurality of expert models to the plurality of tasks, the gate weights including first task gate weights controlling the contribution from the exclusive expert model and the at least one shared expert model to the first task
Ma, p. 1934, “The gating networks are simply linear transformations of the input with a softmax layer:
g
k
x
=
s
o
f
t
m
a
x
W
g
k
x
,
where
W
g
k
∈
R
n
×
d
is a trainable matrix.
n
is the number of experts and
d
is the feature dimension. Each gating network can learn to ‘select’ a subset of experts to use conditioned on the input example. This is desirable for a flexible parameter sharing in the multi-task learning situation. As a special case, if only one expert with the highest gate score is selected, each gating network actually linearly separates the input space into
n
regions with each region corresponding to an expert. The MMoE is able to model the task relationships in a sophisticated way by deciding how the separations resulted by different gates overlap with each other. If the tasks are less related, then sharing experts will be penalized and the gating networks of these tasks will learn to utilize different experts instead.”
In view of FIG. 1(c) above, Ma discloses gating networks with weights,
W
g
k
[a gating network defined by a plurality of gate weights]. Ma says the gating network selects which subset of experts to use for each task [the gate weights controlling a contribution of the plurality of expert models to the plurality of tasks]. As discussed above, Ma discloses the exclusive expert model and the at least one shared expert model for the first task.
for each task: initialize weight parameters in the expert models; initialize the gate weights
Ma, p. 1934, “The gating networks are simply linear transformations of the input with a softmax layer:
g
k
x
=
s
o
f
t
m
a
x
W
g
k
x
,
where
W
g
k
∈
R
n
×
d
is a trainable matrix.”
p. 1935, “we implement Tucker decomposition for learning multi-task models, which is reported to deliver the most reliable results [34]. For example, given input hidden-layer size
m
, output hidden-layer size
n
and task number
k
, the weights
W
…”
p. 1936, “we train each method on training dataset 400 times with random parameter initialization and report the results on the test dataset”
On p. 1934, Ma discloses weights for the gate network [gate weights]. Then on p. 1935, Ma discloses weights for their multi-task models [weight parameters in the expert models]. Lastly, on p. 1936, Ma discloses initialization of the parameters for their model. Putting it together, Ma discloses initializing the weight parameters of their expert models and the gate weights of their gating networks.
provide training inputs to the multi-task learning network
Ma, p. 1936, “we train each method on training dataset 400 times with random parameter initialization and report the results on the test dataset”
As depicted above in FIG. 1 (c), Ma depicts input into their Multi-gate MoE model [provide…inputs to the multi-task learning network], and p. 1936 of Ma discloses training. Therefore, it follows Ma discloses training inputs to be provided to their Multi-gate MoE model.
determine a loss following a forward pass over the multi-task learning network
Ma, p. 1937, “For the Shared-Bottom model, we implement the shared bottom network as a feed forward neural network with several fully-connected layers with ReLU activation”
p. 1935, “it’s worth to observe that the lowest losses of all the three models are comparable”
Me discloses performed forward pass and compute losses for each training epoch.
back propagate losses and update the weight parameters for the expert models and the gate weights
Ma, p. 1937, “All models are optimized using mini-batch Stochastic Gradient Descent (SGD) with batch size 1024”
p. 1935, “r size n and task number k, the weights W, which is a m × n × k tensor, is derived from the following equation: W = Xr1 i1 Xr2 i2 Xr3 i3 S(i1,i2,i3) · U1 (:,i1) ◦ U2 (:,i2) ◦ U3 (:,i3), where tensor S of size r1 × r2 × r3, matrix U1 of size m × r1, U2 of size n × r2, and U3 of size k × r3 are trainable parameters. All of them are trained together via standard backpropagation. r1, r2 and r3 are hyper-parameters”
Ma discloses using SGD and standard backpropagation to update parameters for the expert and gate functions [back propagate losses and update the weight parameters for the expert models and the gate weights].
store a final set of weight parameters for use in a trained model for multiple tasks
Ma, p. 1937, “We show the results after training 2 million steps (10 billion examples with batch size 1024), 4 million steps and 6 million steps. MMoE outperforms other models in terms of both metrics. L2- Constrained and Cross-Stitch are worse than the Shared-Bottom model”
Ma discloses after training evaluating model corresponding to storing final set of weight parameters.
Ma does not explicitly disclose:
an exclusive expert model only connected to a first task of the plurality of tasks
However, in the same field, analogous art Tang teaches:
an exclusive expert model only connected to a first task of the plurality of tasks
PNG
media_image2.png
486
553
media_image2.png
Greyscale
Tang, p. 273, FIG. 4 depicts expert models that are exclusively connected [an exclusive expert model only connected to] to a tower A [a first task of the plurality of tasks].
Ma, Tang, and the instant application are analogous art because they are all directed to mixture-of-expert models.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Ma with Tang in order to create a more robust model. “In this paper, we propose a novel MTL model called Progressive Layered Extraction (PLE), which separates task-sharing and task specific parameters explicitly and introduces an innovative progressive routing manner to avoid the negative transfer and seesaw phenomenon, and achieve more efficient information sharing and joint representation learning. Offline and online experiment results on the industrial dataset and public benchmark datasets show significant and consistent improvements of PLE over SOTA MTL models” (Tang, p. 278).
Regarding Claim 2:
As discussed above Ma in view of Tang teach [the] system as claimed in claim 1, and Ma further discloses:
wherein the at least one processor is configured to provide input to the trained model to perform the multiple tasks
Ma, p. 1934, “Each gating network can learn to ‘select’ a subset of experts to use conditioned on the input example. This is desirable for flexible parameter sharing in the multi-task learning situation”
p. 1934, “All the models are trained with the Adam optimizer and the learning rate is grid searched from [0.0001, 0.001, 0.01]. For each model-correlation pair setting, we have 200 runs with independent random data generation and model initialization. The average results are shown in figure 4”
Ma discloses perform multiple task using the training model [provide input to the trained model to perform the multiple tasks].
Regarding Claim 3:
As discussed above Ma in view of Tang teach [the] system as claimed in claim 1, and Ma further discloses:
wherein each expert model comprises one or more neural networks layers
Ma, p. 1931 “In our paper, each expert is a feed-forward network”
Ma discloses that the expert models are feed forward neural network [each expert model comprises one or more neural networks layers].
Regarding Claim 4:
As discussed above Ma in view of Tang teach [the] system as claimed in claim 3, and Ma further discloses:
wherein one of:
temporal data is provided as input and the expert models comprise recurrent layers; or
non-temporal data is provided as input and the expert models comprise dense layers
Ma, p. 1931, “In our paper, each expert is a feed-forward network. We then introduce a gating network for each task”
p. 1934, “we add a separate gating network
g
k
for each task
k
. More precisely, the output of task
k
is
y
k
=
h
k
f
k
x
,
where
f
k
x
=
∑
i
=
1
n
g
k
x
i
f
i
x
”
p. 1937, “we conduct experiments on a large-scale content recommendation system in Google Inc., where the recommendations are generated from hundreds of millions of unique items for billions of users. Specifically, given a user’s current behavior of consuming an item, this recommendation system targets at showing the user a list of relevant items to consume next”
Ma discloses each expert is a feed forward neural network [the expert models comprise dense layers] wherein the system use no sequence and timestamps data which is user, context features corresponding to non-temporal data [non-temporal data is provided as input]
Regarding Claim 5:
As discussed above Ma in view of Tang teach [the] system as claimed in claim 1, and Tang further discloses:
wherein the gate weights are applicable to gate functions comprising an exclusivity mechanism for setting expert models to be exclusively connected to one task
Tang, p. 273, “In CGC, shared experts and task-specific experts are combined through a gating network for selective fusion…More precisely, the output of task
k
’s gating network is formulated as:
g
k
x
=
w
k
x
S
k
x
…
w
k
x
is a weighting function…
S
k
x
is a selected matrix composed of all selected vectors including shared experts and task
k
’s specific experts…”
Tang discloses a gating network function which includes a weighting function [the gate weights are applicable to gate functions] and a selection matrix. The selection matrix of the gating network function selects shared experts and/or task-specific experts [an exclusivity mechanism for setting expert models to be exclusively connected to one task]. This corresponds to the claimed language because the gating network function can select experts such as task-specific experts. In other words, selected experts are exclusively connected.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Ma with Tang in order to create a more robust model. “In this paper, we propose a novel MTL model called Progressive Layered Extraction (PLE), which separates task-sharing and task specific parameters explicitly and introduces an innovative progressive routing manner to avoid the negative transfer and seesaw phenomenon, and achieve more efficient information sharing and joint representation learning. Offline and online experiment results on the industrial dataset and public benchmark datasets show significant and consistent improvements of PLE over SOTA MTL models” (Tang, p. 278).
Regarding Claim 6:
As discussed above Ma in view of Tang teach [the] system as claimed in claim 1, and Tang further discloses:
wherein the gate weights are applicable to gate functions comprising an exclusion mechanism for setting expert models to be connected such that they are excluded from some tasks
Tang, p. 273, “In CGC, shared experts and task-specific experts are combined through a gating network for selective fusion…More precisely, the output of task
k
’s gating network is formulated as:
g
k
x
=
w
k
x
S
k
x
…
w
k
x
is a weighting function…
S
k
x
is a selected matrix composed of all selected vectors including shared experts and task
k
’s specific experts…”
Tang discloses a gating network function which includes a weighting function [the gate weights are applicable to gate functions] and a selection matrix. The selection matrix of the gating network function selects shared experts and/or task-specific experts [an exclusion mechanism for setting expert models to be connected such that they are excluded from some tasks]. This corresponds to the claimed language because the gating network function can select certain experts. In other words, non-selected experts are excluded.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Ma with Tang in order to create a more robust model. “In this paper, we propose a novel MTL model called Progressive Layered Extraction (PLE), which separates task-sharing and task specific parameters explicitly and introduces an innovative progressive routing manner to avoid the negative transfer and seesaw phenomenon, and achieve more efficient information sharing and joint representation learning. Offline and online experiment results on the industrial dataset and public benchmark datasets show significant and consistent improvements of PLE over SOTA MTL models” (Tang, p. 278).
Regarding Claim 7:
As discussed above Ma in view of Tang teach [the] system as claimed in claim 1, and Ma further discloses:
wherein the steps for each task are repeated for different inputs until a stopping criterion is satisfied
Ma, p. 1933, “repeat step (1) and (2) hundreds of times with datasets generated independently but control the list of task correlation scores and the hyper-parameters the same”
Ma discloses each tasks are repeated for the dataset for hundreds of times [the steps for each task are repeated for different inputs until a stopping criterion is satisfied].
Regarding Claims 11-17:
Claims 11-17 correspond to claims 1-7 and are rejected for at least the same reasons as given in the rejections of claims 1-7. In particular, 11:1, 12:2, 13:3, 14:4, 15:5, 16:6, 17:7.
Claims 8-10 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Ma in view of Tang as applied to claims 1 and 11 above, respectively, and further in view of Wang et al. (“Graph-Driven Generative Models for Heterogeneous Multi-Task Learning”), hereinafter Wang.
Regarding Claim 8:
As discussed above Ma in view of Tang teach [the] system as claimed in claim 1, and Ma further discloses:
wherein the at least one processor is configured to perform…optimization to balance the tasks on a gradient level
Ma, p. 1937, “All models are optimized using mini-batch Stochastic Gradient Descent (SGD) with batch size 1024”
Ma discloses optimization using gradient level [optimization to balance the tasks on a gradient level].
Ma in view of Tang do not explicitly disclose:
…a two-step optimization…
However, in the same field, analogous art Wang teaches:
…a two-step optimization…
Wang, p. 979, “The GCN serves as a generator of latent representations for the sub-graphs, while the VAEs are specified to address the different tasks. The model is then optimized jointly over the objectives for all tasks to encourage the GCN to produce representations that can be used simultaneously by all of them”
p. 985, “Compared with only performing topic modeling, i.e., GD-VAE (T), considering more tasks brings improvements, and the proposed GD-VAE achieves the best performance”
Wang discloses first step and second steps are VAEs and GCN train separately and optimized.
Ma, Tang, Wang, and the instant application are analogous art because they are all directed to multi-task models.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Ma and Tang with Wang because of the following, learn heterogeneous tasks within framework “easily extended to new tasks by specifying the corresponding generative processes” and “better leverage information across tasks, and achieve state-of-the-art results on clinical topic modeling, procedure recommendation, and admission-type prediction” (Wang, p. 986).
Regarding Claim 9:
As discussed above Ma and Tang in view of Wang teach [the] system as claimed in claim 8, and Wang further discloses:
wherein the two-step optimization comprises a modified model-agnostic meta-learning where task specific layers are not frozen during an intermediate update
Wang, p. 980, FIG. 1, “Each task operates on a different sub-graph from the admission graph. The shared GCN (fφ) learns embeddings for ICD codes and admissions, and the embeddings pass through task-specific VAEs”
p. 980, “At test time, the GCN is used to represent sub-graphs, i.e., collections of shared ICD codes, specialized admissions and their interactions, that feed into different task-specific VAEs. We test our model on the three tasks described above. Experimental results show that the jointly learned representation for the admission graph indeed improves the performance of all tasks relative to the individual task model”
Wang discloses update VAEs and GCN corresponding to two step optimization, both task modules are updated corresponding to task specific layers are not frozen.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Ma and Tang with Wang because of the following, learn heterogeneous tasks within framework “easily extended to new tasks by specifying the corresponding generative processes” and “better leverage information across tasks, and achieve state-of-the-art results on clinical topic modeling, procedure recommendation, and admission-type prediction” (Wang, p. 986).
Regarding Claim 10:
As discussed above Ma in view of Tang teach [the] system as claimed in claim 1, but do not explicitly disclose:
wherein at least one individual expert model comprises another multi-task learning network
However, in the same field, analogous art Wang teaches:
wherein at least one individual expert model comprises another multi-task learning network
Wang, p. 980. “For the k-th task, pθk (·) represents a generative model (i.e., a stochastic decoder) with parameters θk, and p(zk) is the prior distribution for latent code zk. The corresponding inference network for zk consists of two parts: (i) a deterministic encoder fφ(·) shared across all tasks to encode each xk into xˆk = fφ(xk) independently; and (ii) an encoder with parameters ψk to stochastically map xˆk into latent code zk”
Wang teaches decoder (expert) generate multiple outputs (corresponding to expert comprises multiple task).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Ma and Tang with Wang because of the following, learn heterogeneous tasks within framework “easily extended to new tasks by specifying the corresponding generative processes” and “better leverage information across tasks, and achieve state-of-the-art results on clinical topic modeling, procedure recommendation, and admission-type prediction” (Wang, p. 986).
Regarding Claims 18-20:
Claims 18-20 correspond to claims 8-10 and are rejected for at least the same reasons as given in the rejections of claims 8-10. In particular, 18:8, 19:9, 20:10.
Response to Arguments
Applicant's arguments filed July 27, 2026 (“Remarks”), have been fully considered and are not persuasive..
35 U.S.C. § 101:
Remarks, pp. 7-8. Applicant’s arguments with respect to independent claims 1 and 11 have been fully considered and are persuasive. The rejections of claims 1-20 have been withdrawn. In particular, the technological improvement of exclusion and exclusivity provided by paras. [0031]-[0033] of the specification are described and reflected within the claim. Furthermore, when the claim is considered as a whole, the claims are directed towards improving machine-learning itself, as supported by the specification, rather than being directed towards the mathematical concept and mental process.
35 U.S.C. § 103:
Remarks, pp. 8-10. Applicant argues Ma, Wang, and Tang fail to disclose the recited architecture. Examiner respectfully disagrees. As shown by Applicant, FIG. 1D of the present application depicts an Expert 0 that is exclusively connected to Task A. As discussed under section 103, Tang also explicitly depicts in FIG. 4 Expert Models A that are exclusively connected to a Task A. Therefore, Tang discloses the recited architecture.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to STEVEN PHUNG whose telephone number is (703)756-1499. The examiner can normally be reached Monday-Thursday: 9:00AM-4:00PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, KAMRAN AFSHAR can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/STEVEN PHUNG/Examiner, Art Unit 2125
/KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125