Prosecution Insights
Last updated: August 16, 2026
Application No. 18/425,715

PRIVACY-PRESERVING INTERPRETABLE SKILL LEARNING FOR HEALTHCARE DECISION MAKING

Non-Final OA §101§103
Filed
Jan 29, 2024
Priority
Feb 01, 2023 — provisional 63/442,475 +1 more
Examiner
LEE, WILLIAM MICHAEL
Art Unit
3682
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
NEC Laboratories America Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-52.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
17 currently pending
Career history
15
Total Applications
across all art units

Statute-Specific Performance

§101
29.2%
-10.8% vs TC avg
§103
48.6%
+8.6% vs TC avg
§102
2.8%
-37.2% vs TC avg
§112
19.4%
-20.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
DETAILED ACTION The action is in response to the original filing on January 29, 2024. Claims 1-20 are pending and have been considered below. Claims 1 and 20 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-6, 8-16, and 18-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1: Step 1 – The claim is directed to a method: a computer-implemented method for training a healthcare treatment machine learning model… Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mathematical concepts (see MPEP 2106.04(a)(2)(I)) and mental processes (see MPEP 2106.04(a)(2)(III)): aggregating local weights from a plurality of clients to update a set of global weights for an imitation-based skill learning model… aggregating local weights to update a set of global weights is a mathematical calculation. clustering a set of local prototype vectors from the plurality of clients to generate a plurality of clusters… clustering vectors requires mathematical calculations such as calculating the distances or similarities between vectors. selecting representative vectors for the plurality of clusters as a set of global prototypes… a human can reasonably perform “selecting representative vectors” within the human mind or with the aid of a pen and paper, which is a mental process. determining client-specific prototype vectors for the plurality of clients based on the representative vectors… a human can reasonably perform “determining client-specific prototype vectors… based on the representative vectors” within the human mind or with the aid of a pen and paper, which is a mental process. Step 2A, Prong 2 – The following limitations are additional elements that fail to integrate the judicial exception into a practical application: distributing the updated set of global weights and the client-specific prototype vectors to the plurality of clients… distributing data to client devices is data outputting (see MPEP 2106.05(g)). Step 2B – These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they only amount to data outputting (MPEP 2106.05(g)) without significantly more. These limitations, taken either alone or in combination, fail to provide an inventive concept. Thus, the claim is not patent eligible. Claims 2-6 and 8-10 recite limitations which further narrow the abstract ideas of claim 1 by specifying more details of the mathematical concepts and mental processes that occur: Regarding claim 2, specifying wherein the set of global weights includes weights of a convolution layer and weights of an imitation learning layer in this manner does not overcome the rejection of claim 1 as modifying “the set of global weights” does not make “aggregating” to not be a mathematical concept. Regarding claim 3, specifying wherein the imitation learning layer implements an action-selection policy based on behavior cloning in this manner does not overcome the rejection of claim 2 as modifying “the imitation learning layer” does not make “aggregating” to not be a mathematical concept. Regarding claim 4, this claim further limits the abstract ideas of claim 1 to be based on a mathematical concept: wherein selecting the representative vectors includes determining respective centroids of the plurality of clusters… determining centroids requires mathematical calculations such as calculating distances between vectors. Regarding claim 5, describing further comprising learning the local weights and the local prototype vectors at the plurality of clients based on initial global weights and initial prototypes is data gathering (see MPEP 2106.05(g)). Regarding claim 6, this claim further limits the abstract ideas of claim 5 to be based on a mathematical concept: wherein the learning includes minimizing an objective function that includes an imitation loss and a plurality of regularization losses… minimizing an objective function, for example, a loss function, requires mathematical calculations. Regarding claim 8, specifying wherein the local prototype vectors correspond to treatment actions that can be performed in a medical context in this manner does not overcome the rejection of claim 1 as modifying “the local prototype vectors” does not make “clustering” to not be a mathematical concept. Regarding claim 9, this claim further limits the abstract ideas of claim 8 to be based on a mental process: selecting a treatment action based on a skill predicted by the imitation-based skill learning model, based on the measured sate information… a human can reasonably perform “selecting a treatment action” based on model outputs within the human mind or with the aid of a pen and paper. Furthermore, measuring a patient’s state information and notifying a medical professional of the treatment action to assist the medical professional in decision-making for patient management is data gathering and outputting (see MPEP 2106.05(g)). Regarding claim 10, specifying wherein the treatment action includes an instruction to a treatment system to automatically administer a treatment to a patient in this manner does not overcome the rejection of claim 9 as modifying “the treatment action” does not make “selecting” to not be a mental process. Regarding claim 11: Step 1 – The claim is directed to a system: a system for training a healthcare treatment machine learning model… Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mathematical concepts (see MPEP 2106.04(a)(2)(I)) and mental processes (see MPEP 2106.04(a)(2)(III)): aggregate local weights from a plurality of clients to update a set of global weights for an imitation-based skill learning model… aggregating local weights to update a set of global weights is a mathematical calculation. cluster a set of local prototype vectors from the plurality of clients to generate a plurality of clusters… clustering vectors requires mathematical calculations such as calculating the distances or similarities between vectors. select representative vectors for the plurality of clusters as a set of global prototypes… a human can reasonably “select representative vectors” within the human mind or with the aid of a pen and paper, which is a mental process. determine client-specific prototype vectors for the plurality of clients based on the representative vectors… a human can reasonably “determine client-specific prototype vectors… based on the representative vectors” within the human mind or with the aid of a pen and paper, which is a mental process. Step 2A, Prong 2 – The following limitations are additional elements that fail to integrate the judicial exception into a practical application: a hardware processor… a processor used as a mere tool to apply an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)). a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to… a memory used as a mere tool to apply an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)). Claims 11-20 are system claims that contains similar limitations to the methods of claims 1-10, respectively. Therefore, claims 11-20 are rejected under substantially the same rationale as claims 1-10, respectively. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4 and 11-14 are rejected under 35 U.S.C. 103 as being unpatentable over Qiao et al. (“A Framework for Multi-Prototype Based Federated Learning: Towards the Edge Intelligence,” 2023, hereinafter Qiao) in view of Liu et al. (“Federated Imitation Learning: A Novel Framework for Cloud Robotic Systems with Heterogeneous Sensor Data,” 2019, hereinafter Liu). Regarding claim 1: Regarding the limitation a computer-implemented method for training a healthcare treatment machine learning model, comprising: aggregating local weights from a plurality of clients to update a set of global weights for an imitation-based skill learning model, Qiao teaches a computer-implemented method for training a healthcare treatment machine learning model (Page 135, Col. 1, Section A, ¶¶1-3 “The essence of the federated training strategy is to train models through the cooperation of distributed clients with data locally… Clients and the edge server collaborate to train a shared model… ω is the model parameters of global model… the local training of each client is to minimize the local loss… The global objective in federated scenarios is to minimize the loss function across heterogeneous clients,” please note that a healthcare treatment machine learning model in the preamble is interpreted to be non-limiting because it recites an intended use or purpose without resulting in a structural or manipulative difference between the claimed invention and prior art; for example, “the federated training strategy… to train models” may be applied to any kind of data, including healthcare treatment data which may include, for example, X-ray images or scanned medical documents), comprising: aggregating local weights from a plurality of clients to update a set of global weights for a… learning model (Col. 2, Section III, ¶1 “In each global iteration, the server plays the role of receiving and averaging the model parameters from clients, and sending the aggregated model parameters back to local clients for the next iteration,” Page 137, Col. 1, Algorithm 1, Line 14 depicts aggregating local weights from a plurality of clients to update a set of global weights for a… learning model). However, Qiao fails to teach an imitation-based skill learning model. Liu, in the same field of endeavor, teaches an imitation-based skill learning model (Page 2, Col. 2, Section B, ¶1 “Local robots learn skills through imitation learning and the cloud server fuses knowledge. We develop a federated learning algorithm to fuse private models into the shared model in the cloud. With the shared model, the cloud server is capable of generating guide models corresponding to requests of local robots. After that, the local robots perform transfer learning based on the guide model,” Page 4, Col. 2, Section D, ¶1 “Layer Transfer means that some layers in the model trained by source data are copied directly and the remaining layer are trained by target data… In image recognition, we usually copy the front layers and retrain the back layers,” Page 5, Col. 1, Fig. 4 and ¶1 “As presented in Fig. 4… we use front layers as feature extractors in the case of imitation learning. The decision model in the cloud can be used as the initial model for local training… In the training of local robots, the feature extraction layer is frozen and only the full connection layers are trained”). Qiao teaches clustering a set of local prototype vectors from the plurality of clients to generate a plurality of clusters (Page 135, Col. 2, Col. 2, Section B, ¶1 “The process of prototype learning is to calculate the mean of embedding vectors as the prototype of each class,” Section III, ¶1 “In the last global iteration, client i and j with heterogeneous datasets Di and Dj transmit not only their model parameters, but also their individual multiple weighted prototypes based on clustering algorithm through feature extractor layers,” wherein “the last global iteration” implies multiple iterations of training or aggregating local weights… has already occurred, Page 136, Col. 1, Fig. 2, Caption: “The overview of the proposed multi-prototype based FL framework (here illustrated with k = 2 prototypes). In T-1 global iterations, clients communicate with the server as typical federated training does. In the T-th round, clients also transmit their computed weighted prototypes to the server,” Section A, ¶¶1-2 “vij is the embedding space of the j-th class. For calculating multiple prototypes of each class, we iteratively cluster vij into k clusters based on k-means algorithm,” wherein “embedding vectors” to be clustered within “the embedding space of the j-th class” encompass a set of local prototype vectors from the plurality of clients, Col. 2, ¶1 “the multiple prototypes of j-th class can be defined as follows: PNG media_image1.png 52 538 media_image1.png Greyscale where uki, j is the output of Clustering(.) of i-th client belonging to j-th class and k is the number of clusters of each class… our clustering algorithm is conducted on the client-side,” Section B, ¶1 “Before the end of global iterations, the clustering results of clients as calculated in (6) are transmitted to the server”). Qiao further teaches selecting representative vectors for the plurality of clusters as a set of global prototypes (Page 136, Col. 2, ¶1 “Following the standard iteration of k-means, we randomly select k centroids in the first iteration. Then, we calculate centroids for each sample in Di,j should belong to. Finally, we recalculate the centroids and repeat the iteration until the centroids do not change or change very little… local prototypes are defined to be equivalent to the centroids of k-means,” Section B, ¶1 “the clustering results of clients as calculated in (6) are transmitted to the server… the distribution of data among clients is usually heterogeneous, which means that clients may have different sample sizes. The more instances in one cluster are, the higher the weight is, and vice versa, which can make the prototypes more representative… a weighted average of these multiple prototypes (centroids of k-means) can be performed, and then aggregated by class at the server,” wherein the “multiple prototypes” encompass representative vectors for the plurality of clusters). Qiao further teaches determining client-specific prototype vectors for the plurality of clients based on the representative vectors (Page 136, Col. 1, ¶1 “the goal of this paper is to use the aggregated weighted prototypes learned by federated training process for model inferences. Formally, PNG media_image2.png 58 557 media_image2.png Greyscale where c is a local prototype of one client, Uj is defined to aggregate the weighted prototypes over all clients belonging to j-th class, and ui denotes one instance in corresponding aggregated prototypes set U-j,” Col. 2, Section B, ¶1 “a weighted average of these multiple prototypes (centroids of k-means) can be performed, and then aggregated by class at the server, as mentioned in (4), which can be further formulated as follows: PNG media_image3.png 97 547 media_image3.png Greyscale where Di,j is the size of instances belonging to class j, and Dki,j is the size of instances belonging to the k-th cluster of class j,” wherein “the aggregated weighted prototypes” aggregated “over all clients belonging to j-th class” which are sent back to the clients encompass client-specific prototype vectors for the plurality of clients). Qiao further teaches and distributing the updated set of global weights and the client-specific prototype vectors to the plurality of clients (Page 135, Col. 1, Section A, ¶¶2-3 “Clients and the edge server collaborate to train a shared model… for each client… the local training of each client is to minimize the local loss… The global objective in federated learning scenarios is to minimize the loss function across heterogeneous clients,” Col. 2, Section III, ¶1 “the server sends both the latest model parameters and aggregated prototypes back to all clients for model inferences,” Page 137, Col. 1, Lines 5 and 14-16 depict distributing the updated set of global weights… to the plurality of clients). Qiao and Liu are analogous art to the claimed invention as both are in the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the imitation-based skill learning model of Liu with the methodology of Qiao. The motivation to do so is to “speed up local training and increase the accuracy of local [devices]” (Liu, Page 5, Col. 1, ¶1). Regarding claim 2, Qiao in view of Liu teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated). Regarding the limitation wherein the set of global weights includes weights of a convolution layer and weights of an imitation learning layer, Qiao teaches wherein the set of global weights includes weights of a convolution layer (Page 137, Col. 1, Section A, ¶3 “For local models, we adopt… A 4-layer CNN network with 2 convolution layers,” while only “local models” are described as having “convolution layers,” it is implicit within the federated learning context that the weights of the local “convolution layers” are aggregated to update the set of global weights, hence wherein the set of global weights includes weights of a convolution layer is implicit). However, Qiao fails to teach and weights of an imitation learning layer. Liu teaches and weights of an imitation learning layer (Page 1, Col. 2, Fig. 1 – “Imitation learning with guides from the cloud” and Caption: “in this work, FIL enables the bottom right robot not only acquires skills by training data, but also gets knowledge from other robots through the cloud robotic system,” Page 2, Col. 2, Section B, ¶1 “Local robots learn skills through imitation learning and the cloud server fuses knowledge. We develop a federated learning algorithm to fuse private models into the shared model in the cloud. With the shared model, the cloud server is capable of generating guide models corresponding to requests of local robots. After that, the local robots perform transfer learning based on the guide model,” Page 5, Col. 1, Fig. 4 and ¶1 “we use front layers as feature extractors in the case of imitation learning. The decision model in the cloud can be used as the initial model for local training… cloud model can play a guiding role… In the training of local robots, the feature extraction layer is frozen and only the full connection layers are trained… it can also adjust the relevant parameters in the process of back propagation,” Col. 2, ¶1 “The policy network mainly consists of convolution layers and fully connected layers,” Page 3, Col. 2, Algorithm 1 depicts fusing weights of an imitation learning layer or the adjusted weights of the “fully connected layers” of “local robots” trained using “imitation learning” at the central “cloud server”). Qiao and Liu are analogous art to the claimed invention as both are in the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the weights of an imitation learning layer of Liu with the set of global weights including weights of a convolution layer of Qiao. The motivation to do so is to “speed up local training and increase the accuracy of local [devices]” (Liu, Page 5, Col. 1, ¶1). Regarding claim 3, Qiao in view of Liu teaches the method of claim 2 (and thus the rejection of claim 2 is incorporated). Qiao fails to teach wherein the imitation learning layer implements an action-selection based policy based on behavior cloning. However, Liu teaches this limitation (Page 2, Col. 2, Section A, ¶1 “Local robots acquire knowledge through imitation learning in FIL. Imitation learning is commonly posed as either behavioral cloning… or as inverse reinforcement learning… The knowledge acquiring approach used in FIL of local robots belongs to behavioral cloning, which focuses on learning the experts policy using supervised learning… Given demonstrations of robots, we will divide these into state-action pairs,” Page 3, Col. 1, Fig. 2 and ¶1 “The three agents use three different types of dataset separately. Datasets of local robots are labeled but will not be sent to the cloud. Three different policy models will be obtained by local training. RGB images will be trained by Agent A, and a private policy model (Private Model A) will be obtained. Actions of robots will be determined by the output which might be some actions or parameters. Similar processes occur in agent B and agent C. Outputs of the three private models are with the same types. The inputs to each model are different… Then the parameters of all three models will be uploaded to the cloud and fused there… the cloud will be capable of generating guide models for different types of input. When a local robot requests a service, the cloud will provide a guide model in correspondence with the type of sensor data,” wherein providing “guide models for different types of input” based on local model output such as “actions or parameters” encompasses an action-selection based policy based on behavior cloning used to train or guide the imitation learning layer through “imitation learning”). Qiao and Liu are analogous art to the claimed invention as both are in the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the action-selection based policy based on behavior cloning of Liu with the methodology of Qiao. The motivation to do so is to “speed up local training and increase the accuracy of local [devices]” (Liu, Page 5, Col. 1, ¶1). Regarding claim 4, Qiao in view of Liu teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated). Qiao further teaches wherein selecting the representative vectors includes determining respective centroids of the plurality of clusters (Page 136, Col. 2, ¶1 “we randomly select k centroids in the first iteration. Then, we calculate centroids for each sample in Di,j should belong to. Finally, we recalculate the centroids and repeat the iteration until the centroids do not change or change very little… local prototypes are defined to be equivalent to the centroids of k-means,” Section B, ¶1 “a weighted average of these multiple prototypes (centroids of k-means) can be performed, and then aggregated by class at the server). Regarding claim 11: Qiao teaches a system for training a healthcare treatment machine learning model (Page 135, Col. 1, Section A, ¶¶1-3 as explained above with respect to claim 1), comprising: a hardware processor (Abstract: “a novel multiple-prototype based federated learning (MPFed) framework is proposed, in which clients communicate with server as typical federated training,” wherein a “server” used to train a model implies the use of a hardware processor, for example, a CPU or GPU). Qiao further teaches a memory that stores a computer program which, when executed by the hardware processor (Page 136, Col. 2, Section C, ¶1 “Further details about model training and inference can be seen in Algorithm 1,” Page 137, Algorithm 1, wherein “Algorithm 1,” or a computer program, is implied to be stored on a memory and is executed by the hardware processor, for example a CPU or GPU, to initiate “model training and inference”), causes the hardware processor to: aggregate local weights (Col. 2, Section III, ¶1, Page 137, Col. 1, Algorithm 1, Line 14 as explained above with respect to claim 1)… Claims 11-14 are system claims that contains similar limitations to the methods of claims 1-4, respectively. Therefore, claims 11-14 are rejected under substantially the same rationale as claims 1-4, respectively. Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Qiao in view of Liu and further in view of Tan et al. (“FedProto: Federated Prototype Learning Across Heterogeneous Clients,” 2022, hereinafter Tan). Regarding claim 5, Qiao in view of Liu teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated). Regarding the limitation further comprising learning the local weights and the local prototype vectors at the plurality of clients based on initial global weights and initial prototypes, Qiao teaches further comprising learning the local weights… at the plurality of clients based on initial global weights (Page 137, Col. 1, Algorithm 1, Lines 2, 5, 7, 11 depicts learning the local weights or “Local model updates” at the plurality of clients based on initial global weights or “global model wt” during training iteration t after initializing “w0”). However, the combination of Qiao and Liu fails to teach learning the local weights and the local prototype vectors at the plurality of clients based on initial global weights and initial prototypes. Tan, in the same field of endeavor, teaches learning the local prototype vectors at the plurality of clients based on initial prototypes (Page 4, Col. 2, Algorithm 1, Lines 1 and 7 depict learning the local prototype vectors or “Ci” for “each client i” based on initial prototypes or “global prototype set {C̄(j)}” initialized at the server). Qiao and Tan are analogous art to the claimed invention as both are in the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the algorithms of Liu and Qiao. The motivation to do so is “to minimize the classification error on the local data” (Tan, Abstract). Claim 15 is a system claim that contains similar limitations to the method of claim 5. Therefore, claim 15 is rejected under substantially the same rationale as claim 5. Claims 6-7 and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Qiao in view of Liu and further in view of Tan, and further in view of Ming et al. (“Interpretable and Steerable Sequence Leaning via Prototypes,” 2019, hereinafter Ming). Regarding claim 6, Qiao in view of Liu and further in view of Tan teaches the method of claim 5 (and thus the rejection of claim 5 is incorporated). Regarding the limitation wherein the learning includes minimizing an objective function that includes an imitation loss and a plurality of regularization losses, Liu teaches an imitation loss (Page 2, Col. 2, Section A, ¶1 “Local robots acquire knowledge through imitation learning in FIL. Imitation learning is commonly posed as either behavioral cloning… or as inverse reinforcement learning… The knowledge acquiring approach used in FIL of local robots belongs to behavioral cloning, which focuses on learning the experts policy using supervised learning… Given demonstrations of robots, we will divide these into state-action pairs. We treat these pairs as i.d.d. examples and finally, we apply supervised learning,” Page 4, Col. 1, ¶1 “Formula (1) … summarized the whole process: PNG media_image4.png 66 431 media_image4.png Greyscale In the Formula (1), Di… is the dataset of the local robot i. L represents the loss function. θ represents parameters of models. xi are original data and yi are labels,” wherein a “loss function” within an “imitation learning” context through “behavioral cloning… using supervised learning” encompasses an imitation loss when given its broadest reasonable interpretation). However, the combination of Qiao, Liu, and Tan fails to teach wherein the learning includes minimizing an objective function that includes… and a plurality of regularization losses. Ming, in the same field of endeavor, teaches wherein the learning includes minimizing an objective function that includes a cross entropy loss and a plurality of regularization losses (Page 3, Col. 2, Section 3.2, ¶1 “Our goal is to learn a ProSeNet that is both accurate and interpretable. For accuracy, we minimize the cross-entropy loss on a training set: CE(Θ, D)… where Θ is the set of all trainable parameters in the model,” Page 4, Col. 1, ¶2 “Full objective: To summarize, the loss we are minimizing is: PNG media_image5.png 64 530 media_image5.png Greyscale where λc, λe, λd and λl1 are hyperparameters that control the strength of the regularizations,” wherein “the regularizations” depicted, or Rc, Re, and Rd in Equation (1), encompass a plurality of regularization losses). Qiao, Liu, and Ming are analogous art to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the imitation loss of Liu and the objective function that includes a plurality of regularization losses of Ming with the federated learning methodology of Qiao. The motivation to do so is to “speed up local training and increase the accuracy of local [devices]” (Liu, Page 5, Col. 1, ¶1) and to design a model that “selects high-quality prototypes which align well with human knowledge and can be interactively refined for better interpretability without loss of performance” (Ming, Abstract). Regarding claim 7, Qiao in view of Liu and further in view of Tan, and further in view of Ming teaches the method of claim 6 (and thus the rejection of claim 6 is incorporated). Regarding the limitation wherein the plurality of regularization losses include a loss that regularizes a segment representation from the imitation-based skill learning model to be as adjacent to a closest prototype as possible, a loss that reverse-regularizes prototype vectors to be as similar to a segment representation as possible, and a loss that enforces a diverse structure of learnable parameterized prototype vectors to avoid redundancy and to improve generalizability of resulting prototypes, Liu teaches the imitation-based skill learning model (Page 2, Col. 2, Section B, ¶1, Page 4, Col. 2, Section D, ¶1, Page 5, Col. 1, Fig. 4 and ¶1, see claim 1). However, Liu fails to teach wherein the plurality of regularization losses include a loss that regularizes a segment representation from the imitation-based skill learning model to be as adjacent to a closest prototype as possible, a loss that reverse-regularizes prototype vectors to be as similar to a segment representation as possible, and a loss that enforces a diverse structure of learnable parameterized prototype vectors to avoid redundancy and to improve generalizability of resulting prototypes, Liu teaches the imitation-based skill learning model. Ming teaches wherein the plurality of regularization losses include a loss that regularizes a segment representation from a model (Page 3, Col. 1, Section 3.1, Fig. 1 and ¶1 “We aim to learn representative prototype sequences (not necessarily exist in the training data) that can be used as classification references and ana logical explanations. For a new input sequence, its similarities with each representative sequences are measured in the learned latent space. Then, the prediction of the new instance can be derived and explained by its similar prototype sequences,” Col. 2, ¶2 “For a given input sequence… the sequence encoder r maps the entire sequence into a single embedding vector with fixed length”) to be as adjacent to a closest prototype as possible (Page 4, Col. 2, ¶1 “To improve interpretability, Li et al. also proposed… the clustering regularization Rc… Rc encourages a clustering structure in the latent space by minimizing the squared distance between an encoded instance and its closest prototype: PNG media_image6.png 74 427 media_image6.png Greyscale where X is the set of all sequences in the training set D,” wherein “an encoded instance” encompasses a segment representation), a loss that reverse-regularizes prototype vectors to be as similar to a segment representation as possible (Page 4, Col. 2, ¶1 “The evidence regularization Re encourages each prototype vector to be as close to an encoded instance as possible: PNG media_image7.png 95 417 media_image7.png Greyscale wherein an “evidence regularization” that encourages prototype vectors to “be as close to an encoded instance as possible” instead of the other way around, for example, as demonstrated by the “clustering regularization” described above, encompasses a loss that reverse-regularizes…), and a loss that enforces a diverse structure of learnable parameterized prototype vectors to avoid redundancy and to improve generalizability of resulting prototypes (Page 3, Col. 2, Section 3.2, ¶2 “Diversity… It would be confusing to have multiple similar prototypes in the explanations and also inefficient in utilizing model parameters. We prevent such phenomenon through a diversity regularization term that penalizes on prototypes that are close to each other: PNG media_image8.png 84 422 media_image8.png Greyscale Where dmin is a threshold that classifies whether two prototypes are close or not…Rd is a soft regularization that exerts a larger penalty on smaller pairwise distances”). Qiao, Liu, and Ming are analogous art to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the imitation-based skill learning model of Liu and the regularization losses of Ming with the federated learning methodology of Qiao. The motivation to do so is to “speed up local training and increase the accuracy of local [devices]” (Liu, Page 5, Col. 1, ¶1) and to design a model that “selects high-quality prototypes which align well with human knowledge and can be interactively refined for better interpretability without loss of performance” (Ming, Abstract). Claims 16-17 are system claims that contains similar limitations to the methods of claims 6-7, respectively. Therefore, claims 16-17 are rejected under substantially the same rationale as claims 6-7, respectively. Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Qiao in view of Liu and further in view of Kleinerman et al. (“Treatment selection using prototyping in latent-space with application to depression treatment,” 2021, hereinafter Kleinerman). Regarding claim 8, Qiao in view of Liu teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated). Regarding the limitation wherein the local prototype vectors correspond to treatment actions that can be performed in a medical context, Qiao teaches the local prototype vectors (Page 135, Col. 2, Col. 2, Section B, ¶1, Page 136, Col. 1, Section A, ¶¶1-2 as explained above with respect to claim 1). However, the combination of Qiao and Liu fails to teach wherein the local prototype vectors correspond to treatment actions that can be performed in a medical context. Kleinerman, in the same field of endeavor, teaches wherein prototype vectors correspond to treatment actions that can be performed in a medical context (Page 2, ¶3 “we propose a novel deep learning-based approach… by simultaneously identifying prototypes (sub groups) of patients as well as approximating outcome prediction in a personalized manner. More specifically, our approach aims at finding “actionable” prototypes, meaning that they differ not only in their characteristics but, importantly, in their expected responses to the available courses of treatment. Our approach uses a novel deep-learning architecture and a multifaceted loss function which balances between the accuracy of the prediction on the individual level and the cohesiveness of the identified prototypes in terms of predicted treatment outcomes,” Page 2, ¶2 “our goal is to find meaningful prototypes with respect to the treatment outcome”). Qiao and Kleinerman are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the prototypes corresponding to treatment actions of Kleinerman with the local prototype vectors within the federated learning context of Qiao. The motivation to do so is to “increase the interpretability of the network’s results… interpretability can drive physicians trust in an automated treatment selection system” (Kleinerman, Page 17, ¶2). Claim 18 is a system claim that contains similar limitations to the method of claim 8. Therefore, claim 18 is rejected under substantially the same rationale as claim 8. Claims 9-10 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Qiao in view of Liu and further in view of Kleinerman, and further in view of Dalli et al. (US 20220114417 A1, hereinafter Dalli). Regarding claim 9, Qiao in view of Liu and further in view of Kleinerman teaches the method of claim 8 (and thus the rejection of claim 8 is incorporated). Kleinerman teaches further comprising: measuring a patient’s state information (Page 3, Section 1.1, ¶3 “Our clinical dataset combines data from several clinical trials, and describes 4754 MDD patients who were treated as part of clinical trials of antidepressant treatment. Each patient in these datasets is described by sociodemographic information and clinical symptomatology at baseline, the treatment they received, and the outcome of the treatment after 12 weeks”). Regarding the limitation selecting a treatment action based on a skill predicted by the imitation-based skill learning model, based on the measured state information, Liu teaches a skill predicted (Page 1, Col. 1, Fig. 1, Caption: “FIL enables the bottom right robot not only acquires skills by training data, but also gets knowledge from other robots through the cloud robotic system,” Page 5, Col. 2, Section B, ¶1 “Evaluation for the shared-model generating method in FIL… the robot (car) will challenge the tasks such as avoiding collisions and making timely turns,” wherein a skill predicted…, when given its broadest reasonable interpretation, encompasses any output produced by the imitation-based skill learning model based on skills it has acquired) by the imitation-based skill learning model (Page 2, Col. 2, Section B, ¶1, Page 4, Col. 2, Section D, ¶1, Page 5, Col. 1, Fig. 4 and ¶1, see claim 1). However, the combination of Qiao and Liu fails to teach selecting a treatment action based on a skill predicted by the imitation-based skill learning model, based on the measured state information. Kleinerman teaches selecting a treatment action based on a probability predicted by a model, based on the measured state information (Page 1, Section 1, ¶1 “At the heart of much [Precision Medicine] research and practice stands the challenge of effective personalization, such as selecting an optimal treatment for each individual patient,” Page 6, Section 3, ¶1 “We are given a data set of N samples D = {(x1, t1, y1)…, (xN, tN, yN)}, where xi describes a patient sampled from a given distribution x and represented as a d dimensional feature vector xi ∈ Rd, ti ∈ T indicates the treatment received by patient x-i received from a finite set of k > 1 treatment options, and yi ∈ Y indicates the observed outcome of the treatment,” ¶4 “The optimal treatment selection policy, π*, assigns t*i for each patient xi such that it maximizes the desired outcome probability (i.e., remission). Formally, for a patient x, PNG media_image9.png 60 521 media_image9.png Greyscale Where Pr(r|x, t) is the probability of remission for patient x given treatment t. Naturally, the true probability is unknown,” Pages 6-7, Section 4, ¶1 “our true objective is to “approximate” the optimal policy π* by deriving a treatment selection policy as follows: PNG media_image10.png 49 528 media_image10.png Greyscale Page 7, ¶1 “By approximating the optimal policy we mean that we seek to minimize the following loss function: PNG media_image11.png 65 594 media_image11.png Greyscale ¶2 “However Pr(r|x, t) is unknown. As such, any approximation thereof need not necessarily minimize the above loss. To overcome this hurdle, we propose to approximate Pr(r|x, t) in an unorthodox way such that it would potentially prove more useful for minimizing the above loss indirectly… Our method leverages the assumption that patients may be divided into sub-groups which vary significantly in their reactions to treatments… we implement our approximation approach with a neural network based architecture, that: 1) identifies prototypes of patients; and 2) predicts the remission probability for each patient-treatment pair based on their resemblances to identified prototypes”). Regarding the limitation notifying a medical professional of the treatment action to assist the medical professional in decision-making for patient management, Kleinerman teaches the treatment action to assist the medical professional in decision-making for patient management (Page 2, ¶¶2-3 “machine-assisted treatment selection approaches can be largely classified into one of two paradigms… Algorithms from both treatment selection paradigms have demonstrated significant benefits across a wide range of medical applications. However, in some medical settings, both paradigms suffer from important limitations,” Pages 3-4, Section 1.1, ¶3 “Given that the current standard of treatment for many psychiatric disorders, including MDD, is an educated “trial-and-error” approach, our approach could help bring about a much desired leap forward in treatment effectiveness in terms of increased remission rates and reduced length of the process of finding optimal treatments at the individual patient level,” Page 4, Section 2, ¶¶2-3 “In order to mitigate these limitations we propose a novel model, which we will call Differential Prototypes Neural Network (DPNN for short)… the prototypes are constantly tuned during the training process in order to guarantee that the prototypes will approximate an optimal treatment selection policy,” Page 17, Section 7, ¶4 “We believe that the DPNN approach can enable physicians to gain insights through the learned prototypes and their “resemblance” to each individual patient,” wherein the “machine-assisted treatment selection” is implied to be used to assist the medical professional or “physicians” in decision-making for patient management). However, the combination of Qiao, Liu, and Kleinerman fails to teach notifying a medical professional of the treatment action… Dalli, in the same field of endeavor, teaches notifying a medical professional of a treatment action (Fig. 2 – 904, 908, ¶277 “the EIGS system is effectively acting as a guideline execution engine,” ¶278 “The explainer 908… has determined that the CT and PET scan data together with other relevant explanatory information from the model 904 determines that there is a medium probability of lung cancer cells and assigns a diagnostic code of “D47.09” and a recommendation for a follow-up biopsy procedure,” Fig. 11 – 9141, 9142, ¶279 “The EIGS system may output an explanation 9141 containing the fused image data from the CT and PET scan, together with diagnostic text output, as illustrated in FIG. 3. An exemplary EIGS also outputs a recommendation for a biopsy in its interpretation 9142, which is presented as part of the output in FIG. 3,” Fig. 3 depicts “Referring Physician” and “Electronically Signed By” fields, implying that the outputted “recommendation for a biopsy” is intended for a medical professional to approve and/or administer). Qiao, Liu, Kleinerman, and Dalli are analogous art to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the predicted skill and imitation-based skill learning model of Liu, the patient state measurement and treatment selection of Kleinerman, and the notification of a medical professional of Dalli with the methodology of Qiao. The motivation to do so is to “speed up local training and increase the accuracy of local [devices]” (Liu, Page 5, Col. 1, ¶1), to “increase the interpretability of the network’s results… interpretability can drive physicians trust in an automated treatment selection system” (Kleinerman, Page 17, ¶2), and “to enable practical and useful actionable explanations to be generated” (Dalli, Abstract). Regarding claim 10, Qiao in view of Liu and further in view of Kleinerman, and further in view of Dalli teaches the method of claim 9 (and thus the rejection of claim 9 is incorporated). Dalli further teaches wherein the treatment action includes an instruction to a treatment system to automatically administer a treatment to a patient (¶273 “The EIGS system may control a robotic needle biopsy system that accurately performs biopsies with minimal invasion for patients,” ¶279 “Hidden from the output… is information to be used by the robotic needle biopsy system to execute the recommended procedure, after the appropriate authorization has been given by the system user and the appropriate consent has been received from the patient”). Qiao and Dalli are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the automatic treatment instruction of Dalli with the methodology Qiao. The motivation to do so is to “to enable practical and useful actionable explanations to be generated” (Dalli, Abstract). Claims 19-20 are system claims that contains similar limitations to the methods of claims 9-10, respectively. Therefore, claims 19-20 are rejected under substantially the same rationale as claims 9-10, respectively. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM M LEE whose telephone number is (571)272-4761. The examiner can normally be reached Mon-Fri. 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WILLIAM M LEE/ Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Jan 29, 2024
Application Filed
Jul 27, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month