Prosecution Insights
Last updated: August 17, 2026
Application No. 17/509,507

FEDERATED LEARNING DATA SOURCE SELECTION

Non-Final OA §101§103§112
Filed
Oct 25, 2021
Examiner
MAC, GARY
Art Unit
2127
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
4 (Non-Final)
43%
Grant Probability
Moderate
4-5
OA Rounds
0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 43% of resolved cases
43%
Career Allowance Rate
9 granted / 21 resolved
-12.1% vs TC avg
Strong +44% interview lift
Without
With
+43.6%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
14 currently pending
Career history
52
Total Applications
across all art units

Statute-Specific Performance

§101
38.0%
-2.0% vs TC avg
§103
42.4%
+2.4% vs TC avg
§102
6.5%
-33.5% vs TC avg
§112
11.6%
-28.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 21 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s argument filed 02/06/2026 have been fully considered but they are not persuasive. Applicant’s Argument: On page 11-13 of Applicant’s response, applicant states that amended claim does not recite a mental process and the claims recite an improvement to how the machine learning model itself operates. Specifically, claim 1 recites a specific federated-learning control loop that improves how the model is trained and maintained. Examiner’s Response: Applicant’s argument is not persuasive. The claim as a whole is directed to an abstract idea of a mental process that can be performed in the human mind. The process of removing at least one data source based on the influential score can be performed in the human mind. The claim limitation “wherein removing the at least one data source stops data updates from the central server to the at least one data source” merely discloses an outcome of the process of removing the data source. An important consideration in determining whether a claim improves technology is the extent to which the claim covers a particular solution to a problem or a particular way to achieve a desired outcome, as opposed to merely claiming the idea of a solution or outcome (see MPEP 2106.05(a)). The amended claims do not provide sufficient details to describe any technological improvement. If the specifications explicitly set forth an improvement but in a conclusory manner (see MPEP 2106.04(d)(1): a bare assertion of an improvement without the detail necessary to be apparent to a person of ordinary skill in the art), the examiner should not determine the claim improves technology. Claim 1 does not disclose any details regarding how the training is performed or the determination of the influential score other than a high level of generality. Applicant’s Argument: On page 14-17 of Applicant’s response, applicant states that Mars fails to disclose the server-to-client model-update cutoff. Mars does not distribute model updates to clients. Li also fails to disclose stopping updates to the source. Mars and Li also fail to disclose the claimed retraining process. Examiner’s Response: Applicant’s argument is not persuasive. Claim 1 recites “wherein removing the at least one data source stops data updates from the central server to the at least one data source”. Data updates and model updates are both recited in claim 1 and there should be a clear distinction between the two terms. Data updates would be any information related to the training data. Model updates would be any information related to the machine learning model such as updated model parameters. Claim 1 recites a central server receiving training data from a plurality of data sources. The claim limitation “wherein removing the at least one data source stops data updates from the central server to the at least one data source” is interpreted to mean that the central server stops sending a request signal for additional training data to the removed data source. Mars (col. 12, lines 10-13) teaches the server may can signal the specific external training data source to stop transmitting machine learning data. Therefore, Mars teaches a server stops data updates to a data source because the server stops sending a request to the data source to obtain additional training data. The claims do not disclose any details of model updates between the server and data sources. Mars discloses a method for intelligently training a machine learning model. Mars (col. 9, lines 5-32) discloses a monitoring module to detect predefined triggering events or conditions for training updates. A training update may be triggered when the accuracy of the model does not satisfy a threshold. It is implied that a model may require multiple training updates to satisfy the one or more predefined conditions. Thus, Mars teaches the retraining limitation as recited in the claims. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-9 and 11-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 1, 11, and 12 recite “the retraining including collecting, at the central server, model updates received from the subset of the plurality of data sources, updating the training dataset based on the collected model updates”. It is not clear what constitutes as model updates. The claim recites a plurality of data sources that contains training data and the training data is provided to a central server. Although, the claim recites a federated learning system, the claims do not provide any details whether the data sources contain a model and what are the model updates from the data sources. It is not clear how the model updates are used to update the training dataset. Examiner interprets the claims to mean that the model updates are additional training data to improve the model on the central server. Thus, updating the training dataset means that the new training data are added into the existing training dataset to create an updated dataset to train the model. Claims 2-9 and 13-20 are dependent claims to independent claims 1 and 12. Thus, the dependent claims are rejected on the same basis as the parent claims. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-9, and 11-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding Claim 1: Subject Matter Eligibility Analysis Step 1: Claim 1 recites “A method, comprising” and is thus a process, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: “testing,” (a mental process that can be performed in the human mind with the aid of pen and paper, i.e. judgement) “based on the influential score of each of the plurality of data sources, removing, ” (a mental process that can be performed in the human mind, i.e. judgement; selecting a data source with low influential score to be removed) “, updating the training dataset based on the collected model updates, updating the machine-learning model, ” (a mental process that can be performed in the human mind, i.e. judgement) Claim 1 therefore recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: "training, at a central server of a federated learning system, a machine-learning model utilizing a training dataset including data received from a plurality of data sources connected to the central server, ” (mere instructions to apply the exception using a generic computer component - see MPEP 2106.05(f)) "” (merely specifies a particular technological environment in which the abstract idea is to take place, ie. a field of use, and thus does not integrate the abstract idea into a practical application nor cannot provide significantly more than the abstract idea itself - see MPEP 2106.05(h)) “testing, at the central server, the accuracy of the machine-learning model against the validation dataset ” (mere instructions to apply the exception using a generic computer component - see MPEP 2106.05(f)) “retraining, at the central server, the machine-learning model utilizing a reduced training dataset including the data from the subset of the plurality of data sources, the retraining including wherein the reduced training dataset improves the accuracy of the machine-learning model in predicting the plurality of annotated datapoints of the validation dataset, while utilizing less training data” (mere instructions to apply the exception using a generic computer component - see MPEP 2106.05(f)) “ collecting, at the central server, model updates received from the subset of the plurality of data sources” (This step is directed to data gathering, which is understood to be insignificant extra solution activity - see MPEP 2106.05(g)) The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are mere insignificant extra solution activity in combination of generic computer functions being implemented with generic computer elements in a high level of generality to perform the disclosed abstract idea above. Therefore, Claim 1 is directed to the abstract idea. Subject Matter Eligibility Analysis Step 2B: "training, at a central server of a federated learning system, a machine-learning model utilizing a training dataset including data received from a plurality of data sources connected to the central server, ” (mere instructions to apply the exception using a generic computer component - see MPEP 2106.05(f)) "” (merely specifies a particular technological environment in which the abstract idea is to take place, ie. a field of use, and thus does not integrate the abstract idea into a practical application nor cannot provide significantly more than the abstract idea itself - see MPEP 2106.05(h)) “testing, at the central server, the accuracy of the machine-learning model against the validation dataset ” (mere instructions to apply the exception using a generic computer component - see MPEP 2106.05(f)) “retraining, at the central server, the machine-learning model utilizing a reduced training dataset including the data from the subset of the plurality of data sources, the retraining including wherein the reduced training dataset improves the accuracy of the machine-learning model in predicting the plurality of annotated datapoints of the validation dataset, while utilizing less training data” (mere instructions to apply the exception using a generic computer component - see MPEP 2106.05(f)) “ collecting, at the central server, model updates received from the subset of the plurality of data sources” (This step is directed to transmitting or receiving information, which is understood to be insignificant extra solution activity and well understood, routine and conventional activity of transmitting and receiving data as identified by the court - see MPEP 2106.05(d)) The additional elements as disclosed above alone or in combination do not recite significantly more than the abstract idea itself as they are mere insignificant extra solution activity in combination of generic computer functions being implemented with generic computer elements in a high level of generality to perform the disclosed abstract idea above. Therefore, Claim 1 is subject-matter ineligible. Regarding Claim 12: The claim recites an article of manufacture (“A computer program product, comprising”) that performs the method as described in claim 1. Therefore, claim 12 is rejected for the same reasons as disclosed for claim 1. The limitations for additional elements of claim 12 are analyzed below. Subject Matter Eligibility Analysis Step 2A Prong 1: Please see Step 2A Prong 1 analysis of claim 1 Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: “a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code executable by a processor to perform operations comprising” (mere instructions to apply the exception using a generic computer component - see MPEP 2106.05(f)) Regarding Claims 2 and 13: Subject Matter Eligibility Analysis Step 2A Prong 1: “wherein the testing comprises evaluating the respective data of the data source against each annotated datapoint within the validation dataset” (a mental process, i.e. evaluation) Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: None Regarding Claims 3 and 14: Subject Matter Eligibility Analysis Step 2A Prong 1: “wherein the testing comprises generating rankings of the plurality of data sources for each annotated datapoint” (a mental process, i.e. judgement) Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: None Regarding Claims 4 and 15: Subject Matter Eligibility Analysis Step 2A Prong 1: “wherein the testing comprises aggregating the rankings of the plurality of data sources across the plurality of annotated datapoints of the validation dataset” (a mental process, judgment) Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: Regarding Claims 5 and 16: Subject Matter Eligibility Analysis Step 2A Prong 1: “wherein the testing comprises computing the accuracy of the machine-learning model trained with the data source in predicting the plurality of annotated datapoints utilizing an influence function comprising a loss function component, a metrics of the data source, and a gradient of the loss function component, wherein the loss function component and the gradient of the loss function component is computed at the central server and wherein a result of the metrics of the data source is provided to the central server from a corresponding data source” (a mathematical calculation, see par. 57 in the Specification) Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: None Regarding Claims 6 and 17: Subject Matter Eligibility Analysis Step 2A Prong 1: “further comprising constructing a bipartite graph comprising the plurality of annotated datapoints and the plurality of data sources by generating weighted edges between annotated datapoints and data sources based upon the influence of the data source” (a mental process with the aid of pen and paper, i.e. evaluation) Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: None Regarding Claims 7 and 18: Subject Matter Eligibility Analysis Step 2A Prong 1: “further comprising selecting the subset of the plurality of data sources based upon a coverage budget provided by a user, wherein the coverage budget is selected from the group consisting of a first constraint on a number of the plurality of data sources and a second constraint on a number of the plurality of annotated datapoints” (a mental process, i.e. judgement) Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: None Regarding Claims 8 and 19: Subject Matter Eligibility Analysis Step 2A Prong 1: “further comprising selecting a least number of data sources that covers the validation dataset” (a mental process, i.e. judgement) Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: None Regarding Claims 9 and 20: Subject Matter Eligibility Analysis Step 2A Prong 1: None Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: “wherein the respective data comprises a derivative of local data stored at the data source to preserve a privacy of the local data stored at the data source” (merely specifies a particular technological environment in which the abstract idea is to take place, ie. a field of use, and thus does not integrate the abstract idea into a practical application nor cannot provide significantly more than the abstract idea itself - see MPEP 2106.05(h)) Regarding Claim 11: The claim recites a system (“An apparatus, comprising”) that performs the method as described in claim 1. Therefore, claim 11 is rejected for the same reasons as disclosed for claim 1. The limitations for additional elements of claim 11 are analyzed below. Subject Matter Eligibility Analysis Step 2A Prong 1: Please see Step 2A Prong 1 analysis of claim 1 Subject Matter Eligibility Analysis Step 2A Prong 2 & 2B: “at least one processor” (mere instructions to apply the exception using a generic computer component - see MPEP 2106.05(f)) “a computer readable storage medium having computer readable program code embodied therewith and executable by the at least one processor to cause the at least one processor to perform operations comprising” (mere instructions to apply the exception using a generic computer component - see MPEP 2106.05(f)) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 7, 11-15, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Mars (US10296848B1) in view of Li, “Resolving Conflicts in Heterogeneous Data by Truth Discovery and Source Reliability Estimation”. Regarding claim 1, Mars teaches: “A method, comprising: training, at a central server of a federated learning system, a machine-learning model utilizing a training dataset including data received from a plurality of data sources connected to the central server, wherein the central server includes a validation dataset having a plurality of annotated datapoints ”([col. 2, lines 55-61;col. 5, lines 30-53; col. 10, lines 6-19; col. 11, lines 8-26; col. 13, lines 32-62, Figure 2], The ML management console (central server) collects training data from a plurality of external sources. The ML management console comprises a collection of seed samples (validation dataset) that can be used as input data to retrieve a plurality of labeled training samples from external data sources. After the console collects and process the ML training data, the training data is deployed to a ML model for execution. Slot identification engine may generate slot labels for the user input queries and slot identification engine may assign multiple labels to the user input. The seed samples are a type of user input that is received by the model and the user query may have a plurality of slot labels (plurality of annotated datapoints) assigned it the input.) “testing, at the central server, the accuracy of the machine-learning model ” ([col. 9, lines 39-45; col. 11, lines 50-53; col. 12, lines 36-58; col. 13, lines 53-67; col. 14, lines 1-5], The training data may have metadata that identifies which external data source it originated from. The fit score (influential score) for each of the training data samples generally represents how well a given training data samples fits the machine learning model or one or more of the seed training data samples. The system may test the performance of the model and measure one or more operational metrics of the model. When the training data sample is poor or bad, the system may re-evaluate the training data sample set by calculating the fit score for the training data. The model may be a classification model that outputs labels for the input data.) “based on the influential score of each of the plurality of data sources, removing, at the central server, at least one data source of the plurality of data sources from contributing to the training dataset, wherein the at least one data source includes less influence, relative to a subset of the plurality of data sources, on the machine-learning model accurately predicting each annotated datapoint ” ([col. 12, lines 4-58; col. 13, lines 5-31], A threshold may be set for each of the external training data sources to limit the amount of data collected from the data source. After the fit score has been determine for the training data samples, the system may remove training data samples that have a fit score below a certain threshold. A training data source that has a low level of quality may be requested to stop transmitting machine learning training data compared to training data sources that have a higher level of quality.) “retraining, at the central server, the machine-learning model utilizing a reduced training dataset including the data from the subset of the plurality of data sources, the retraining including collecting, at the central server, model updates received from the subset of the plurality of data sources, updating the training dataset based on the collected model updates, updating the machine-learning model, and iterating until a convergence criterion is satisfied, wherein the reduced training dataset improves the accuracy of the machine-learning model in predicting the plurality of annotated datapoints ” ([col. 9, lines 5-45; col. 13, lines 32-67; col. 14, lines 1-5; Figure 2], After a subset of the data is determined from the external data sources, the system loads the training data samples into the ML model to update the configuration of the model. The system continuously receives data samples to update and execute the ML model until the ML model satisfies one or more predetermined conditions, such as maintaining a level of accuracy of 80%. In some embodiments, the system continues to collect training data and updating the ML model for a number of iterations until the conditions are satisfied. The system tests the performance of the machine learning model based on one or more operational metrics to measure the improvement with the updated training data samples. The model may be a classification model that outputs labels for the input data.) Mars does not explicitly disclose an implementation of “a validation dataset having a plurality of annotated datapoints configured to test an accuracy of the machine-learning model”, and “testing ... the accuracy of the machine-learning model against the validation dataset”. However, Li discloses in the same field of endeavor: “training, at a central server of a federated learning system, a machine-learning model utilizing a training dataset including data received from a plurality of data sources connected to the central server, wherein the central server includes a validation dataset having a plurality of annotated datapoints configured to test an accuracy of the machine-learning model” ([pg. 2, col. 1, par. 3; pg. 3, Section 2.2, par. 1-5; pg. 8, Section 3.2.1, par. 2-5; pg. 3, Section 2.1, Table 1 & 2; pg. 4, Section 2.2, Algorithm 1], The framework describes an optimization method for determining source reliability for a plurality of data sources. Weather forecast data is collected from 3 different sources and it is validated against the ground truth dataset (validation dataset). The CRH framework is a machine learning model that consist of an optimization process to minimize an objective function. Algorithm 1 shows the framework of performing multiple iterations to train the model until a convergence criterion (accuracy of the machine learning model) is satisfied.) “testing, at the central server, the accuracy of the machine-learning model against the validation dataset to determine an influential score for each of the plurality of data sources based upon respective data from each of the plurality of data sources included in the training dataset, wherein a respective influential score of a data source indicates an influence of the data source on the machine-learning model accurately ” ([pg. 2, col. 1, par. 3; pg. 3, Section 2.2, par. 1-5; pg. 8, Section 3.2.1, par. 2-5; pg. 3, Section 2.1, Table 1 & 2], The framework consists of minimizing an objective function that contains weights of multi-source input that reflects the reliability degree (influential score) of the sources. The source reliability is validated against a ground truth dataset. The model consists of a loss function that measures the deviation of the source data from the truth. The model performs an optimization process based on a convergence criterion to minimize the weighted deviation from the truths to the multi-source input.) “... wherein the at least one data source includes less influence, relative to a subset of the plurality of data sources, on the machine-learning model accurately ” ([pg. 2, col. 1, par. 3; pg. 3, Section 2.2, par. 1-5; pg. 8, Section 3.2.1, par. 2-5; pg. 3, Section 2.1, Table 1 & 2], Observations made by an unreliable source will be determined to have a lower source weight. A reliable source consists of a high source weight value and provides closer observations to the ground truth. The model determines the deviation between the truth and the observations.) “... training dataset improves the accuracy of the machine-learning model in ” ([pg. 2, col. 1, par. 3; pg. 3, Section 2.2, par. 1-5; pg. 8, Section 3.2.1, par. 2-5; pg. 3, Section 2.1, Table 1 & 2], Observations made by an unreliable source will be determined to have a lower source weight. A reliable source consists of a high source weight value and provides closer observations to the ground truth. The model determines the deviation between the truth and the observations.) It would be obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of “a validation dataset having a plurality of annotated datapoints configured to test an accuracy of the machine-learning model”, and “testing ... the accuracy of the machine-learning model against the validation dataset” from Li into the teaching of Mars. The source reliability determination from Li can be implemented into Mars to generate a score for each data source and ranking the data sources based on the computed score. Doing so can train a ML model to determine an accurate estimation of source reliability by comparing observations to true information (Li, abstract). Regarding claim 12, Mars teaches: Claim 12 recites an article of manufacture (“A computer program product, comprising”) that performs the same process as described in Claim 1. Therefore claim 12 is rejected under the same reasons mention for claim 1. However, claim 12 has additional limitations and the claim elements are addressed below by Mars: “a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code executable by a processor to perform operations comprising” ([col. 14, lines 6-20], A processor executes instructions (program code) from a computer-readable medium to perform the methods described by the reference.) Regarding claims 2 and 13, Mars teaches: “wherein the testing comprises evaluating the respective data of the data source against each annotated datapoint within the validation dataset” ([col. 12, lines 36-58], The system processes the training data to determine a fit score (influential score) to rank each of the training data samples from external sources. The fit score describes how well a training data sample matches the seed samples (validation set) of a training request and the overall data quality in representing the ML model task.) Regarding claims 3 and 14, Mars teaches: “wherein the testing comprises generating rankings of the plurality of data sources for each annotated datapoint” ([col. 11, lines 35-44; col. 12, lines 36-58], The system uses the calculated fit score to generate a ranking for each of the data samples that comes from the plurality of data sources. The ranking represents how valuable a particular data source is for training the ML model. The training data samples from various external data source can be stored into distinct datastores for specific processing of the set of data.) Regarding claims 4 and 15, Mars teaches: “wherein the testing comprises aggregating the rankings of the plurality of data sources across the plurality of annotated datapoints of the validation dataset” ([col. 11, lines 35-44; col. 13, lines 1-4], The training data samples from various external data source can be stored into distinct datastores for specific processing of the set of data. The fit score may be calculated for a set of training samples that originated from the same data source. The system lists (aggregate) the rank based on the fit score in descending or ascending order.) Regarding claims 7 and 18, Mars teaches: “further comprising selecting the subset of the plurality of data sources based upon a coverage budget provided by a user, wherein the coverage budget is selected from a group consisting of a first constraint on a number of the plurality of data sources and a second constraint on a number of the plurality of annotated datapoints”([col. 10, lines 17-27; col. 12, lines 4-35], A user interface is provided where the administrator (user) is able to provide a specific number (first constraint) of external training data sources from a list of external data sources. It is implied that the user may have prior knowledge to gather data from known data sources that can highly benefit the ML model training and limit the data gather from a selected list of external data sources. In the user interface, the administrator may also provide an input value to define the threshold for limiting the number of training data samples (second constraint) receive from the external data sources.) Regarding claim 11, Mars teaches: Claim 11 recites a system (“An apparatus, comprising”) that performs the same process as described in Claim 1. Therefore claim 11 is rejected under the same reasons mention for claim 1. However, claim 11 has additional limitations and the claim elements are addressed below by Mars: “at least one processor” ([col. 14, lines 6-20], A processor executes instructions (program code) from a computer-readable medium to perform the methods described by the reference.) “a computer readable storage medium having computer readable program code embodied therewith and executable by the at least one processor to cause the at least one processor to perform operations comprising” ([col. 14, lines 6-20], A processor executes instructions (program code) from a computer-readable medium to perform the methods described by the reference.) Claims 5 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Mars (US10296848B1) in view of Li, “Resolving Conflicts in Heterogeneous Data by Truth Discovery and Source Reliability Estimation” and Hall (US20230162049A1). Regarding claims 5 and 16, Mars in view of Li teaches: “wherein the testing comprises computing the accuracy of the machine-learning model trained with the data source in predicting the plurality of annotated datapoints utilizing an influence function comprising a loss function component, a metrics of the data sourceMars teaches the system (central server) may compute an operational metric (accuracy) of the ML model in making accurate predictions or classifying labels accurately. The system retrieves training data samples from external data sources. The system validates the performance of the ML model based on the training data samples by determining and comparing operational metrics to determine the reliability of the external data sources. The computation of the operational metrics is performed in a central server system after receiving the training data samples from external data sources. Li further teaches a model that determines source reliability in providing information that is related to the truth and uses a loss function that measures the truth and observation.) Mars in view of Li does not explicitly disclose an implementation of a computing function “a gradient of the loss function component”. However, Hall discloses in the same field of endeavor: “wherein the computing comprises computing an accuracy of the data source in predicting the annotations utilizing an influence function comprising a loss function component, a metrics of the data source component, and a gradient of the loss function component, wherein the loss function component and the gradient of the loss function component is computed at the central server and wherein a result of the metrics of the data source component is provided to the central server from a corresponding data source” ([0169, 0185-0187], The training data for the ML model may come from a variety of data sources. In one embodiment, data may be collected at a central server to perform the method of determining the predictive power of the dataset from a particular data source. A balanced accuracy metric or a log loss function may be used to evaluate the dataset for each data source to determine the accuracy and data quality of a dataset.) It would be obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of a computing function “a gradient of the loss function component” from Hall into the teaching of Mars in view of Li. Doing so can train a global ML model with high quality data and remove the noisy data in the dataset using metrics to determine which data are consistently providing incorrect predictions (Hall, abstract). Claims 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Mars (US10296848B1) in view of Li, “Resolving Conflicts in Heterogeneous Data by Truth Discovery and Source Reliability Estimation” and Rounthwaite (US20100325133A1). Regarding claims 6 and 17, Mars in view of Li teaches: “further comprising A fit score assesses the training data samples and each data source may be evaluated based on a level of quality score. The scores describe the data source and training data that can improve the machine learning model training.) Mars in view of Li does not explicitly disclose an implementation of representing the data quality as a bipartite graph. However, Rounthwaite discloses in the same field of endeavor: “wherein the selecting comprises constructing a bipartite graph comprising the plurality of ” ([0021], A bipartite graph is generated to show a visual representation of the relationship between two types of entities. A weighted edge defines how relevant a first set of nodes map to a second set of nodes.) It would be obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of representing the data quality as a bipartite graph from Rounthwaite into the teaching of Mars in view of Li. Doing so can use a visual graphical representation to define high correlations between 2 different sets of data (Rounthwaite, par. 21). Claims 8 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Mars (US10296848B1) in view of Li, “Resolving Conflicts in Heterogeneous Data by Truth Discovery and Source Reliability Estimation” and McDonald (US20170212241A1). Regarding claims 8 and 19, Mars in view of Li teaches: “further comprising selecting Mars teaches a training data seed samples (validation dataset) is a set of dataset that a user would like to obtain more training data that are similar to the seed samples from a plurality of external data sources.) Mars in view of Li does not explicitly disclose an implementation of selecting the least number of data sources. However, McDonald discloses in the same field of endeavor: “wherein the selecting comprises selecting the least number of data sources that covers the The system evaluates the data from multiple data sources and selects a data source that meets a certain criterion, such as a data quality metric.) It would be obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of selecting the least number of data sources from McDonald into the teaching of Mars in view of Li. Doing so can select the data with the lowest uncertainty value to obtain the best results (Thompson, par. 33). Claims 9 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Mars (US10296848B1) in view of Li, “Resolving Conflicts in Heterogeneous Data by Truth Discovery and Source Reliability Estimation” and Choudhury (US20210150269A1). Regarding claims 9 and 20: Mars in view of Li does not explicitly disclose an implementation of “the respective data comprises a derivative of local data stored at the data source to preserve a privacy of the local data stored at the data source”. However, Choudhury discloses in the same field of endeavor: “wherein the respective data comprises a derivative of local data stored at the data source to preserve a privacy of the local data stored at the data source” ([0039-0040], The training data at each local data source have certain attributes filtered out prior to training the federated learning model. Attributes like direct identifiers are not used to train the model.) It would be obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of “the respective data comprises a derivative of local data stored at the data source to preserve a privacy of the local data stored at the data source” from Choudhury into the teaching of Mars in view of Li. Doing so can protect sensitive information from being leaked by a data source during training of a global federated learning model (Choudhury, abstract). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Xue, “Toward Understanding the Influence of Individual Clients in Federated Learning”, discloses a framework for quantifying each individual client influences in the collaborative training process. Lee (US20220158888A1) discloses a method of validating client models using a validation dataset to determine whether the client is an abnormal client in federated learning. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to GARY MAC whose telephone number is (703)756-1517. The examiner can normally be reached Monday - Friday 8:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached on (571) 270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GARY MAC/Examiner, Art Unit 2127 /ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127
Read full office action

Prosecution Timeline

Show 6 earlier events
Jul 21, 2025
Response after Non-Final Action
Jul 28, 2025
Response after Non-Final Action
Aug 19, 2025
Request for Continued Examination
Aug 26, 2025
Response after Non-Final Action
Nov 06, 2025
Non-Final Rejection mailed — §101, §103, §112
Feb 06, 2026
Response Filed
May 27, 2026
Final Rejection mailed — §101, §103, §112
Jul 27, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699910
EXPLAINABILITY FOR ARTIFICIAL INTELLIGENCE-BASED DECISIONS
3y 11m to grant Granted Aug 04, 2026
Patent 12688426
METHOD AND DEVICE FOR COMPRESSING NEURAL NETWORK
4y 8m to grant Granted Jul 21, 2026
Patent 12626130
METHOD AND DEVICE FOR COMPRESSING NEURAL NETWORK
4y 5m to grant Granted May 12, 2026
Patent 12608643
GENERATING WORKFLOW REPRESENTATIONS USING REINFORCED FEEDBACK ANALYSIS
4y 7m to grant Granted Apr 21, 2026
Patent 12596907
NEURAL NETWORK OPERATION APPARATUS AND METHOD
4y 8m to grant Granted Apr 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
43%
Grant Probability
86%
With Interview (+43.6%)
4y 4m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 21 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month