Prosecution Insights
Last updated: September 17, 2026
Application No. 18/765,288

GENERATION AND SECURING OF REFERENCE DATASETS FOR ARTIFICIAL INTELLIGENCE ALGORITHMS

Non-Final OA §103
Filed
Jul 07, 2024
Priority
Jul 08, 2023 — provisional 63/512,614
Examiner
YI, HYUNGJUN B
Art Unit
Tech Center
Assignee
Asher Orion Group LLC
OA Round
1 (Non-Final)
32%
Grant Probability
At Risk
1-2
OA Rounds
2y 2m
Est. Remaining
58%
With Interview

Examiner Intelligence

Grants only 32% of cases
32%
Career Allowance Rate
9 granted / 28 resolved
-27.9% vs TC avg
Strong +26% interview lift
Without
With
+25.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
24 currently pending
Career history
62
Total Applications
across all art units

Statute-Specific Performance

§101
28.3%
-11.7% vs TC avg
§103
54.6%
+14.6% vs TC avg
§102
11.9%
-28.1% vs TC avg
§112
4.4%
-35.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 28 resolved cases

Office Action

§103
DETAILED ACTION This action is responsive to the claims filed on 07/07/2024. Claims 1-20 are pending for examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 05/04/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over ITU et al., (ITU-T Focus Group on AI for Health, FG-AI4H DEL07, Artificial Intelligence for Health Evaluation Considerations (March 2023)), hereafter referred to as ITU in view of Petrick et al. (Nicholas A. Petrick, Weijie Chen, Jana G. Delfino, et al. "Regulatory considerations for medical imaging AI/ML devices in the United States: concepts and challenges," Journal of Medical Imaging 10(5), 051804 (23 Jun 2023) doi.org/10.1117/1.JMI.10.5.051804), hereafter referred to as Petrick. Claim 1: ITU teaches: A method, comprising: providing a computer system, the computer system comprising a processor system and a memory system in operative connection with the processor system, (ITU, page 11, paragraph 2, “The proposed system consists of three major components: an administrative backend, a public frontend, and an execution environment (cf. Figure 5). The administrative backend comprises a web application where administrators are able to develop and publish benchmarking tasks and can inspect the benchmarking results. The public frontend constitutes the public interface to the proposed benchmarking tasks.”, Page 11, section 9.1.1, “The administrative backend of the platform serves as the management interface for the administrative staff. It is here where new AI health benchmarking tasks will be developed and published. It also provides the data storage for the system: a database, which contains the benchmarking tasks, and a data set store for managing data sets.”; Page 14, section 9.1.3, “The execution environment consists of the execution manager service and an execution server pool on which execution clients can be run. The execution manager service orchestrates the benchmarking of the submitted software solutions and interfaces with the internal interface to retrieve queued submissions as well as the private test data sets. The execution server pool is a set of servers on which the actual execution of the software solutions is performed.”, ITU’s execution server pool and execution clients provide the claimed processor system that executes submitted AI software. ITU’s database and dataset store provide the claimed memory system operatively connected to those processor resources. The generic processor-and-memory language of claim 1 is therefore taught, or at minimum would have been an ordinary implementation of ITU’s expressly disclosed server-based architecture.) the memory system having one or more reference datasets which have confirmed truth stored thereon, (ITU, page 12, paragraph 1, “A new benchmarking task must at least consist of a name, a description, a private test data set, and a deadline. Furthermore, administrators may add further documentation, examples, a public data set and other auxiliary information. The private test data set will remain undisclosed and will be used to validate the submitted solutions. It represents the "gold standard" with the included true labels or annotations as "ground truth". In contrast, public data (if applicable) will be made available to all participants.”; Page 13, “Data set store – The data set store should be able to efficiently manage large amounts of unstructured data, e.g., by employing a binary large object (BLOB) store. The internal interface will store uploaded data sets here. Encryption of the data and version control are mandatory.”, The private test datasets stored in ITU’s dataset store correspond to the claimed reference datasets. The included true labels or annotations expressly provide “confirmed truth,” because ITU identifies them as the gold standard and ground truth against which submitted model outputs are evaluated.) wherein access to data of the one or more reference datasets is secured such that, at least one of: (i) the data of the one or more reference datasets cannot be accessed for use in training artificial intelligence algorithms or (ii) the data of the one or more reference datasets is altered to diminish the viability thereof in training artificial intelligence algorithms, (ITU, page 8, section 8.1, “Hence, the test data set must be unpublished and be withheld from the AI developer, and the validation should happen in a completely closed environment without connection to the Internet. A model that has been over-fitted to the test data, can achieve an excellent test result without actually being able to perform well in practice when fresh, unknown data points are coming in and need to be processed. For instance, the model could simply memorize the association between test data points and corresponding labels (if they are known) and then correctly reproduce the labels during the test / benchmark but be helpless in the real world when the model has to infer the label from fresh data points without knowing the label in advance.”; page 13, “While the public data sets should be downloadable by participants, the private test data sets must never be disclosed to the participants. This is particularly important because the private test data serve for testing the generalization capabilities of the AI-based solutions in a fair manner. Crucially, and in contrast to some "data science challenges", not only the true labels or annotations of the test data, but the entire test data sets including the "features", or "raw data" must remain undisclosed. This has the following reason: Access to the private test data might tempt some participants to take unfair advantage by (asking human experts to label the test data and then) tuning their AI solution to produce good results on these test data ("overfitting"). Yet, this overfitting does not ensure that the solutions can generalize to other, previously unseen data,”, ITU secures and completely withholds the private reference data specifically to prevent participants from labeling the data and tuning or training their models on it. This directly teaches alternative (i). Because claim 1 requires only “at least one of” alternatives (i) or (ii), it is unnecessary for the reference also to teach alteration under alternative (ii).) and providing a portal system to access the computer system over a network via a computing device (ITU, page 14, section 9.1.2, “The public frontend is comprised of a public-facing web client, which is a website that is the portal to all proposed benchmarking tasks. The public-facing web client interfaces with the internal interface to provide AI developers and other interested parties access to the published benchmarking tasks. Users are able to create a user account. Unauthenticated users and authenticated users, which are not signed up for a benchmarking task, will be presented a list of current benchmarking tasks and are able to search for benchmarking tasks by keywords. For each benchmarking task, a details page exists which contains the description of the benchmarking task and a deadline for submissions. Authenticated users will be able to sign up for a benchmarking task on the details page. Participants who have signed up for a benchmarking task will be able to access further documentation and examples, as well as the public (training or example) data set of the benchmarking task (if available). Furthermore, the website should provide participants with detailed information about the submission process.”, ITU’s public-facing web client is expressly described as a “website” and “the portal” to the benchmarking tasks. It communicates with the internal platform interface and permits remote users operating computing devices to interact with the benchmarking system over a network.) so that a user can access one of the one or more reference datasets to evaluate the performance of an artificial intelligence model in characterizing the one of the one or more reference datasets, (ITU, page 8, section 8.1, paragraph 3, “The benchmarking platform with validation in a closed environment works – in brief – as follows. The developer submits the to-be tested and already trained AI/ML model to the platform. In a closed environment, the model is provided with the test data points, processes these data, creates the corresponding output which is then compared by functionality of the platform with the "ground truth", using standardized quality criteria and metrics (as defined by the topic groups in their topic description documents [DEL10] and applicable sub-documents). The validation results are returned to the AI/ML developer and the benchmark organizer.”; Page 11, paragraph 2, “They can upload their AI solution for the benchmarking task to the platform in a standardized packaging format. The software projects contained in these solutions must adhere to a standardized interface for benchmarking. Once a solution has been handed in, it will be queued for benchmarking, which will be performed automatically by the execution environment. Together with the aforementioned documentation, the participants should be provided a minimal example simulating the execution environment, with the exact syntax that is later used in the validation but using exemplary input data only (and not the actual test data). In this way, the AI developers can check themselves whether their implementation works because it adheres to the standardized interface, which would greatly ease any kind of troubleshooting. The execution environment will report back the results to the administrative backend, where they will be aggregated into a central overview (customizable tables, visualisations and/or rankings with potentially multiple ranking schemes).”, ITU permits a user to access the private test dataset in a controlled manner by uploading a model that is executed against the dataset inside the platform. The model characterizes the test data, the characterization is compared with ground truth, and the resulting performance information is returned to the user.) Petrick, in the same field of medical artificial intelligence, teaches the following which ITU fails to teach: wherein the artificial intelligence model was independently trained using a dataset other than the one or more reference datasets. (Petrick, page 051804-6, paragraph 1, “A central principle for performance evaluation is that the test dataset should be independent of the training dataset (different patients and different clinical sites) to avoid biases in performance assessment and demonstrate performance generalizability. Violation of this principle will result in optimistically biased performance estimates and potentially unacceptable real-world performance.32 Although this principle is simple, there are subtle ways in which the independence principle can be violated. A typical example of violation arises from the failure to recognize that multiple images (or image regions) from the same patient are correlated. The images from the same patient are different, but not independent. Therefore, to avoid information leakage between the training and testing datasets, the images from each patient should appear in only one of these datasets…. Figure 2 shows the need for the AI/ML develop ment to be independent from the performance assessment conducted as part of a device submis sion including having the test data site and time independent from that of the AI/ML training and tuning data.”, Petrick expressly teaches that the dataset used to evaluate a finalized AI model must be separate and independent from the dataset used to train and tune that model. Thus, the evaluated model has been trained using data other than the claimed reference/test datasets.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of ITU with the teachings of Petrick, a motivation of which would have been to improve the reliability, representativeness, and clinical relevance of ITU’s standardized medical AI benchmarking process by applying known medical-imaging AI evaluation practices. ITU recognizes that meaningful AI benchmarking requires carefully designed, high-quality test datasets that account for differences in patient populations, measurement devices, and data sources, and further seeks standardized and comparable evaluation of AI models by a trusted third party. Petrick teaches complementary techniques for accomplishing those objectives in medical-imaging AI evaluation, including maintaining independence between training and test datasets, selecting sufficiently representative cohorts and subgroups, evaluating performance across different disease presentations and image-acquisition devices and protocols, comparing performance with predicate devices or other appropriate comparators, and assessing robustness and generalizability. Accordingly, one of ordinary skill would have been motivated to incorporate Petrick’s known dataset-design and medical-imaging evaluation practices into ITU’s benchmarking platform to obtain more accurate, unbiased, generalizable, and clinically meaningful performance assessments of medical AI models, with a reasonable expectation of success because Petrick’s techniques perform their established function of improving the quality and regulatory usefulness of AI/ML performance evaluation. Claim 2: ITU and Petrick teaches the limitations of claim 1, ITU further teaches: The method of claim I wherein the memory system further has stored thereon an evaluation algorithm which is executable by the processor system to evaluate performance of artificial intelligence models in characterizing the one of the one or more reference datasets, and wherein the method further comprises: (ITU, page 15, paragraph 2, “When the execution client is ready, the private test data are retrieved from the administrative backend. The true labels or annotations of the test data must be strictly withheld from the AI model that must have access only to the unlabelled / unannotated data points (i.e., "features"). Subsequently, the software solution is uploaded to the execution client, unpacked, and run. Now, the solution generates output variables y = f(x) from the test data x. A strictly separate functionality, that the to-be-tested AI model cannot interfere with (!), then calculates the benchmarking metrics by comparing the reported results with the true labels or annotations using the respective statistical metrics. These metrics need to be specified for each benchmarking task (accuracy, F1 score, precision, recall, ROC/AuC, Jaccard index, etc.).”, ITU’s “strictly separate functionality” is an evaluation algorithm executed by the benchmarking system. It evaluates the model by comparing the model-generated characterizations with true labels or annotations using predetermined statistical metrics.) executing the evaluation algorithm to evaluate the performance of the artificial intelligence model in characterizing the one of the one or more reference datasets; (ITU, page 15, paragraph 2, “When the execution client is ready, the private test data are retrieved from the administrative backend... A strictly separate functionality, that the to-be-tested AI model cannot interfere with (!), then calculates the benchmarking metrics by comparing the reported results with the true labels or annotations using the respective statistical metrics. These metrics need to be specified for each benchmarking task (accuracy, F1 score, precision, recall, ROC/AuC, Jaccard index, etc.).”, The execution client executes the submitted model on private reference data and executes the separate evaluation functionality to calculate performance metrics by comparison with ground truth.) and providing output of evaluation information characterizing the performance of the artificial intelligence model generated by the evaluation algorithm via the portal system to the computing device. (ITU, page 14, paragraph 2, “After validation of the submissions, it should be possible to display the results in customizable tables and visualisations to the participants, to benchmark organizers, and to selected expert evaluators. Results might be aggregated in optional and possibly anonymous rankings, with potentially multiple ranking schemes.”, ITU provides the calculated evaluation results to participants through the public-facing website and portal in tables, visualizations, and rankings. Those results constitute evaluation information characterizing model performance and are delivered to the user’s networked computing device.) Claim 3: ITU and Petrick teaches the limitations of claim 1, ITU further teaches: The method of claim 1 wherein the artificial intelligence model is stored in the memory system of the computer system and is executable by the processor system. (ITU, page 14, section 9.1.3, “Once a solution has been submitted by a participant, the software solution package is queued for benchmarking by the administrative backend. The execution manager service will go through the queued submissions and schedule them to run on an execution client in the execution server pool. When the execution client has completed the computations of the submitted software solution on the private test data as well as the calculation of the benchmarking metrics, it will report the results back to the execution manager service, which will in turn communicate them back to the administrative backend. Participants submit software solutions in the form of a standardized package. The packaging format could, for example, be a compressed file (ZIP, TAR GZIP etc.), which contains all files necessary to execute the solution. The package should contain a top-level executable (an EXE file or a batch file for a Windows environment, or a Linux executable or a Bash script file for a linux environment), which is run by the execution client.”; Page 15, paragraph 2, “Subsequently, the software solution is uploaded to the execution client, unpacked, and run. Now, the solution generates output variables y = f(x) from the test data x.”, The submitted model package is retained and queued by the platform, uploaded into the execution client, unpacked, and executed by the server. The platform storage and execution client respectively correspond to the claimed memory and processor systems.) Claim 4: ITU and Petrick teaches the limitations of claim 3, ITU further teaches: The method of claim 3 wherein the artificial intelligence model is transmitted to the computer system via the portal system (ITU, page 14, paragraph 2, “A submission will consist of two parts: the software solution and the documentation. The software solution must be packaged in a standardized format, which contains everything needed to execute the solution. The documentation should be contained in a single document (e.g., PDF or text file). The details page of a benchmarking task provides the means to upload solutions and documentations. The upload must be performed via an encrypted communication channel (e.g., HTTPS), because the solutions may contain secret intellectual property/trade secrets of the participant.”; page 15, section 9.1.3, “When the execution client is ready, the private test data are retrieved from the administrative backend. The true labels or annotations of the test data must be strictly withheld from the AI model that must have access only to the unlabelled / unannotated data points (i.e., ‘features’). Subsequently, the software solution is uploaded to the execution client, unpacked, and run. Now, the solution generates output variables y = f(x) from the test data x. A strictly separate functionality, that the to-be-tested AI model cannot interfere with (!), then calculates the benchmarking metrics by comparing the reported results with the true labels or annotations using the respective statistical metrics.”, ITU teaches that the AI developer or participant transmits its submitted AI solution/software solution through the benchmarking portal using an encrypted HTTPS communication channel. ITU makes clear that the submitted “software solution” is the executable implementation of the “AI model” being tested because, in describing the same execution workflow, ITU states that the private test data are provided to the “AI model,” after which the “software solution is uploaded to the execution client, unpacked, and run” to generate outputs from those test data. Accordingly, ITU’s submitted AI/software solution corresponds to the claimed artificial intelligence model being transmitted to the computer system.) in an encrypted, containerized format from a manufacturer thereof and stored on the memory system. (ITU, page 10, “The to-be tested AI/ML model must also be protected from undesired access aimed for example at getting hold of the source code itself (intellectual property) or of the training data via the model (e.g., through model inversion or adversarial attacks). Access to the models would also make it possible for data-providers to unnoticeably manipulate the test data (to make it harder for competitors to achieve good results, e.g., by adding adversarial noise). Finally, from the benchmarking perspective, it must also be assured that neither model-provider nor data-provider can manipulate or interfere with the validation procedure. Both these last points could be addressed by preventing access to the original source, for example via encapsulation into containerization software (e.g., docker or apache mesos) or pre-compilation.”; Page 15, paragraph 1, “This can be best implemented by using some sort of containerization software, Docker, Mesos, Singularity/Apptainer, etc., or through virtualization, VMWare, Virtual Box, HyperV, etc. Each solution package must contain—besides the actual software—information about the desired execution environment… It creates a new execution client by starting a new container/virtual machine using the selected base image… Subsequently, the software solution is uploaded to the execution client, unpacked, and run.”, ITU teaches transmitting and storing a standardized AI software package for execution within a Docker-, Mesos-, or equivalent containerized environment. In combination with ITU’s encrypted portal upload, this teaches or at least suggests transmission and storage of the model in the claimed encrypted, containerized manner.) Claim 5: ITU and Petrick teaches the limitations of claim 4, ITU further teaches: The method of claim 4 wherein a federated learning platform is used to transmit the artificial intelligence model to the computer system. (ITU, page 10, paragraph 1-2, “Benchmarking in a closed environment can potentially also be conducted in a federated fashion (unlike the centralized concept described in clause 8.1 and without transferring samples such as in clause 8.2). In the federated approach to benchmarking, the test data sets remain where they had been acquired, for instance, in different hospitals. The to-be-tested (and already trained) AI/ML model is sent to the locations where the data are stored (to the hospitals in this example). Here, the model is benchmarked against the local test data with standardized test procedures and metrics. The results are returned to the party that is organizing the benchmark.”, ITU expressly teaches a federated platform architecture in which the AI model, rather than the test data, is transmitted to distributed computers at the locations where the reference datasets are stored.) Claim 6: ITU and Petrick teaches the limitations of claim 1, ITU further teaches: The method of claim 1 wherein the user is provided access to the one or more reference datasets to execute a plurality of artificial intelligence models to characterize the one of the one or more reference datasets (ITU, page 4, paragraph 1, “As growing numbers of AI models become available for use researchers, patients, clinicians, and policy makers require a framework to understand whether the models are safe, purpose-fit and cost effective, and to compare model performance with current standards of care, and between each other.”; Page 9, “Moreover, all (potentially competing) AI models must be benchmarked at the very same moment, in the case of validation via interface. Otherwise, test data received by "model A" could be stored and included in a separate "model B" by the same developer, which could take part in the benchmark at a later time point.”; Page 10, “The security mechanisms need to assure that all models have processed the very same data. Only then, the results are comparable, and the procedure is trustworthy.”, ITU permits participating users to submit and execute multiple competing AI models against the same secured reference data so that their performance is comparable. Each submitted model processes and characterizes the same data through the controlled benchmarking environment.) to evaluate the performance of each of the plurality of artificial intelligence models in characterizing the one of the one or more reference datasets, (ITU, page 4, paragraph 1, “As growing numbers of AI models become available for use researchers, patients, clinicians, and policy makers require a framework to understand whether the models are safe, purpose-fit and cost effective, and to compare model performance with current standards of care, and between each other.”; page 7, section 8, “Independent benchmarking by a trusted third party using agreed-upon, standardized test procedures and metrics on separate high-quality test data from different sources is the core idea for the technical validation step, pursued by the ITU/WHO Focus Group on ‘AI for health’ (cf. Figure 3). This approach could be a valuable complement to in-house or local technical tests and subsequent clinical trials. It does not put test subjects at risk, can be repeated in the case of model / software updates, can be based on large amounts of high-quality test data from different sources and sites, and is fast. Moreover, the approach can lead to comparable and transparent results using standardized testing procedures with meaningful test objectives, test tasks, quality criteria, and test metrics defined by a community of experts.”; page 10, section 8.3, “Moreover, the security mechanisms need to assure that all models have processed the very same data. Only then, the results are comparable, and the procedure is trustworthy.”, ITU expressly teaches evaluating the performance of a plurality of AI models, stating that its evaluation framework is used “to compare model performance … between each other.” ITU further teaches that all of the models process the same test data and that standardized test procedures, quality criteria, and metrics are applied to produce comparable results. Thus, each model of the plurality is individually benchmarked against the same reference/test dataset, thereby evaluating the respective performance of each artificial intelligence model in characterizing that dataset, and the resulting performance measurements are then comparable across the plurality of models.) Petrick further teaches: wherein each of the artificial intelligence models was independently trained using a dataset other than the accessed one of the one or more reference datasets. (Petrick, page 051804-6, paragraph 1, “A central principle for performance evaluation is that the test dataset should be independent of the training dataset (different patients and different clinical sites) to avoid biases in performance assessment and demonstrate performance generalizability. Violation of this principle will result in optimistically biased performance estimates and potentially unacceptable real-world performance.32 Although this principle is simple, there are subtle ways in which the independence principle can be violated. A typical example of violation arises from the failure to recognize that multiple images (or image regions) from the same patient are correlated. The images from the same patient are different, but not independent. Therefore, to avoid information leakage between the training and testing datasets, the images from each patient should appear in only one of these datasets…. Figure 2 shows the need for the AI/ML develop ment to be independent from the performance assessment conducted as part of a device submis sion including having the test data site and time independent from that of the AI/ML training and tuning data.”, Petrick expressly teaches that the dataset used to evaluate a finalized AI model must be separate and independent from the dataset used to train and tune that model. Thus, the evaluated model has been trained using data other than the claimed reference/test datasets.) Claim 7: ITU and Petrick teaches the limitations of claim 1, Petrick further teaches: The method of claim 6 wherein one of the plurality of artificial intelligence models is a predicate artificial intelligence model. (Petrick, page 051804-2, “Radiology has been a pioneer in adopting AI/ML-enabled devices into the clinical environment. Radiological AI/ML devices are numerous and expanding with applications aiming to improve the efficiency, accuracy, or consistency of the medical image interpretation process across a wide range of radiological tasks and imaging modalities. A non-exhaustive list of AI/ML-enabled medical devices authorized for marketing in the United States and identified through FDA’s publicly available information can be found on the FDA webpage. Table 1 defines some common types or classes of medical imaging AI/ML that have been submitted to the FDA. Each product type, and its associated product code, contains one or more devices authorized for marketing in the United States with most of these devices considered SaMD.”; page 051804-3, section 2.1.1, “Most medical image processing devices, with or without AI/ML, are currently classified as class II. Aside from some exemptions and in addition to general controls, to market a class II device, a manufacturer must describe and test their device according to all applicable special controls and demonstrate substantial equivalence between their new device and a legally marketed device, i.e., the predicate device.”; page 051804-7, sections 2.4–2.4.1, “Standalone testing also provides a performance benchmark for comparing AI/ML devices from the same or different manufacturers. This benchmark can reduce the need for clinical performance testing in future regulatory submissions… Selection of performance metrics is crucial for benchmarking an AI/ML model and for comparing performance with predicate devices or other appropriate comparators.”, Petrick teaches that legally marketed medical imaging devices include AI/ML-enabled devices and that standalone testing is used to compare AI/ML devices from the same or different manufacturers. Petrick further expressly teaches benchmarking an AI/ML model by comparing its performance with a predicate device. Thus, where the legally marketed predicate device is one of the AI/ML-enabled medical devices described by Petrick, the predicate device includes an operative AI/ML model, and that operative model constitutes the claimed “predicate artificial intelligence model.” Accordingly, Petrick teaches or at least suggests including, among the plurality of AI models being comparatively benchmarked, the AI/ML model of a legally marketed predicate device as the predicate artificial intelligence model.) Claim 8: ITU and Petrick teaches the limitations of claim 1, ITU further teaches: The method of claim 1wherein access to the portal system is provided via a web browser. (ITU, page 14, section 9.1.2, “The public frontend is comprised of a public-facing web client, which is a website that is the portal to all proposed benchmarking tasks. The public-facing web client interfaces with the internal interface to provide AI developers and other interested parties access to the published benchmarking tasks. Users are able to create a user account. Unauthenticated users and authenticated users, which are not signed up for a benchmarking task, will be presented a list of current benchmarking tasks and are able to search for benchmarking tasks by keywords.”, ITU expressly characterizes the portal as a public-facing website and web client. A user’s access to that website through a conventional web browser teaches or at minimum would have been the ordinary and expected implementation of the claimed browser access.) Claim 9: ITU and Petrick teaches the limitations of claim 1, ITU further teaches: The method of claim 1 wherein each of the one or more reference datasets is created based upon predetermined rules defining construction (ITU, page 3, paragraph 1, “The process commences with a clear definition of the respective task that the AI/ML models under assessment are expected to perform and of the intended use. Suitable quality criteria with corresponding meaningful procedures and metrics are specified for the initial technical test / benchmarking steps. The required test data properties are defined. The test data set is collected according to these requirements and assessed for quality, realism, representativeness, and fidelity to the target population.”; Page 5, paragraph 1, “Careful attention must be paid to define a population of interest and systematically collect samples (test cases) which cover this population. It is very much a question of design of experiments and careful choices of test cases. A proper sampling paradigm / scheme (that says we need exactly more of, e.g., "male; 10-15 year-old", "female; 70-80 year-old; smoker") would help do a data-informed and targeted data search.”, ITU requires defining the intended task, target population, data properties, sampling scheme, quality criteria, and representativeness requirements before collecting the test data. Those predefined requirements constitute predetermined rules governing construction and population of each reference dataset.) and statistical characterization of the data to be populated therein. (ITU, page 8, paragraph 1, “The topic description documents also specify the metrics (accuracy, F1 score, precision, recall, ROC/AUC, Jaccard index, etc.) for every benchmarking task. With these metrics, the predictive performance of a classification or regression model can be assessed.”, ITU expressly requires rules addressing the distribution and balance of the dataset, number of cases, statistical power, generalizability, and applicable statistical metrics. These requirements teach predetermined statistical characterization of the data selected for inclusion in the reference datasets.) Claim 10: ITU and Petrick teaches the limitations of claim 1, ITU further teaches: The method of claim 9 further comprising making information available to at least the user describing the process of creation of the one or more reference datasets. (ITU, page 14, paragraph 1, “Participants who have signed up for a benchmarking task will be able to access further documentation and examples, as well as the public (training or example) data set of the benchmarking task (if available). Furthermore, the website should provide participants with detailed information about the submission process. A submission will consist of two parts: the software solution and the documentation. The software solution must be packaged in a standardized format, which contains everything needed to execute the solution. The documentation should be contained in a single document (e.g., PDF or text file). The details page of a benchmarking task provides the means to upload solutions and documentations. The upload must be performed via an encrypted communication channel (e.g., HTTPS), because the solutions may contain secret intellectual property/trade secrets of the participant. After validation of the submissions, it should be possible to display the results in customizable tables and visualisations to the participants, to benchmark organizers, and to selected expert evaluators.”, ITU makes task documentation, examples, instructions, and results available to users through the public-facing portal.) Claim 11: ITU and Petrick teaches the limitations of claim 1, Petrick further teaches: The method of claim 1 wherein each of the one or more reference datasets is an image dataset (Petrick, page 051804-2, table 1, “Radiological computer assisted detection/diagnosis software for lesions suspicious for cancer… Software algorithm device to assist users in digital pathology—An in vitro diagnostic device intended to evaluate acquired scanned pathology whole slide images.”; Page 051804-3, paragraph 3, “Radiological AI/ML devices are numerous and expanding with applications aiming to improve the efficiency, accuracy, or consistency of the medical image interpretation process across a wide range of radiological tasks and imaging modalities.”, Petrick teaches reference and testing datasets comprising mammographic, MR, CT, ultrasound, radiographic, and pathology images associated with detecting or characterizing lesions suspicious for cancer.) wherein an oncologic condition associated with each image dataset thereof is confirmed by at least one of pathologic confirmation or radiographic confirmation. (Petrick, page 051804-8, section 2.5, “Here, we define a medical imaging reader study as a study in which readers, e.g., radiologists or pathologists, review and interpret medical images for a specified clinical task, e.g., diagnosis, and provide an objective interpretation, such as a rating of the likelihood that a condition is present.”; Page 051804-9, paragraph 1, “The design of an MRMC reader study involves a number of considerations including patient data collection, establishment of a reference standard, recruitment and training of readers, and the study design along with other factors.”; Page 051804-9, section 3, “Data governance concerns relate to devel oping effective policies and protocols for storing, securing, and maintaining data quality, including images, metadata, and reference standard (truth) labels.”, Petrick teaches image datasets having objective diagnostic reference-standard or truth labels established for medical image cases reviewed by radiologists or pathologists. Radiologist confirmation of the condition corresponds to radiographic confirmation, while pathologist evaluation of pathology images corresponds to pathologic confirmation.) Claim 12: ITU and Petrick teaches the limitations of claim 1, Petrick further teaches: The method of claim 11 wherein the one or more reference datasets are divided into tranche types comprising a cancer subtype reference dataset tranche, (Petrick, page 051804-7, section 2.3.2, “For a testing dataset, it is more important to match the study data with the target population, but there may be some flexibility under FDA’s least burdensome principle in studies with controlled design and informed interpretation of results. It may be necessary to include rare cases or patient subgroups, or it may be statistically efficient to stratify, or enrich, the sampling across patient subgroups. If the enrichment is based solely on the disease condition, the impact may only be on the prevalence of the study set.”; pages 051804-7–051804-8, section 2.4.2, “Subgroup assessment is an important component of standalone testing. In this case, model performance is assessed on individual or combined subgroups to better understand where the model may have performance limitations. Studies may need to be sized to yield a certain statistical precision of the performance on different subgroups, or studies may simply report the performance on different subgroups with accompanying confidence intervals… Depending on the clinical task and context, subgroup analyses found in AI/ML submissions are based on patient demographics, e.g., patient age, sex, race; image acquisition conditions, e.g., acquisition device and protocol; and disease type/presentation, e.g., disease subtype and lesion size/shape; or a combination of these characteristics.”, Petrick teaches that a reference/testing dataset may be stratified into separately evaluated patient subgroups and expressly identifies “disease subtype” as one characteristic by which those subgroups are defined. The Examiner interprets the claimed “cancer subtype reference dataset tranche” as a distinct subset or subgroup of the oncology reference dataset containing cases associated with a particular cancer/disease subtype. Thus, in the context of the oncologic reference datasets of claim 11, Petrick’s separately stratified testing subgroup defined according to “disease subtype” corresponds to the claimed cancer subtype reference dataset tranche.) an imaging acquisition reference dataset tranche, (Petrick, page 051804-7, section 2.4.2, “Depending on the clinical task and context, subgroup analyses found in AI/ML submissions are based on patient demographics, e.g., patient age, sex, race; image acquisition conditions, e.g., acquisition device and protocol; and disease type/presentation, e.g., disease subtype and lesion size/shape; or a combination of these characteristics.”; Page 051804-5, section 2.3, “When collecting image data for devel oping or testing AI/ML models, it is important to also acquire appropriate clinical information, e.g., patient demographics, family history, reason for the exam; disease specific information, e.g., disease type and lesion size; image acquisition information, e.g., patient prep, device manufacturer and model, protocol, and reconstruction method; and other clinical test results in order to characterize and understand training and testing limitations and model generalizability.”, Petrick teaches organizing or stratifying test cases according to image-acquisition device, protocol, manufacturer, model, preparation, and reconstruction method. Such an image-acquisition subgroup reads on the claimed imaging-acquisition tranche.) and an imaging robustness reference dataset tranche. (Petrick, page 051804-8, section 2.4.3, “A repeatability or reproducibility study refers to standalone assessments that investigate differences in AI/ML output from reimaging a patient with the same or different acquisition devices and conditions. Such studies are commonly included in in vitro diagnostic device sub missions and may have value for submissions of AI/ML devices.48–50 However, these designs have been less common in medical imaging applications, for example, when the study would have required exposing the patient to additional ionizing radiation. When appropriate, like scan ning pathology slides multiple times with whole slide imaging systems or taking repeated pic tures of a skin lesion with the camera of a mobile device, repeatability and reproducibility studies can demonstrate AI/ML device robustness and generalizability. More robust AI/ML devices will have higher repeatability/reproducibility across real-world use cases.”, Petrick teaches a distinct collection of reimaged cases generated under the same or different devices and acquisition conditions for testing repeatability, reproducibility, and robustness. That collection corresponds to the claimed imaging-robustness reference-dataset tranche.) Claim 13: ITU and Petrick teaches the limitations of claim 1, Petrick further teaches: The method of claim 9 wherein the predetermined rules are used in creating the reference datasets to enable evaluation the artificial intelligence model in determining at least one of (i) bias across subgroups, (Petrick, page 051804-5, section 2.3, paragraph 2, “The datasets, both training and test, should include relevant cohorts and subgroups containing enough patients/cases to facilitate robust algorithm training and facilitate subgroup analyses as discussed in Sec. 2.4. Research has shown that as the training set is gradually increased from a small size, overfitting initially decreases dramatically, with diminishing returns as the dataset size gets larger.31 The rate at which adding more data improves performance depends on the complexity of the AI/ML model and the complexity of the data space. Estimation of the test dataset size for adequate precision and study power is a classical problem in statistics, and pilot data are extremely helpful for estimating the dataset sizes appropriately.”; Page 051804-7, section 2.4.2, “Subgroup assessment is an important component of standalone testing. In this case, model performance is assessed on individual or combined subgroups to better understand where the model may have performance limitations. Studies may need to be sized to yield a certain statistical precision of the performance on different subgroups, or studies may simply report the performance on different subgroups with accompanying confidence intervals.”, Petrick teaches predetermined dataset-construction requirements concerning subgroup inclusion, sufficient sample size, statistical precision, and study power so that model performance and limitations can be evaluated separately across those subgroups. Differences in performance across the defined subgroups reveal subgroup bias.) (ii) fairness, (Petrick, page 051804-9, section 3, “Data governance concerns relate to devel oping effective policies and protocols for storing, securing, and maintaining data quality, including images, metadata, and reference standard (truth) labels.67 Algorithm robustness concerns generally include how to reduce algorithmic bias and improve fairness across patients, groups, and sites. FDA is particularly concerned with how sponsor studies, often based on limited patient, group and site diversity, generalize to actual clinical practice across the United States. All AI/ML algorithms are biased to some extent because the data used to train the models are intrinsically a function of the population groups, disease conditions, clinical environments, and imaging technologies available during the collection process.67 Therefore, robustness includes not only devel oping methods to measure and reduce important sources of bias but also defining fairness criteria for the AI/ML under appropriate operational conditions as well as integrating the appropriate level of transparency.”, Petrick expressly teaches constructing and evaluating medical-AI datasets and studies to measure bias and improve or define fairness across patients, demographic groups, clinical sites, disease conditions, and imaging technologies.) and (ii) response to data variation. (Petrick, page 051804-8, section 2.4.3, “A repeatability or reproducibility study refers to standalone assessments that investigate differences in AI/ML output from reimaging a patient with the same or different acquisition devices and conditions. Such studies are commonly included in in vitro diagnostic device sub missions and may have value for submissions of AI/ML devices.48–50 However, these designs have been less common in medical imaging applications, for example, when the study would have required exposing the patient to additional ionizing radiation. When appropriate, like scanning pathology slides multiple times with whole slide imaging systems or taking repeated pictures of a skin lesion with the camera of a mobile device, repeatability and reproducibility studies can demonstrate AI/ML device robustness and generalizability. More robust AI/ML devices will have higher repeatability/reproducibility across real-world use cases.”, Petrick teaches designing the test data to vary patient subpopulations, clinical sites, acquisition devices, protocols, and imaging conditions, and then measuring differences in AI/ML output caused by those variations. This directly enables evaluation of the model’s response to data variation.) Claims 14-18 recite limitations substantially similar to claims 1-5, as such a similar analysis applies. Claim 19 recites limitations substantially similar to claim 9, as such a similar analysis applies. Claim 20 recites limitations substantially similar to claim 13, as such a similar analysis applies. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Karargyris, A., Umeton, R., Sheller, M. J., Aristizabal, A., George, J., Wuest, A., ... & Mattson, P. (2023). Federated benchmarking of medical artificial intelligence with MedPerf. Nature machine intelligence, 5(7), 799-810. Homeyer, A., Geißler, C., Schwen, L. O., Zakrzewski, F., Evans, T., Strohmenger, K., ... & Zerbe, N. (2022). Recommendations on compiling test datasets for evaluating artificial intelligence solutions in pathology. Modern Pathology, 35(12), 1759-1769. Pollard, V. T., Ryan, M. W., & Mohanty, A. (2022). FDA issues good machine learning practice guiding principles. The Journal of Robotics, Artificial Intelligence & Law, 5. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HYUNGJUN B YI whose telephone number is (703)756-4799. The examiner can normally be reached M-F 9-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached on (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /H.B.Y./Examiner, Art Unit 2124 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Jul 07, 2024
Application Filed
Aug 20, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12651178
MACHINE LEARNING TECHNIQUES FOR ASSOCIATING NETWORK ADDRESSES WITH INFORMATION OBJECT ACCESS LOCATIONS
5y 4m to grant Granted Jun 09, 2026
Patent 12619888
END-TO-END SYSTEMS AND METHODS FOR CONSTRUCT SCORING
1y 7m to grant Granted May 05, 2026
Patent 12536429
INTELLIGENTLY MODIFYING DIGITAL CALENDARS UTILIZING A GRAPH NEURAL NETWORK AND REINFORCEMENT LEARNING
4y 7m to grant Granted Jan 27, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
32%
Grant Probability
58%
With Interview (+25.8%)
4y 4m (~2y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 28 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month