Prosecution Insights
Last updated: August 17, 2026
Application No. 18/392,732

DATA PROCESSING DEVICE AND DATA PROCESSING METHOD

Non-Final OA §101§103
Filed
Dec 21, 2023
Priority
Jul 07, 2021 — continuation of PCTJP2021025545
Examiner
ZENG, WENWEI
Art Unit
Tech Center
Assignee
Mitsubishi Electric Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
20 currently pending
Career history
18
Total Applications
across all art units

Statute-Specific Performance

§101
44.1%
+4.1% vs TC avg
§103
49.2%
+9.2% vs TC avg
§102
3.4%
-36.6% vs TC avg
§112
3.4%
-36.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on December 21, 2023; August 27, 2024; May 8, 2025; and April 20, 2026 were considered by the examiner. The submissions are in compliance with the provisions of 37 CFR 1.97. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1- 6 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (mental process) without significantly more. Claim 1: Regarding claim 1, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites 1. A data processing device comprising processing circuitry to generate a plurality of pieces of candidate input data by putting together a plurality of pieces of trained input data used for first training in a machine learning model and a plurality of pieces of untrained input data not used for the first training … , and a device or system is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components: to generate a plurality of pieces of candidate input data by putting together a plurality of pieces of trained input data used for first training in a machine learning model and a plurality of pieces of untrained input data not used for the first training … (mental process, a person can mentally evaluate and generate input data, see MPEP 2106.04(a)(2)(III)) generate a plurality of pieces of candidate intermediate data by putting together trained intermediate data given …. and untrained intermediate data given... (mental process, a person can mentally evaluate and generate candidate intermediate data by placing together trained and untrained intermediate data, see MPEP 2106.04(a)(2)(III)), select one piece of candidate intermediate data from among the plurality of pieces of the candidate intermediate data, ... preferentially selecting the one piece of candidate intermediate data having a greater degree of heterogeneity when used for second training subsequent to the first training as compared with selected intermediate data including selected trained intermediate data that is the trained intermediate data already selected and selected untrained intermediate data that is the untrained intermediate data already selected …, (mental process, a person can mentally evaluate and select one piece of candidate intermediate data, see MPEP 2106.04(a)(2)(III)), and to select one piece of candidate input data, from among the plurality of pieces of the candidate input data, corresponding to the one piece of candidate intermediate data as data to be used at a time of the second training, (mental process, a person can mentally evaluate and select with pen and paper which piece of candidate input data that relates to candidate intermediate data that is used for a second round of training a model, see MPEP 2106.04(a)(2)(III)), If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: A data processing device comprising processing circuitry … (In step 2A, prong 2, this recites a generic computer component being used as a tool. – see MPEP 2106.05(f)), the processing circuitry … (In step 2A, prong 2, this recites a generic computer component being used as a tool. – see MPEP 2106.05(f)), by inputting the plurality of pieces of the trained input data into the machine learning model … (In step 2A, prong 2, inputting recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), by inputting the plurality of pieces of untrained input data into the machine learning model … (In step 2A, prong 2, inputting recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional elements v and vi recite generic computer component being used as a tool to perform abstract ideas, which are not indicative of significantly more. The additional elements vii and viii recite mere data gathering and are considered insignificant extra-solution activities. In step 2B, these insignificant extra-solution activities are well understood routine and conventional activities, which include receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Claim 2: Regarding claim 2, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “2. A data processing device comprising processing circuitry to select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data given …” , and a device or system is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components: select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data given … preferentially selecting the one piece of candidate intermediate data having a greater degree of heterogeneity when used for second training subsequent to the first training as compared with selected intermediate data which is selected untrained intermediate data that is the untrained intermediate data already selected, (mental process, a person can mentally evaluate and select one piece of candidate intermediate data from many candidate intermediate data, see MPEP 2106.04(a)(2)(III)), and to select one piece of candidate input data corresponding to the one piece of candidate intermediate data selected from among a plurality of pieces of candidate input data which are the plurality of pieces of untrained input data to be used at a time of the second training. (This recites a mental process, a person can mentally evaluate and select with pen and paper which piece of candidate input data that relates to untrained input data that is used for a second round of training a model , see MPEP 2106.04(a)(2)(III)), If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: A data processing device comprising processing circuitry… (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)), … by inputting, to a machine learning model, a plurality of pieces of untrained input data not used for first training in the machine learning model, (In step 2A, prong 2, inputting recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), … the processing circuitry, (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional elements iii and v generic computer component being used as a tool to perform abstract ideas, which are not indicative of significantly more. The additional element iv recites mere data gathering and is considered insignificant extra-solution activity. In step 2B, this insignificant extra-solution activity is well understood routine and conventional activity, which includes receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)), Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Claim 3: Regarding claim 3, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “3. A data processing device comprising processing circuitry to select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data …” , and a device or system is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components: select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data given … more preferentially selecting the one piece of candidate intermediate data having a greater degree of heterogeneity when used for second training subsequent to the first training as compared with trained intermediate data and selected intermediate data which is selected untrained intermediate data that is the untrained intermediate data already selected;… (mental process, a person can mentally evaluate and select one piece of candidate intermediate data from other pieces of candidate intermediate data, see MPEP 2106.04(a)(2)(III)), … and select one piece of candidate input data corresponding to the one piece of candidate intermediate data selected from among a plurality of pieces of candidate input data which are the plurality of pieces of untrained input data to be used at a time of the second training, (This recites a mental process, a person can mentally evaluate and select with pen and paper which piece of candidate input data that relates to untrained input data that is used for a second round of training a model , see MPEP 2106.04(a)(2)(III)), If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: A data processing device comprising processing circuitry … (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)), … by inputting, to a machine learning model, a plurality of pieces of untrained input data not used for first training in the machine learning model, (In step 2A, prong 2, inputting recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), the processing circuitry… (this recites a generic computer component being used as a tool. – see MPEP 2106.05(f)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional elements iii and v recite generic computer component being used as a tool to perform abstract ideas, which are not indicative of significantly more. The additional element iv recites mere data gathering and is considered insignificant extra-solution activity. In step 2B, this insignificant extra-solution activity is well understood routine and conventional activity, which includes receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)), Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Claims 4 - 6: Regarding claims 4-6, “A data processing method performed by a processing circuitry comprising: …”, and a method is one of the four statutory categories of invention in step 1 of the 101-analysis set forth in MPEP 2106. Since claims 4-6 recite similar limitations as corresponding independent claims 1-3 listed above, they are rejected for similar reasons under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 2, 4, and 5 are rejected under 35 U.S.C. 103 over Nguyen, C. et al. US10867246B1 , published on December 15, 2020, (hereafter, Nguyen) in view of Schneider, S. in “Deep Learning Based Computer Vision for Animal Re-Identification”, published in June 2020 , available at https://atrium.lib.uoguelph.ca/bitstream/10214/18056/3/Schneider_Stefan_202006_PhD.pdf , (hereafter, Schneider), further in view of Zhao, Z. et al. in “ Dsal: Deeply supervised active learning from strong and weak labelers for biomedical image segmentation,” from IEEE journal of biomedical and health informatics, 25(10), pages 3744-3751, published on January 18, 2021, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9326423, (hereafter, Zhao). Claim 1: Regarding claim 1, Nguyen teaches “1. A data processing device comprising processing circuitry to generate a plurality of pieces of candidate input data by putting together a plurality of pieces of trained input data used for first training in a machine learning model…” See Nguyen in col. 1, lines 35-36, describe " The neural network is trained using a first training dataset." Here, Nguyen shows for first training in a machine learning model, (i.e. putting together a plurality of pieces of trained input data). Further, Nguyen teaches “to select one piece of candidate intermediate data from among the plurality of pieces of the candidate intermediate data,” See Nguyen in col. 7, lines 14-22, describe "The deep learning module 120 selects a subset of samples from the input dataset. The subset of samples is selected based on distances between vectors representing the samples. Further details of the process for selecting the subsets are shown in FIG. 5 and described herein. The deep learning module 120 stores the subset of samples in the training data store 360 as a training dataset. The deep learning module 120 provides 460 the samples of the subset as a training dataset for retraining the neural network during the next iteration of the process". When Nguyen mentions the " subset of samples is selected based on distances between vectors representing the samples" , this shows selecting candidate intermediate data from plurality of intermediate data, where intermediate data is construed to mean vector representations or other similar representations of the data. Intermediate data is construed to include any form of latent representations in machine learning that represent the connection between raw, unstructured inputs (like text, images, or audio) and the final result. Further, Nguyen teaches “the processing circuitry more preferentially selecting the one piece of candidate intermediate data having a greater degree of heterogeneity when used for second training subsequent to the first training as compared with selected intermediate data including selected trained intermediate data that is the trained intermediate data already selected ;” See Nguyen in page 11, column 5, lines 63-67 through column 6, lines 1-9 describe “The sample selection module 340 selects samples that are determined to have low similarity to the clusters to which the samples are assigned. Accordingly, the sample selection module 340 excludes samples that are determined to have high similarity to the clusters to which they are assigned.” Nguyen describes low similarity relates to items that are less similar to one another and thus more diverse or (i.e. having a greater degree of heterogeneity). See also Nguyen column 1, lines 40-43, describe “The samples of the subset are provided as a second training dataset for retraining the neural network. This process may be repeated.” Here, Nguyen shows for data to be used for second training, this is part of a second training dataset. Further, see Nguyen from col. 7, lines 14-16, and 17-19, where Nguyen describes “The subset of samples is selected based on distances between vectors representing the samples...The deep learning module 120 stores the subset of samples in the training data store 360 as a training dataset.” Nguyen here describes selecting the trained intermediate data which was data that was already selected. Also, see Nguyen in page 9, col. 1, lines 35-41, describe "The neural network is trained using a first training dataset. The neural network is executed to process samples from an input dataset. A vector representation of each sample processed by the neural network is obtained from a particular hidden layer of nodes. A subset of samples from the input dataset is obtained based on distances between vectors representing the samples." Here, Nguyen shows that for trained intermediate data, the intermediate data relates to a vector representation of each sample of the trained neural network on a first training dataset. The selection is processing only the samples from an input. Further, Nguyen teaches “and to select one piece of candidate input data, from among the plurality of pieces of the candidate input data, corresponding to the one piece of candidate intermediate data as data to be used at a time of the second training.” See Nguyen in col. 7, lines 14-22, describe "The deep learning module 120 selects a subset of samples from the input dataset. The subset of samples is selected based on distances between vectors representing the samples. Further details of the process for selecting the subsets are shown in FIG. 5 and described herein. The deep learning module 120 stores the subset of samples in the training data store 360 as a training dataset. The deep learning module 120 provides 460 the samples of the subset as a training dataset for retraining the neural network during the next iteration of the process". Here, Nguyen shows that the selection of a subset of samples from the input dataset (i.e. select one piece of candidate input data), and each of the subset of samples is chosen based on distances between vectors representing the samples (where vectors representing the samples show one piece of candidate intermediate data). Intermediate data is construed by examiner to mean any form of intermediate representations such as vector representations or feature embeddings. Using the "samples of the subset as a training dataset for retraining the neural network" model "during the next iteration" is construed to mean using data for a second training. However, Nguyen did not teach “and a plurality of pieces of untrained input data not used for the first training” or “ to generate a plurality of pieces of candidate intermediate data by putting together trained intermediate data given by inputting the plurality of pieces of the trained input data into the machine learning model and untrained intermediate data given by inputting the plurality of pieces of untrained input data into the machine learning model;” or “and selected untrained intermediate data that is the untrained intermediate data already selected,” In an analogous art, Schneider teaches “and a plurality of pieces of untrained input data not used for the first training” See Schneider in page 67, describe “Quantify generalization to new untrained locations We want to test if a model’s classification accuracy differs between images taken from trained locations vs. images taken from untrained locations, across a sparse number (36 in our case) of unique geographic locations of varying environments by measuring performance on trained and untrained locations. Untrained locations mimic expected camera trap usage, as biologists often deploy cameras to new locations over time. We test this using a k-fold validation split”, See Schneider in page 65 describe " Practically, ecologists will often want to classify images from a camera situated at a new location whose images were not included in the training set (Meek et al., 2013)." When Schneider mentions the "images were not included in the training set," this shows that these are data that are not used for an initial training method and are untrained input data. See Schneider in page 65 describe " Practically, ecologists will often want to classify images from a camera situated at a new location whose images were not included in the training set (Meek et al., 2013)." When Schneider mentions the "images were not included in the training set," this shows that these images are data that are not used for an initial training method and therefore is considered untrained input data not run for the first round of training. Further, see Schneider in page 26 describe in first full paragraph "Considering traditional CNNs for re-ID requires a large number of labeled data for each individual and re-training the network for every new individual sighted, both of which are infeasible requirement for animal re-ID. In 1993, Bromley et al. introduced a suitable neural network architecture for this problem, titled a Siamese network, which learns to detect if two input images are similar or dissimilar (Bromley et al., 1994). Once trained, Siamese networks require only one labeled input image of an individual in order to accurately reidentify if an individual in a second image." Since the challenge arises in receiving adequate labelled data to be used for training a model, Schneider shows that using a network architecture called a Siamese network, helps detect only one labeled input image instead of gathering enough data for a model. See Schneider in page 49 describe " The completed generated image was then augmented using colour changes, blurring, grayscale, dropout, and brightening/ darkening. The 25,000 simulated images were split into a 90/10 training/validation split for hyperparameter tuning, and the results reported results on the remaining 123 unseen testing images. Image were resized to 224 * 224 in size, and pixel values were normalized between 0 and 1 as were number of fish considering a regression." Schneider mentions the 123 unseen images are considered data not used for a first round of training. Also, see Schneider in page 67 describe in item 2, Quantify generalization to new untrained locations, the researchers “want to test if a model’s classification accuracy differs between images taken from trained locations vs. images taken from untrained locations, across a sparse number (36 in our case) of unique geographic locations of varying environments by measuring performance on trained and untrained locations. Untrained locations mimic expected camera trap usage, as biologists often deploy cameras to new locations over time. We test this using a k-fold validation split.” Here, Schneider shows using an untrained images (36 mentioned), are untrained input data not used for first round of training. Examiner construes first to mean an initial round or a round of training a model. Further, Schneider teaches “to generate a plurality of pieces of candidate intermediate data by putting together trained intermediate data given by inputting the plurality of pieces of the trained input data into the machine learning model and untrained intermediate data given by inputting the plurality of pieces of untrained input data into the machine learning model;” See Schneider in page 84, in section 4.6 Discussion, describe "In this paper, we tested the reliability of deep learning methods on a modestly sized and unbalanced ecological camera trap data set using both trained and untrained location...For that purpose, we documented a gradient of performance relative to the amount of training data one has as a guide for ecologists considering deep learning methods for their smaller scale data sets. Our work offers a series of conclusions. First, our model performs well for smaller scale data sets considering trained locations, with DenseNet201 achieving 95.6% accuracy. " Here, Schneider shows using trained and untrained sets as input to model, and includes placing both trained and untrained locations where deep learning methods refers to taking the two types of both trained and untrained datasets into a machine learning model which is DenseNet201 in this case. See Schneider from page 62 discuss for more details. Further, see Schneider in page 27, section 2.5 Deep Learning for Species and Animal Re-ID, mention in one method using an “algorithm pre-processes the image by extracting a shell pattern, converting it to grey scale, unravelling the data into a raw input vector, and then training a simple feedforward network”. Here, Schneider shows that raw input vector is viewed as an intermediate form of the data or intermediate data. The examiner construes intermediate data to mean any form of data that is a processed version of the dataset before inputting the data into a model. Here, Schneider shows that the data was pre-processed, and this data is part of the both trained and untrained datasets described in page 84. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Nguyen and incorporate into the teachings of Schneider because both references teach generating candidate input data of trained data used for first training, and untrained input data not used for first training. One of ordinary skill in the art would be motivated to do so because “ For example, MobileNetV2 is designed for efficiency related to size and speed to work on a mobile device and on older PC hardware. It performed competitively with 93.1% accuracy and F1 score of 0.754,” (See Schneider in page 81, section 4.5 Results). However, Nguyen in view of Schneider did not teach “and selected untrained intermediate data that is the untrained intermediate data already selected,” In an analogous art, Zhao teaches “and selected untrained intermediate data that is the untrained intermediate data already selected,” See Zhao in page 4, in section B. Active learning process, describe " We define the AL process as follows: given a small labeled dataset L0 and unlabeled pool U0, in each iteration t, find valuable samples from Ut−1 for labeling and updating the model. Therefore, the core of AL is the criterion for informative sample selection. As mentioned in previous sections, our criterion utilizes the knowledge within the networks. In DS U-Net, some hidden layers are supervised by optimizing the final loss function, and well-learned layer parameters always generate accurate predictions from the hidden layers." Here, Zhao mentions selecting a sample that includes an unlabeled pool, which is the pool of unlabeled data, where unlabeled here also relates to untrained since the model has not yet learned from its ground-truth labels. The data is considered intermediate data because some of these data are part of the hidden layers of the model, which are intermediate parts of the model that processes the data. Further, see Zhao in page 1, abstract mention "They also tend to ignore the intermediate knowledge within networks. In this work, we propose a deep active semi-supervised learning framework, DSAL, combining active learning and semi-supervised learning strategies... In DSAL, a new criterion based on deep supervision mechanism is proposed to select informative samples with high uncertainties and low uncertainties for strong labelers and weak labelers respectively. The internal criterion leverages the disagreement of intermediate features within the deep learning network for active sample selection, which subsequently reduces the computational costs". Here, Zhao describes looking at intermediate data. See Zhao in page 5, section C. Annotation from strong and weak labelers, describe "The pseudocode of the proposed AL process is summarized in Algorithm 1. In each iteration of AL, the proposed algorithm selects both uncertain samples and certain samples from unlabeled dataset U0 based on confidence score and uncertainty score. After that, the segmentation model is fine-tuned with the enlarged dataset. Then the updated model is evaluated on an out-of-bag testing dataset. The process is repeated until exhausting the cost budget or reaching satisfactory performance." Here, Zhao mentions the selecting of samples from an unlabeled data, which unlabeled is construed to also include untrained since the algorithm did not learn from this data. This data of samples also includes intermediate features mentioned from the abstract, which is related to intermediate data that was already selected. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Nguyen and Schneider with the teachings of Zhao by using the teachings by Nguyen and Schneider of generating candidate input data of trained data used for first training, and untrained input data not used for first training, with Zhao’s teaching of selecting untrained intermediate data that is the untrained intermediate data already selected. One of ordinary skill in the art would be motivated to do so because by integrating Zhao’s framework into the methods of Nguyen and Schneider, one with ordinary skill in the art would achieve the goal of providing “classifications at different levels are performed together, which helps the shallower layers to learn discriminative features more efficiently,” (see Zhao in page 4, section III. Proposed method), and “In order to tackle the paucity of labeled data, some approaches different from traditional supervised learning have been developed [9]. Utilizing unlabeled data, semi-supervised learning (SSL) methods [10] involve a self-training process to produce pseudo labels, in which, model updates and pseudo annotations, given the labeled data and model parameters, are performed in an alternating manner. Using unlabeled data in SSL leads to further improvement in model performance... By doing this, active learning (AL) paradigms can be explored to select valuable samples to be labeled with high quality, ” (see Zhao in page 1, I. Introduction, paragraph 2). Claim 2: Regarding claim 2, Nguyen teaches "2. A data processing device comprising processing circuitry..." See Nguyen in col. 5, lines 48-51, where Nguyen describes "For example, the neural network 200 may be executed by one or more processors different from the processors that execute the clustering module 320 or the labelling module 350." Here, Nguyen mentions using processors. Further, Nguyen teaches "to select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data given by inputting, to a machine learning model …," See Nguyen in col. 5, lines 17-18, 23-27, lines 28- 38, describe "In an embodiment, the system receives a dataset in which most of the samples are unlabeled. .. The system performs clustering on these embeddings to identify which data samples are needed to be labeled. These data samples are then labeled, and are added to the labeled sample set, which is provided as input data for the next training iteration. The clustering module 320 receives a set of embeddings from the embedding selection module 310. Each embedding represents a sample that was provided as input during a particular iteration to the neural network 200. The clustering module 320 performs clustering of the received set of embeddings and generates a set of clusters, each cluster comprising one or more embeddings. The clustering module 320 uses a distance metric that represents a distance between any two embeddings and identifies clusters of embeddings that represent embeddings that are close to each other according to the distance metric". Here, Nguyen shows the pieces of intermediate data represented by embeddings, where each embedding represents a sample that was "provided as input during a particular iteration to the neural network model. Nguyen mentions the system receives a data set (i.e. input data) and that “system performs clustering on these embeddings to identify which data samples are needed to be labeled” shows that before labeling, and performing any training, these data samples are considered to be untrained input data. Also, see Nguyen in col. 7, lines 14-22, describe "The deep learning module 120 selects a subset of samples from the input dataset. The subset of samples is selected based on distances between vectors representing the samples." Here, Nguyen describes selecting from the subset of samples of distances between vectors, where distances between vectors representing samples relate to intermediate data, (i.e. selecting one piece of candidate intermediate data), since the examiner construes intermediate data to mean data such as vector representations or data in the middle of being processed before providing output data. Here, Nguyen describes that the unlabeled samples from col. 5, lines 17-18 are also part of the subset of samples that the system used in col. 7, lines 14-22 to generate distances between vectors representing the samples (i.e. intermediate data), which Nguyen shows that the unlabeled or untrained samples also are a part of untrained intermediate data. Further, Nguyen teaches “the processing circuitry more preferentially selecting the one piece of candidate intermediate data having a greater degree of heterogeneity when used for second training subsequent to the first training as compared with selected intermediate data…” See Nguyen in page 11, column 5, lines 63-67 through column 6, lines 1-9 describe " The sample selection module 340 selects samples that are determined to have low similarity to the clusters to which the samples are assigned. Accordingly, the sample selection module 340 excludes samples that are determined to have high similarity to the clusters to which they are assigned." Nguyen describes low similarity relates to items that are less similar to one another and thus more diverse or (i.e. having a greater degree of heterogeneity). Further, see in Nguyen, col. 5, lines 17-38 describe "In an embodiment, the system receives a dataset in which most of the samples are unlabeled. In an iteration, the neural network is trained on only the labeled samples from the original sample dataset. At the end of each iteration, the trained neural network runs a forward pass on the entire dataset to generate embeddings representing sample data at a particular layer. The system performs clustering on these embeddings to identify which data samples are needed to be labeled. These data samples are then labeled, and are added to the labeled sample set, which is provided as input data for the next training iteration. The clustering module 320 receives a set of embeddings from the embedding selection module 310. Each embedding represents a sample that was provided as input during a particular iteration to the neural network 200. The clustering module 320 performs clustering of the received set of embeddings and generates a set of clusters, each cluster comprising one or more embeddings. The clustering module 320 uses a distance metric that represents a distance between any two embeddings and identifies clusters of embeddings that represent embeddings that are close to each other according to the distance metric." See Nguyen in col. 5, lines 63-67 mention "The sample selection module 340 selects a subset of samples from the samples provided by the embedding selection module 310. The sample selection module 340 selects the samples based on the scores determined for the samples.” Here, Nguyen talks about samples provided by the embedding, which relates to selected intermediate data since embedding is interpreted to be a vector or feature representation of data that is not yet an output. Also, see Nguyen in col. 7, lines 50-61, "Accordingly, the sample selection module 340 includes samples of the input data in the subset if the sample has a score indicating low similarity of the sample to the cluster associated with the sample. The labelling module 350 associates each sample of the subset with a label representing an expected result that should be output by the neural network for the sample. In an embodiment, the labelling module receives the associations between samples and the labels from users via a user interface. The deep learning module 120 stores the labelled samples as a training dataset in the training data store 360 for training the neural network 200 for the next iteration." Here, Nguyen shows that that using the intermediate data from the sample selection module 340, which selects intermediate data that is heterogeneous from col. 5, lines 63-67 through column 6, lines 1-9, Nguyen mentions this type of data will be used for training in the next iteration or second training. See Nguyen in col. 6 lines 1-35 for more details. Further, Nguyen teaches “to select one piece of candidate input data corresponding to the one piece of candidate intermediate data selected from among a plurality of pieces of candidate input data which are the plurality of pieces of untrained input data to be used at a time of the second training.” See Nguyen in col. 5, lines 17-18, 23-27, lines 28- 38, describe “In an embodiment, the system receives a dataset in which most of the samples are unlabeled... The system performs clustering on these embeddings to identify which data samples are needed to be labeled. These data samples are then labeled, and are added to the labeled sample set, which is provided as input data for the next training iteration...The clustering module 320 receives a set of embeddings from the embedding selection module 310. Each embedding represents a sample that was provided as input during a particular iteration to the neural network 200. The clustering module 320 performs clustering of the received set of embeddings and generates a set of clusters, each cluster comprising one or more embeddings. The clustering module 320 uses a distance metric that represents a distance between any two embeddings and identifies clusters of embeddings that represent embeddings that are close to each other according to the distance metric". Here, Nguyen shows the pieces of intermediate data represented by subset of embeddings, where each embedding represents a sample that was "provided as input during a particular iteration to the neural network model." For the "next training iteration" relates to second round of training. Nguyen mentions “system performs clustering on these embeddings to identify which data samples are needed to be labeled” shows that before labeling, and performing any training, these data samples are considered to be untrained input data. Nguyen mentions that the system receives a data set (i.e. input data) and that the “system performs clustering on these embeddings to identify which data samples are needed to be labeled” shows that before labeling, and performing any training, these data samples are considered to be untrained input data, which are also a part of intermediate data. Also, see Nguyen in col. 7, lines 14-22, describe "The deep learning module 120 selects a subset of samples from the input dataset. The subset of samples is selected based on distances between vectors representing the samples." Here, Nguyen describes selecting from the subset of samples of distances between vectors, (i.e. selecting one piece of candidate intermediate data), since the examiner construes intermediate data to mean data such as vector representations or data in the middle of being processed before providing output data. Here, Nguyen describes that the unlabeled samples from col. 5, lines 17-18 are also part of the subset of samples that the system used in col. 7, lines 14-22 to generate distances between vectors representing the samples (i.e. intermediate data), which Nguyen shows that the unlabeled or untrained samples also are a part of untrained intermediate data. However, Nguyen did not teach "a plurality of pieces of untrained input data not used for first training in the machine learning model," or “…which is selected untrained intermediate data that is the untrained intermediate data already selected;” In an analogous art, Schneider teaches "a plurality of pieces of untrained input data not used for first training in the machine learning model," See Schneider in page 65 describe " Practically, ecologists will often want to classify images from a camera situated at a new location whose images were not included in the training set (Meek et al., 2013)." When Schneider mentions the "images were not included in the training set," this shows that these images are data that are not used for an initial training method and therefore is considered untrained input data not run for the first round of training. Further, see Schneider in page 26 describe in first full paragraph "Considering traditional CNNs for re-ID requires a large number of labeled data for each individual and re-training the network for every new individual sighted, both of which are infeasible requirement for animal re-ID. In 1993, Bromley et al. introduced a suitable neural network architecture for this problem, titled a Siamese network, which learns to detect if two input images are similar or dissimilar (Bromley et al., 1994). Once trained, Siamese networks require only one labeled input image of an individual in order to accurately reidentify if an individual in a second image." Since the challenge arises in receiving adequate labelled data to be used for training a model, Schneider shows that using a network architecture called a Siamese network, helps detect only one labeled input image instead of gathering enough data for a model. See Schneider in page 49 describe " The completed generated image was then augmented using colour changes, blurring, grayscale, dropout, and brightening/ darkening. The 25,000 simulated images were split into a 90/10 training/validation split for hyperparameter tuning, and the results reported results on the remaining 123 unseen testing images. Image were resized to 224 * 224 in size, and pixel values were normalized between 0 and 1 as were number of fish considering a regression." Schneider mentions the 123 unseen images are considered data not used for a first round of training. Also, see Schneider in page 67 describe in item 2, Quantify generalization to new untrained locations, the researchers “want to test if a model’s classification accuracy differs between images taken from trained locations vs. images taken from untrained locations, across a sparse number (36 in our case) of unique geographic locations of varying environments by measuring performance on trained and untrained locations. Untrained locations mimic expected camera trap usage, as biologists often deploy cameras to new locations over time. We test this using a k-fold validation split.” Here, Schneider shows using an untrained images (36 mentioned), are untrained input data not used for first round of training. Examiner construes first to mean an initial round or a round of training a model. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Nguyen and incorporate into the teachings of Schneider because both references teach to select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data and select one piece of candidate input data corresponding to the one piece of candidate intermediate data selected from among a plurality of pieces of candidate input data which are the plurality of pieces of untrained input data to be used at a time of the second training. One of ordinary skill in the art would be motivated to do so because “ For example, MobileNetV2 is designed for efficiency related to size and speed to work on a mobile device and on older PC hardware. It performed competitively with 93.1% accuracy and F1 score of 0.754,” (See Schneider in page 81, section 4.5 Results). However, Nguyen in view of Schneider did not teach “which is selected untrained intermediate data that is the untrained intermediate data already selected,” In an analogous art, Zhao teaches “which is selected untrained intermediate data that is the untrained intermediate data already selected,” See Zhao in page 4, in section B. Active learning process, describe " We define the AL process as follows: given a small labeled dataset L0 and unlabeled pool U0, in each iteration t, find valuable samples from Ut−1 for labeling and updating the model. Therefore, the core of AL is the criterion for informative sample selection. As mentioned in previous sections, our criterion utilizes the knowledge within the networks. In DS U-Net, some hidden layers are supervised by optimizing the final loss function, and well-learned layer parameters always generate accurate predictions from the hidden layers." Here, Zhao mentions selecting a sample that includes an unlabeled pool, which is the pool of unlabeled data, where unlabeled here also relates to untrained since the model has not yet learned from its ground-truth labels. The data is considered intermediate data because some of these data are part of the hidden layers of the model, which are intermediate parts of the model that processes the data. Further, see Zhao in page 1, abstract mention "They also tend to ignore the intermediate knowledge within networks. In this work, we propose a deep active semi-supervised learning framework, DSAL, combining active learning and semi-supervised learning strategies... In DSAL, a new criterion based on deep supervision mechanism is proposed to select informative samples with high uncertainties and low uncertainties for strong labelers and weak labelers respectively. The internal criterion leverages the disagreement of intermediate features within the deep learning network for active sample selection, which subsequently reduces the computational costs". Here, Zhao describes looking at intermediate data. See Zhao in page 5, section C. Annotation from strong and weak labelers, describe "The pseudocode of the proposed AL process is summarized in Algorithm 1. In each iteration of AL, the proposed algorithm selects both uncertain samples and certain samples from unlabeled dataset U0 based on confidence score and uncertainty score. After that, the segmentation model is fine-tuned with the enlarged dataset. Then the updated model is evaluated on an out-of-bag testing dataset. The process is repeated until exhausting the cost budget or reaching satisfactory performance." Here, Zhao mentions the selecting of samples from an unlabeled data, which unlabeled is construed to also include untrained since the algorithm did not learn from this data. This data of samples also includes intermediate features mentioned from the abstract, which is related to intermediate data that was already selected. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Nguyen and Schneider with the teachings of Zhao by using the teachings by Nguyen and Schneider of generating candidate to select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data and select one piece of candidate input data corresponding to the one piece of candidate intermediate data selected from among a plurality of pieces of candidate input data which are the plurality of pieces of untrained input data to be used at a time of the second training, with Zhao’s teaching of selecting untrained intermediate data that is the untrained intermediate data already selected. One of ordinary skill in the art would be motivated to do so because by integrating Zhao’s framework into the methods of Nguyen and Schneider, one with ordinary skill in the art would achieve the goal of providing “classifications at different levels are performed together, which helps the shallower layers to learn discriminative features more efficiently,” (see Zhao in page 4, section III. Proposed method), and “In order to tackle the paucity of labeled data, some approaches different from traditional supervised learning have been developed [9]. Utilizing unlabeled data, semi-supervised learning (SSL) methods [10] involve a self-training process to produce pseudo labels, in which, model updates and pseudo annotations, given the labeled data and model parameters, are performed in an alternating manner. Using unlabeled data in SSL leads to further improvement in model performance... By doing this, active learning (AL) paradigms can be explored to select valuable samples to be labeled with high quality, ” (see Zhao in page 1, I. Introduction, paragraph 2). Claim 4: Regarding claim 4, the claim recites similar limitations as corresponding independent claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 5: Regarding claim 5, the claim recites similar limitations as corresponding independent claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Claims 3 and 6 are rejected under 35 U.S.C. 103 over Nguyen, in view of Yoo, D. et al. in "Learning Loss for Active Learning," published for a conference held from June 15-20, 2019 in Conference on Computer Vision and Pattern Recognition (CVPR), available at http://openaccess.thecvf.com/content_CVPR_2019/papers/Yoo_Learning_Loss_for_Active_Learning_CVPR_2019_paper.pdf, (hereafter, Yoo), and further in view of Zhao. Claim 3: Regarding claim 3, Nguyen teaches “3. A data processing device comprising processing circuitry …” See Nguyen in col. 5, lines 48-51, where Nguyen describes "For example, the neural network 200 may be executed by one or more processors different from the processors that execute the clustering module 320 or the labelling module 350." Here, Nguyen mentions using processors, which are part of a processing circuitry. Further, Nguyen teaches “the processing circuitry more preferentially selecting the one piece of candidate intermediate data having a greater degree of heterogeneity when used for second training subsequent to the first training as compared with trained intermediate data …” See Nguyen in page 11, column 5, lines 63-67 through column 6, lines 1-9 describe “The sample selection module 340 selects samples that are determined to have low similarity to the clusters to which the samples are assigned. Accordingly, the sample selection module 340 excludes samples that are determined to have high similarity to the clusters to which they are assigned.” Nguyen describes low similarity relates to items that are less similar to one another and thus more diverse or (i.e. having a greater degree of heterogeneity). See Nguyen column 1, lines 40-43, describe “The samples of the subset are provided as a second training dataset for retraining the neural network. This process may be repeated.” Here, Nguyen shows for data to be used for second training, this is part of a second training round. Further, see Nguyen from col. 7, lines 14-16, and 17-19, where Nguyen describes “The subset of samples is selected based on distances between vectors representing the samples...The deep learning module 120 stores the subset of samples in the training data store 360 as a training dataset.” Nguyen here describes selecting the trained intermediate data which was data that was already selected. Also, see Nguyen in page 9, col. 1, lines 35-41, describe "The neural network is trained using a first training dataset. The neural network is executed to process samples from an input dataset. A vector representation of each sample processed by the neural network is obtained from a particular hidden layer of nodes. A subset of samples from the input dataset is obtained based on distances between vectors representing the samples." Here, Nguyen shows that for trained intermediate data, the intermediate data relates to a vector representation of each sample of the trained neural network on a first training dataset. The selection is processing only the samples from an input. Further, see Nguyen teach “… and to select one piece of candidate input data corresponding to the one piece of candidate intermediate data selected from among a plurality of pieces of candidate input data which are the plurality of pieces of untrained input data to be used at a time of the second training,” See Nguyen in col. 5, lines 17-18, 23-27, lines 28- 38, describe “In an embodiment, the system receives a dataset in which most of the samples are unlabeled... The system performs clustering on these embeddings to identify which data samples are needed to be labeled. These data samples are then labeled, and are added to the labeled sample set, which is provided as input data for the next training iteration...The clustering module 320 receives a set of embeddings from the embedding selection module 310. Each embedding represents a sample that was provided as input during a particular iteration to the neural network 200. The clustering module 320 performs clustering of the received set of embeddings and generates a set of clusters, each cluster comprising one or more embeddings. The clustering module 320 uses a distance metric that represents a distance between any two embeddings and identifies clusters of embeddings that represent embeddings that are close to each other according to the distance metric". Here, Nguyen shows the pieces of intermediate data represented by subset of embeddings, where each embedding represents a sample that was "provided as input during a particular iteration to the neural network model." For the "next training iteration" relates to second round of training. Nguyen mentions “system performs clustering on these embeddings to identify which data samples are needed to be labeled” shows that before labeling, and performing any training, these data samples are considered to be untrained input data. Later, see Nguyen column 1, lines 40-43, describe “The samples of the subset are provided as a second training dataset for retraining the neural network. This process may be repeated.” Here, Nguyen shows for data to be used for second training, this is part of a second training round. Further, Nguyen teaches “to select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data given by inputting, to a machine learning model,” See Nguyen in col. 5, lines 17-18, 23-27, lines 28- 38, describe “In an embodiment, the system receives a dataset in which most of the samples are unlabeled... The system performs clustering on these embeddings to identify which data samples are needed to be labeled. These data samples are then labeled, and are added to the labeled sample set, which is provided as input data for the next training iteration...The clustering module 320 receives a set of embeddings from the embedding selection module 310. Each embedding represents a sample that was provided as input during a particular iteration to the neural network 200. The clustering module 320 performs clustering of the received set of embeddings and generates a set of clusters, each cluster comprising one or more embeddings. The clustering module 320 uses a distance metric that represents a distance between any two embeddings and identifies clusters of embeddings that represent embeddings that are close to each other according to the distance metric". Here, Nguyen shows the pieces of intermediate data represented by subset of embeddings, where each embedding represents a sample that was "provided as input during a particular iteration to the neural network model." For the "next training iteration" relates to second round of training. Nguyen mentions system receives a data set (i.e. input data) and that “system performs clustering on these embeddings to identify which data samples are needed to be labeled” shows that before labeling, and performing any training, these data samples are considered to be untrained input data. Also, see Nguyen in col. 7, lines 14-22, describe "The deep learning module 120 selects a subset of samples from the input dataset. The subset of samples is selected based on distances between vectors representing the samples." Here, Nguyen describes selecting from the subset of samples of distances between vectors, where vectors representing samples relate to intermediate data (i.e. selecting one piece of candidate intermediate data), since the examiner construes intermediate data to mean data such as vector representations or data in the middle of being processed before providing output data. Here, Nguyen describes that the unlabeled samples from col. 5, lines 17-18 are also part of the subset of samples that the system used in col. 7, lines 14-22 to generate distances between vectors representing the samples (i.e. intermediate data), which Nguyen shows that the unlabeled or untrained samples also are a part of untrained intermediate data. However, Nguyen did not teach “a plurality of pieces of untrained input data not used for first training in the machine learning model,” or “… and selected intermediate data which is selected untrained intermediate data that is the untrained intermediate data already selected,” In an analogous art, Yoo teaches “… a plurality of pieces of untrained input data not used for first training in the machine learning model,” See Yoo mention in page 95, section 3.1, Overview, “In this section, we formally define the active learning scenario with the proposed loss prediction module. In this scenario, we have a set of models composed of a target model Θtarget and a loss prediction module Θloss. The loss prediction module is attached to the target model as illustrated in Figure 1-(a). The target model conducts the target task as ŷ = Θtarget(x), while the loss prediction module predicts the loss l^=Θloss(h). Here, h is a feature set of x extracted form several hidden layers of Θtarget. In most real-world learning problems, we can gather a large pool of unlabeled data UN at once. The subscript N denotes the number of data points. Then, we uniformly sample K data points at random from the unlabeled pool, and ask human oracles to annotate them to construct an initial labeled dataset L0K. The subscript 0 means it is the initial stage. This process reduces the size of the unlabeled pool as U0N−K. Once the initially labeled dataset L0K is obtained, we jointly learn an initial target model Θ0target and an initial loss prediction module Θ0loss. After initial training, we evaluate all the data points in the unlabeled pool by the loss prediction module to obtain data-loss pairs {(x,l^)|x∈U0N−K}.” Yoo mentions gathering a large pool of unlabeled data and sample K data points from the unlabeled pool to later label these points. While some sample data points were labeled ( i.e. select one piece of candidate data ) and could participate in the model training process, other data points were also not labeled for an initial round of training. These non-labeled data are considered pieces of untrained input data not used for first training in the machine learning model. Yoo then “evaluate all the data points in the unlabeled pool” , which shows inputting untrained input data not used for first training in the machine learning model. Here, Yoo mentions that unlabeled samples cannot participate in supervised training since they lack labels, so these unlabeled samples are not used in a first round of training to train the model. Yoo also described from section 3.1, page 95 that “h is a feature set of x extracted form several hidden layers” , and see Yoo in page 95, part of section 2. Related research mention “This method is directly applicable to any task and network architecture since it depends on intermediate features rather than the task-specific outputs “ Yoo here mentions using hidden layers or intermediate features, which are part of untrained intermediate data. Also, the term intermediate data is construed to mean any form of feature representations, vector embeddings, hidden layer or middle layer steps of the model. See Yoo in figure 1, part b for more details. PNG media_image1.png 557 537 media_image1.png Greyscale It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Nguyen and incorporate into the teachings of Yoo because both references teach to select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data given by inputting, to a machine learning model, a plurality of pieces of untrained input data not used for first training in the machine learning model. One of ordinary skill in the art would be motivated to do so because “The learning process is efficient as the loss prediction module Θsloss has been designed to contain a small number of parameters but to utilize rich mid-level representations h of the target model. This loss prediction module will pick the most informative data points and ask human oracles to annotate them for the next active learning stage s + 1,” (see Yoo describe in pages 96-97, section 3.3 Learning loss). However, Nguyen in view of Yoo did not teach “… and selected intermediate data which is selected untrained intermediate data that is the untrained intermediate data already selected,” In an analogous art, Zhao teaches “…and selected intermediate data which is selected untrained intermediate data that is the untrained intermediate data already selected,” See Zhao in page 4, in section B. Active learning process, describe " We define the AL process as follows: given a small labeled dataset L0 and unlabeled pool U0, in each iteration t, find valuable samples from Ut−1 for labeling and updating the model. Therefore, the core of AL is the criterion for informative sample selection. As mentioned in previous sections, our criterion utilizes the knowledge within the networks. In DS U-Net, some hidden layers are supervised by optimizing the final loss function, and well-learned layer parameters always generate accurate predictions from the hidden layers." Here, Zhao mentions selecting a sample that includes an unlabeled pool, which is the pool of unlabeled data, where unlabeled here also relates to untrained since the model has not yet learned from its ground-truth labels to incorporate them before training. The data is considered intermediate data because some of these data are part of the hidden layers of the model, which are intermediate parts of the model that processes the data. Further, see Zhao in page 1, abstract mention "They also tend to ignore the intermediate knowledge within networks. In this work, we propose a deep active semi-supervised learning framework, DSAL, combining active learning and semi-supervised learning strategies... In DSAL, a new criterion based on deep supervision mechanism is proposed to select informative samples with high uncertainties and low uncertainties for strong labelers and weak labelers respectively. The internal criterion leverages the disagreement of intermediate features within the deep learning network for active sample selection, which subsequently reduces the computational costs". Here, Zhao describes looking at intermediate data. See Zhao in page 5, section C. Annotation from strong and weak labelers, describe "The pseudocode of the proposed AL process is summarized in Algorithm 1. In each iteration of AL, the proposed algorithm selects both uncertain samples and certain samples from unlabeled dataset U0 based on confidence score and uncertainty score. After that, the segmentation model is fine-tuned with the enlarged dataset. Then the updated model is evaluated on an out-of-bag testing dataset. The process is repeated until exhausting the cost budget or reaching satisfactory performance." Here, Zhao mentions the selecting of samples from an unlabeled data, which unlabeled is construed to also include untrained since the algorithm did not learn from this data. This data of samples also includes intermediate features mentioned from the abstract, which is related to intermediate data that was already selected. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Nguyen and Yoo with the teachings of Zhao by using the teachings by Nguyen and Yoo to teach to select one piece of candidate intermediate data from among a plurality of pieces of candidate intermediate data which are untrained intermediate data given by inputting, to a machine learning model, a plurality of pieces of untrained input data not used for first training in the machine learning model, and incorporate with Zhao’s teaching of selecting untrained intermediate data that is the untrained intermediate data already selected. One of ordinary skill in the art would be motivated to do so because by integrating Zhao’s framework into the methods of Nguyen and Yoo, one with ordinary skill in the art would achieve the goal of providing “classifications at different levels are performed together, which helps the shallower layers to learn discriminative features more efficiently,” (see Zhao in page 4, section III. Proposed method), and “In order to tackle the paucity of labeled data, some approaches different from traditional supervised learning have been developed [9]. Utilizing unlabeled data, semi-supervised learning (SSL) methods [10] involve a self-training process to produce pseudo labels, in which, model updates and pseudo annotations, given the labeled data and model parameters, are performed in an alternating manner. Using unlabeled data in SSL leads to further improvement in model performance... By doing this, active learning (AL) paradigms can be explored to select valuable samples to be labeled with high quality, ” (see Zhao in page 1, I. Introduction, paragraph 2). Claim 6 : Regarding claim 6, the claim recites similar limitations as corresponding independent claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENWEI ZENG whose telephone number is (571)272-7111. The examiner can normally be reached Monday-Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WenWei Zeng/Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Dec 21, 2023
Application Filed
Jul 27, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month