Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 11 is objected to because of the following informalities: Claim 11 recites incorrect claim number that the claim depends on. Appropriate correction is required. For examination purposes, the examiner will consider claim 11 as a dependent claim of claim 10 because claim 11 recites the first machine learning model, which is first introduced in claim 10.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title.
Claims 1-3, 5, 7, 9-16, 26-27, 29, 31 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1,
Step 1:
Claim 1 recites a method, one of the four statutory categories of patentable subject matter.
Step 2A, Prong I:
Claim 1 further recites the limitations of:
“dividing each epoch, of a first set of training data epochs in an original dataset having an input space, into a second set, N, of batches” The limitations recites a mental process. A person can mentally perform the division process to divide a training data set into many batches for an epoch. For example, given a collection of images that can be used to train a machine learning model, the person can mentally or manually determine how many images will be used to train during one iteration, which corresponds to a batch, and the person can mentally divide to determine the corresponding number of batches according to the given set of images.
“generating a third set, K, of subsets of samples by selecting, within each batch from every second set, N, of batches, a respective plurality of subsets of one or more samples, each subset being configured to be different from another subset” The limitations recites a mental process. A person can mentally select a subset of samples within a batch of training data. For example, the person can select some images from a batch of images, where in those selected images correspond to the subset of samples, and the selection can be mentally performed within a human’s mind.
“determining, …, a fourth set of clusters of data using the determined third set, K, of subsets of samples as input” The limitations recites a mental process. A person can mentally determine clusters of data from a subset of data. For example, given a collection of images, which corresponds to a subset of samples (e.g., a page from an album of images), the person can mentally select some images out and put them together, which corresponds to a cluster of data. Such determination is a mental process and can be performed within a human’s mind.
“selecting a fifth set of clusters from the fourth set of clusters based on a criterion of relevance” The limitations recites a mental process. A person can mentally select a set of clusters from another set of clusters based on some relevant criteria. For example, given many collections of images, a person can mentally select those collections of images that fit their interest such as collections of images about dogs. Such determination is a mental process and can be performed within a human’s mind.
Step 2A, Prong II:
Claim 1 recites the following additional elements:
“A computer-implemented method, performed by a first node, the method being for handling data augmentation, the first node operating in a communications system” These additional elements are a high-level recitation of generic computer components used as a tool, and does not provide integration into a practical application.
“… determining, using machine learning, a fourth set of clusters of data …” This additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. The limitation recites the black-box application of applying a machine learning model to determine clusters of data without providing further information regarding the type of machine learning model or an unconventional machine learning algorithm/model to determine clusters of data. The machine learning model, as used herein, also does not provide an improvement toward machine learning model practice, or an improvement to any computer elements.
“generating samples in each cluster of the selected fifth set of clusters, and refraining from generating samples in clusters of the fourth set of clusters being excluded from the fifth set of clusters” This additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. The limitation recites the black-box application of generating samples by applying selected set of clusters and only generating samples using the selected clusters. The limitation merely recites the application of clusters to generate samples data without details regarding how the samples are generated based on the clusters, or any unconventional practice to generate samples from the clusters or improvement toward any computer elements.
“generating a sixth set of augmented samples in the input space of the original dataset, by using the generated samples and applying a reverse projection approach to transform the generated samples into transformed samples of the input space of the original dataset, added to the original dataset” This additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. The limitation recites the black-box application of a machine learning technique being a reverse projection approach to augment data for further training the machine learning model. The limitation merely recites the application of a technique to augment data for training, without providing details regarding how the technique is implemented or performing the technique in an unconventional manner or improvement toward any computer elements.
Step 2B:
When considered individually or in combination, the additional limitations and elements of claim 1 do not amount to significantly more than the judicial exception for the same reasons discussed above as to why the additional limitations do not integrate the abstract idea into a practical application. The additional elements outlined in Step 2A performing functions as designed simply accomplish execution of the abstract ideas.
The additional element “A computer-implemented method, performed by a first node, the method being for handling data augmentation, the first node operating in a communications system” is a high-level recitation of generic computer components used as a tool, and does not amount to significantly more than the judicial exception for the same reasons discussed above as to why the additional limitations do not integrate the abstract idea into a practical application.
The additional element “… determining, using machine learning, a fourth set of clusters of data …” recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not amount to significantly more than the judicial exception for the same reasons discussed above.
The additional element “generating samples in each cluster of the selected fifth set of clusters, and refraining from generating samples in clusters of the fourth set of clusters being excluded from the fifth set of clusters” recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not amount to significantly more than the judicial exception for the same reasons discussed above.
The additional element “generating a sixth set of augmented samples in the input space of the original dataset, by using the generated samples and applying a reverse projection approach to transform the generated samples into transformed samples of the input space of the original dataset, added to the original dataset” recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not amount to significantly more than the judicial exception for the same reasons discussed above.
In conclusions from above for the elements considered as a mental process, elements reciting high-level recitation of generic computer components used as a tool, and elements reciting a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f) are carried over and do not provide significantly more than the abstract idea. Looking at the limitations in combination and the claims as a whole does not change this conclusion and the claim is ineligible.
Therefore, additional limitations of claim 1 do not amount to significantly more than the judicial exception.
Thus, claim 1 recites abstract ideas with additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception.
Therefore, claim 1 is not patent eligible.
Regarding claim 2 depends on claim 1, thus the rejection of claim 1 is incorporated.
“The method according to claim 1, wherein the criterion of relevance is a respective indication exceeding a threshold, the respective indication being of a ratio of a respective number of samples of a respective class in each respective cluster of the fourth set of clusters of data, to a respective total number of samples in the respective cluster” This limitation recites a mental process as well as a mathematical concept. The calculation of a ratio of a respective number of samples of a respective class to a respective total number of samples within a cluster of data is a mathematical concept that requires division calculation and can be performed mentally by a person. Furthermore, the comparison of the calculated ratio to a threshold is also a mental process, wherein such comparison is a mental evaluation between two values and can be performed within a human’s mind.
Thus, claim 2 recites abstract ideas rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 2 is not patent eligible.
Regarding claim 3 depends on claim 1, thus the rejection of claim 1 is incorporated.
“The method according to claim 1, wherein the reverse projection approach is one of: a) stochastic, and b) processing each of the generated samples in each cluster of the selected fifth set of clusters, in parallel, in a respective node of a seventh set of nodes” The additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application or amount to significantly more than the judicial exception. The claim merely recites the application of the reverse projection approach to be a stochastic process or a parallel process without reciting any improvement over the conventional technique or unconventional practice using the technique.
Thus, claim 3 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 3 is not patent eligible.
Regarding claim 5 depends on claim 1, thus the rejection of claim 1 is incorporated.
“The method according to claim 1, wherein the generating is performed by using combinatorial sampling” The additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application or amount to significantly more than the judicial exception. The claim merely recites the technique used to generate samples from clusters without providing significantly more such as unconventional method to perform the combinatorial sampling or improvement toward such technique or improvement toward any computer elements.
Thus, claim 5 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 3 is not patent eligible.
Regarding claim 7 depends on claim 1, thus the rejection of claim 1 is incorporated.
“The method according to claim 1, wherein the generating of the sixth set of augmented samples in the input space comprises minimizing sum-of-squares values of variables in the input space” This limitation recites a mental process as well as a mathematical concept. The calculation of minimization sum-of-squares values of variables in the input space is a mathematical concept that can be performed within a human’s mind.
Thus, claim 7 recites abstract ideas rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 7 is not patent eligible.
Regarding claim 9 depends on claim 1, thus the rejection of claim 1 is incorporated.
“The method according to claim 1, further comprising: providing a further indication indicating the generated sixth set of augmented samples to a second node operating in the communications system” The additional element recites an additional element of an insignificant extra-solution activity as identified in MPEP 2106.05(g) of mere data gathering, and does not provide integration into a practical application. The element further recites an additional element of a well-understood, routine, conventional activity as identified in MPEP 2106.05(d) of receiving or transmitting data over a network and does not provide integration into a practical application or amount to significantly more than the judicial exception
Thus, claim 9 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 9 is not patent eligible.
Regarding claim 10 depends on claim 1, thus the rejection of claim 1 is incorporated.
“The method according to claim 1, further comprising: determining a first machine learning model of an event in the communications system using as input the generated sixth set of augmented samples” The additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application or amount to significantly more than the judicial exception. The element recites a black-box application of applying input of augmented samples to a machine learning model without reciting any improvement toward the algorithm of the machine learning model or improvement toward computer elements.
Thus, claim 10 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 10 is not patent eligible.
Regarding claim 11 depends on claim 10, thus the rejection of claim 10 is incorporated.
“The method according to claim 10, further comprising: initiating performance of an action to manage a predicted occurrence of the event according to the determined first machine learning model” The additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application or amount to significantly more than the judicial exception. The element recites a black-box application of performing a management action based on a prediction obtained from the machine learning model without reciting any improvement toward the algorithm of the machine learning model or improvement toward computer elements.
Thus, claim 11 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 11 is not patent eligible.
Regarding claim 12, which recites a method, one of the four statutory categories of patentable subject matter. The applicant is further directed to the rejection of claim 1 above, because the claim recites similar limitations and processing steps.
Regarding claim 13 depends on claim 12, thus the rejection of claim 12 is incorporated. Claim 13 is rejected under the same rationale as claim 10. The applicant is further directed to the rejection of claim 10 above, because the claim recites similar limitations and processing steps.
Regarding claim 14 depends on claim 13, thus the rejection of claim 13 is incorporated. Claim 14 is rejected under the same rationale as claim 11. The applicant is further directed to the rejection of claim 11 above, because the claim recites similar limitations and processing steps.
Regarding claim 15, which recites a system, one of the four statutory categories of patentable subject matter. The applicant is further directed to the rejection of claim 1 above, because the claim recites similar limitations and processing steps.
Regarding claim 16 depends on claim 15, thus the rejection of claim 15 is incorporated. Claim 16 is rejected under the same rationale as claim 2. The applicant is further directed to the rejection of claim 2 above, because the claim recites similar limitations and processing steps.
Regarding claim 26, which recites a system, one of the four statutory categories of patentable subject matter. The applicant is further directed to the rejection of claim 1 above, because the claim recites similar limitations and processing steps.
Regarding claim 27 depends on claim 26, thus the rejection of claim 26 is incorporated. Claim 27 is rejected under the same rationale as claim 10. The applicant is further directed to the rejection of claim 10above, because the claim recites similar limitations and processing steps.
Regarding claim 29 depends on claim 1, thus the rejection of claim 1 is incorporated.
“A computer program product comprising a non-transitory computer readable medium storing a computer program; comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out the method according to claim 1” This additional element is a high-level recitation of generic computer components used as a tool, and does not provide integration into a practical application or amount to significantly more than the judicial exception.
Thus, claim 29 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 29 is not patent eligible.
Regarding claim 31 depends on claim 12, thus the rejection of claim 12 is incorporated.
“A computer program product comprising a non- transitory computer readable medium storing a computer program; comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out the method according to claim 12” This additional element is a high-level recitation of generic computer components used as a tool, and does not provide integration into a practical application or amount to significantly more than the judicial exception.
Thus, claim 21 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 21 is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 6-16, 26-27, 29, 31 are rejected under 35 U.S.C. 103 as being unpatentable over Jayadeva et.al (NPL: EigenSample: A non-iterative technique for adding samples to small datasets) in view of Yang et.al (US 20200302334 A1), further in view of Last et.al (NPL: Oversampling for Imbalanced Learning Based on K-Means and SMOTE).
Regarding claim 1,
Jayadeva teaches or at least suggests “A computer-implemented method, …, the method being for handling data augmentation,” (Page 1 section 1 “it has been shown that augmenting the data set in this manner can improve generalization … In this paper, we suggest an alternative way … Our method, EigenSample” Jayadeva discloses a computer-implemented method for handling data augmentation because EigenSample is expressly directed to augmenting a training dataset by generating synthetic examples that preserve the underlying cluster information of the original dataset. Jayadeva explains that the original data are projected into a lower-dimensional eigenspace, synthetic samples are generated in that space, and the resulting samples are used to augment the training dataset. Accordingly, Jayadeva teaches the claimed computer-implemented data-augmentation method.)
Jayadeva teaches or at least suggests “generating a sixth set of augmented samples in the input space of the original dataset, by using the generated samples and applying a reverse projection approach to transform the generated samples into transformed samples of the input space of the original dataset, added to the original dataset” (Page 1 section 1 “Our method, EigenSample, is based on generating synthetic examples in such a manner that the cluster information in the dataset is minimally distorted. In order to do this, we project the mean-centered dataset into a lower-dimensional space spanned by the highest energy containing eigenvectors … Finally, we project back the synthetic samples generated in this space to the input space by solving an optimization problem”, Page 4 section 3.2 “we demonstrate our approach for reverse projecting the dataset, … One can observe that the images retain the characteristics of the original image sets … Also, the reconstruction error is much lower by using our approach as compared to using the conventional approach of reverse-projection from the eigen sub-space”, and Page 7 section 4 “Our algorithm, EigenSample, for augmenting datasets, is summarized in Algorithm 1” Jayadeva discloses generating augmented samples in the input space of the original dataset by reverse projecting generated samples from the lower-dimensional space back into the original input space. In particular, Jayadeva generates synthetic samples in the lower-dimensional eigenspace and then projects those synthetic samples back to the input space by solving an optimization problem. Jayadeva expressly describes this operation as its approach for “reverse projecting” the dataset and explains that the reconstructed samples retain characteristics of the original data and are used to augment the dataset. Thus, the synthetic samples generated in the projected space correspond to the claimed generated samples, the inverse/reverse-projection operation corresponds to the claimed reverse projection approach, and the resulting reconstructed synthetic samples correspond to the claimed augmented samples in the input space of the original dataset)
Jayadeva does not teach “A computer-implemented method, performed by a first node … the first node operating in a communications system”. However, Yang teaches or at least suggests this part of the limitation (Paragraph 19 “In another aspect, a methodology in the present disclosure reduces accesses to the storage system, and may also reduce communication among compute nodes executing learners”, and Paragraph 22 “In one embodiment, the compute nodes 102, 104 have an ability to simultaneously execute multiple processes, which are referred to as learners” Yang discloses a distributed computing arrangement in which compute nodes execute learner processes and communicate during training. Accordingly, Yang’s compute node corresponds to the claimed first node, and the distributed arrangement in which the compute nodes communicate corresponds to the claimed communications system in which the first node operates.)
Jayadeva does not teach “dividing each epoch, of a first set of training data epochs in an original dataset having an input space, into a second set, N, of batches” However, Yang teaches or at least suggests this limitation (Paragraph 2 “In a training epoch composed by the steps, the whole dataset is loaded, and the training takes a number of training epochs for convergence”, Paragraph 16 “Minibatch refers to a subset of a dataset. Minibatch size can be selected, for example, by a user, and the entire training dataset can be divided into minibatches of the minibatch size”, and Paragraph 31 “Allocating a cache to hold distinct data samples, determining the minibatch workload distribution, and establishing a scheme to assemble a minibatch are repeated until all the minibatch sequences making up an epoch of training dataset are processed” Yang discloses training over multiple epochs, wherein the training dataset is divided into a plurality of minibatches and the minibatch sequences collectively make up an epoch of training. Accordingly, Yang’s training epoch corresponds to the claimed training data epoch, and Yang’s plurality of minibatches corresponds to the claimed second set, N, of batches into which the training data of the epoch is divided.)
Jayadeva does not teach “generating a third set, K, of subsets of samples by selecting, within each batch from every second set, N, of batches, a respective plurality of subsets of one or more samples, each subset being configured to be different from another subset” However, Yang teaches or at least suggests this limitation (Paragraph 27 “A “minibatch sequence” refers to a list of data sample indices that represents a minibatch. In a distributed system, learner processes collectively train with a minibatch in a step. Each learner may acquire the same “minibatch sequence”, and take a slice (or subset), for example, a disjoint slice (or subset), from a minibatch sequence as its local batch … For example, at the beginning of an epoch, the same seed is used by all distributed learners to create a random permutation of the sample indices (a minibatch sequence), which can be subsequently divided into multiple slices (or subsets) … The data samples corresponding to the minibatch slices can be acquired by a data loading procedure in different nodes. For instance, a compute node 210 may load a minibatch slice 204; a compute node 212 may load a minibatch slice 206; and compute node 214 may load a minibatch slice 208” Yang discloses that a minibatch sequence represents a minibatch and may be divided into multiple slices or subsets, with respective compute nodes loading respective minibatch slices. Yang further discloses that the slices or subsets may be disjoint or non-overlapping. Accordingly, Yang’s minibatch corresponds to the claimed batch, the multiple slices or subsets correspond to the claimed third set, K, of subsets of one or more samples, and the disjoint or non-overlapping subsets correspond to the requirement that each subset is different from another subset.)
Before the effective filing date, it would have been obvious to a person of ordinary skill in the art to incorporate the teaching of the EigenSample method to augment data and reverse projecting the augmented data to the training datasets by Jayadeva with the teaching of determining epoch and dividing training datasets into minibatches by Yang. The motivation to do so is referred to in Yang’s disclosure (paragraph 19 “In one aspect, a methodology in the present disclosure need not change the composition of minibatches to reduce communication. For instance, a method and/or system in some embodiments can assemble pre-defined minibatches without using different samples from the specified samples. In another aspect, a methodology in the present disclosure reduces accesses to the storage system, and may also reduce communication among compute nodes executing learners. Yet in another aspect, a methodology in the present disclosure provides an option to further reduce communication traffic incurred from data loading by tolerating imbalance of workload on each learner” Yang discloses that its minibatch-based distributed processing reduces accesses to the storage system and may also reduce communication among compute nodes executing learner processes. Yang further explains that predefined minibatches and workload distribution can reduce communication traffic associated with data loading. Accordingly, incorporating Yang’s minibatch and subset processing into Jayadeva’s data-augmentation method would have predictably improved the efficiency of distributed machine-learning training by reducing storage access and inter-node communication. Furthermore, given that Jayadeva teaches augmenting datasets having a limited number of samples by generating additional synthetic samples. A person of ordinary skill in the art would have been motivated to apply Jayadeva’s augmentation technique to Yang’s smaller sample subsets to increase the number and diversity of training samples available within those subsets while retaining Yang’s distributed-processing efficiencies.)
Jayadeva/Yang does not teach “determining, using machine learning, a fourth set of clusters of data using the determined third set, K, of subsets of samples as input” However, Last teaches or at least suggests this limitation (Page 5 section 3.1 “K-means SMOTE consists of three steps: clustering, filtering, and oversampling. In the clustering step, the input space is clustered into k groups using k-means clustering” Last discloses determining a set of clusters from input sample data using k-means clustering. In Last, the input space is first divided into k groups by the k-means algorithm. Accordingly, Last’s k groups correspond to the claimed fourth set of clusters of data, with the sample data supplied to the clustering operation corresponding to the sample subsets generated by the preceding Yang processing.)
Jayadeva/Yang does not teach “selecting a fifth set of clusters from the fourth set of clusters based on a criterion of relevance” However, Last teaches or at least suggests this limitation (Page 5 section 3.1 “K-means SMOTE consists of three steps: clustering, filtering, and oversampling. In the clustering step, the input space is clustered into k groups using k-means clustering. The filtering step selects clusters for oversampling, retaining those with a high proportion of minority class samples” Last discloses selecting only certain clusters for oversampling by applying a filtering criterion to the clusters. In particular, Last evaluates the class composition of each cluster and retains clusters having a sufficiently high proportion of minority-class samples. Thus, Last’s filtering operation corresponds to the claimed criterion of relevance, and the clusters that pass the filtering operation correspond to the claimed fifth set of selected clusters. In other words, Last asks whether a cluster contains enough samples from the minority class to make that cluster suitable for generation of additional samples; if so, the cluster is selected.)
Jayadeva/Yang does not teach “generating samples in each cluster of the selected fifth set of clusters, and refraining from generating samples in clusters of the fourth set of clusters being excluded from the fifth set of clusters” However, Last teaches or at least suggests this limitation (Page 6 section 3.1 “In the clustering step, the input space is clustered into k groups using k-means clustering. The filtering step selects clusters for oversampling, retaining those with a high proportion of minority class samples. It then distributes 5 the number of synthetic samples to generate, assigning more samples to clusters where minority samples are sparsely distributed. Finally, in the oversampling step, SMOTE is applied in each selected cluster to achieve the target ratio of minority and majority instances … The selection of clusters for oversampling is based on each cluster’s proportion of minority and majority instances. By default, any cluster made up of at least 50 % minority samples is selected for oversampling. This behavior can be tuned by adjusting the imbalance ratio threshold” Last discloses generating samples in each selected cluster by applying SMOTE to the selected clusters. Clusters that do not satisfy the filtering criterion are not included among the selected clusters and therefore are not subjected to the SMOTE oversampling and samples generating operation, which teaches or at least suggests refraining from generating samples in excluded clusters, as claimed. Accordingly, Last teaches generating samples in the selected fifth set of clusters while refraining from generating samples in the clusters excluded from that selected set.)
Before the effective filing date, it would have been obvious to a person of ordinary skill in the art to incorporate the teaching of the EigenSample method to augment data and reverse projecting the augmented data to the training datasets by Jayadeva, and the teaching of determining epoch and dividing training datasets into minibatches by Yang with the teaching of determining clusters based on k-means clustering and SMOTE oversampling by Last. The motivation to do so is referred to Last’s disclosure (Page 2 section 1 “This paper suggests the combination of the k-means clustering algorithm in combination with SMOTE to combat some of other oversampler’s shortcomings with a simple-to-use technique. The use of clustering enables the proposed oversampler to identify and target areas of the input space where the generation of artificial data is most effective. The method aims at eliminating both between-class imbalances and within-class imbalances while at the same time avoiding the generation of noisy samples” Last discloses that combining k-means clustering with SMOTE allows the oversampling process to identify and target regions of the input space where generation of artificial data is most effective while avoiding generation of noisy samples. Therefore, a person of ordinary skill in the art would have been motivated to incorporate Last’s cluster-filtering and selective oversampling technique into the data-augmentation method of Jayadeva, as modified by Yang, so that synthetic samples are generated in clusters identified as suitable for augmentation while avoiding generation in less suitable clusters, thereby improving the quality and effectiveness of the augmented training data.)
Regarding claim 2 depends on claim 1, thus the rejection of claim 1 is incorporated.
Last teaches or at least suggests the limitation “The method according to claim 1, wherein the criterion of relevance is a respective indication exceeding a threshold, the respective indication being of a ratio of a respective number of samples of a respective class in each respective cluster of the fourth set of clusters of data, to a respective total number of samples in the respective cluster” (Page 6 section 3.1 “In the clustering step, the input space is clustered into k groups using k-means clustering. The filtering step selects clusters for oversampling, retaining those with a high proportion of minority class samples. It then distributes 5 the number of synthetic samples to generate, assigning more samples to clusters where minority samples are sparsely distributed. Finally, in the oversampling step, SMOTE is applied in each selected cluster to achieve the target ratio of minority and majority instances … The selection of clusters for oversampling is based on each cluster’s proportion of minority and majority instances. By default, any cluster made up of at least 50 % minority samples is selected for oversampling. This behavior can be tuned by adjusting the imbalance ratio threshold” Last discloses that, after the data are divided into clusters, each cluster is evaluated based on the proportion of minority-class samples contained in that cluster. Last explains that a cluster is selected for oversampling when it contains a sufficiently high proportion of minority-class samples, with the default selection being a cluster having at least 50% minority samples. In other words, for each cluster, Last considers the number of minority-class samples relative to the total number of samples in that cluster and uses that proportion to determine whether the cluster should be selected. Accordingly, Last’s number of minority-class samples in a cluster corresponds to the claimed “respective number of samples of a respective class,” the total number of samples in that cluster corresponds to the claimed “respective total number of samples in the respective cluster,” and the resulting minority-class proportion corresponds to the claimed “respective indication.” Last’s requirement that the cluster have a sufficiently high minority-class proportion, such as at least 50%, corresponds to comparing that indication with a threshold. Thus, Last teaches or at least suggests the claimed relevance criterion based on the ratio of samples of a particular class in a cluster to the total samples in that cluster.)
Regarding claim 3 depends on claim 1, thus the rejection of claim 1 is incorporated.
Jayadeva in view of Yang teaches or at least suggests “The method according to claim 1, wherein the reverse projection approach is one of: a) stochastic, and b) processing each of the generated samples in each cluster of the selected fifth set of clusters, in parallel, in a respective node of a seventh set of nodes” (Jayadeva discloses at Page 4 section 3.2 “we demonstrate our approach for reverse projecting the dataset”, whereas Yang discloses at Paragraph 18 “a minibatch can be assembled for use by one or more learners without any data loading—including loading from the storage system or exchanging among distributed learners. Samples may be re-distributed to balance the workload to optimize parallel computation performance”, and Paragraph 22 “A process 110, for example, executes a learner. Each of the compute nodes 102, 104 in the distributed system may execute one or plurality of learners in parallel”. Jayadeva discloses generating synthetic samples and applying a reverse-projection approach to project those generated samples back into the original input space. Yang discloses a distributed computing arrangement in which samples may be redistributed among multiple compute nodes to balance workload and optimize parallel computation performance and further discloses that the compute nodes may execute learner processes in parallel. Thus, Jayadeva provides the generated samples and reverse-projection processing, while Yang provides the parallel multi-node processing arrangement. Accordingly, it would have been obvious to a person of ordinary skill in the art to perform Jayadeva’s reverse-projection processing of the respective generated samples using Yang’s distributed compute nodes in parallel, such that respective generated samples are processed by respective nodes. Such a modification would predictably permit the reverse-projection workload to be distributed among the nodes and processed concurrently, thereby improving computational efficiency and balancing the processing workload.)
Regarding claim 4 depends on claim 3, thus the rejection of claim 3 is incorporated.
Jayadeva in view of Yang teaches or at least suggests “The method according to claim 3, wherein one of: a. the stochastic approach comprises running each generated sample through an optimization routine based on a sub-gradient, and b. the approach using parallel processing in the seventh set of nodes uses a closed form least squares optimization procedure” (Page 5 section 4 “The optimization problem developed in the previous section would need to be solved by an optimization routine which would conventionally be iterative in nature. However, if we consider a least-squares version of the same, we can develop the solution of the optimization problem to be given by a set of linear equations, which can directly be solved (non-iteratively). This is motivated by the formulation of the least-squares SVM … We consider the least-squares version which minimizes a quadratic penalty on the slack variables”. Jayadeva discloses that its reverse-projection optimization may be reformulated as a least-squares version in which the solution is expressed as a set of linear equations that can be directly solved non-iteratively. Thus, Jayadeva’s least-squares formulation corresponds to the claimed least-squares optimization procedure, and the direct, non-iterative solution corresponds to the claimed closed-form solution. Yang, as discussed with respect to claim 3, discloses performing processing in parallel across multiple compute nodes. Accordingly, when Jayadeva’s directly solvable least-squares reverse-projection procedure is implemented using Yang’s parallel-node arrangement, the combination teaches or at least suggests the claimed approach in which the parallel processing performed by the seventh set of nodes uses a closed-form least-squares optimization procedure.)
Regarding claim 6 depends on claim 1, thus the rejection of claim 1 is incorporated.
Jayadeva teaches or at least suggests the limitation “The method according to claim 1, wherein each cluster of the fourth set of clusters has a respective center, and wherein, the generating samples is performed in at least one of: a. within a first distance from the respective center of a respective cluster, b. within a second distance from a respective sample in a projected space belonging to the respective cluster, and c. within a certain distance between the respective center of the respective cluster and the respective sample” (Page 3 section 2 “We now construct lines joining the cluster centers with each point of the respective cluster, and then compute the mid points of these lines … Since the addition of these points does not disturb the cluster centers, we use them to augment the original data set”. Jayadeva discloses determining cluster centers for the projected samples and then generating additional samples by constructing a line between each cluster center and a respective sample belonging to that cluster and selecting the midpoint of that line. The generated midpoint is then used to augment the original dataset. Accordingly, Jayadeva’s cluster center corresponds to the claimed respective center of the respective cluster, and Jayadeva’s point of the respective cluster corresponds to the claimed respective sample. Because the generated sample is the midpoint of the line joining those two points, it necessarily lies between the cluster center and the respective sample. Thus, Jayadeva inherently teaches or at least suggests generating the sample within a certain distance between the respective center of the respective cluster and the respective sample, as recited in claim 6(c).)
Regarding claim 7 depends on claim 1, thus the rejection of claim 1 is incorporated.
Jayadeva teaches or at least suggests the limitation “The method according to claim 1, wherein the generating of the sixth set of augmented samples in the input space comprises minimizing sum-of-squares values of variables in the input space” (Page 5-6 section 4 equation (18)-(21) “we consider a least-squares version of the same, we can develop the solution of the optimization problem to be given by a set of linear equations, which can directly be solved (non-iteratively). This is motivated by the formulation of the least-squares SVM … We consider the least-squares version which minimizes a quadratic penalty on the slack variables … to be given by (18)–(21)” Jayadeva discloses generating augmented samples in the input space using a least-squares reverse-projection formulation. In particular, equation (18) minimizes the squared norm of the reverse-projected input-space variable zi, together with quadratic penalty terms on the slack variables. Thus, the squared-norm term of equation (18) corresponds to the claimed minimizing sum-of-squares values of variables in the input space, because zi represents the sample projected back into the original input space.)
Regarding claim 8 depends on claim 7, thus the rejection of claim 7 is incorporated.
Jayadeva teaches or at least suggest the limitation “The method according to claim 7, wherein the generating of the sixth set of augmented samples in the input space introduces an error when applying the reverse projection approach to transform the generated samples into transformed samples of the input space of the original dataset, and wherein the generating of the sixth set of augmented samples comprises” (Page 5 section 3.2 “Our approach for reverse projecting the dataset … offers a solution which allows us to choose the extreme we wish to use in obtaining the reverse projection. We use an SVR-like strategy, which allows us to choose parameters C and to obtain a suitably augmented dataset. We introduce smoothness constraints which depend on the values of which allow us to control the quality of the dataset obtained by the reverse projection. The parameter C in the objective function allows us to set the tolerance to which we can obtain the original dataset back. Hence, EigenSample permits us to choose a tradeoff between the projection error and the norm of the solution” Jayadeva discloses that its reverse-projection process involves a projection error when reconstructing or projecting generated samples back into the original input space. Jayadeva further explains that parameters are selected to control the quality of the reverse-projected augmented dataset and to provide a tradeoff between the projection error and the norm of the solution. Thus, Jayadeva’s projection error corresponds to the claimed error introduced when applying the reverse-projection approach to transform the generated samples into transformed samples of the original input space.)
Jayadeva teaches or at least suggests the limitation “a. defining one or more first parameters to constrain the generated sixth set of augmented samples to one or more bounds of the original dataset” (Page 3 section 2 column 1 “Let the lower and upper bounds of the input data samples A be denoted by vectors lb and ub, respectively, where lbi and ubi denote the lower and upper bounds of the ith feature or co-ordinate in the input dataset”. Jayadeva discloses defining lower and upper bound vectors lb and ub, which represent the lower and upper bounds of the respective features or coordinates of the original input dataset. Jayadeva further constrains the reverse-projected solution according to these bounds. Accordingly, lb and ub correspond to the claimed one or more first parameters, and the lower and upper limits they define correspond to the claimed bounds of the original dataset to which the generated augmented samples are constrained.)
Jayadeva teaches or at least suggests the limitation “b. defining one or more tolerance parameters of the error” (Page 3 section 2 column 2 “ε is a hyper-parameter controlling the tolerance of the approximation”. Jayadeva discloses that ε is a hyperparameter that controls the tolerance of the approximation in the reverse-projection formulation. Thus, ε corresponds to the claimed one or more tolerance parameters of the error, because it controls the amount of approximation or projection error permitted in generating the reverse-projected augmented samples.)
Jayadeva teaches or at least suggests the limitation “c. minimizing the sum-of-squares values of the variables in the input space by solving an unconstrained problem based on the error and the defined one or more first parameters and one or more tolerance parameters” (Page 5-6 section 4 equation (18)-(21) “we consider a least-squares version of the same, we can develop the solution of the optimization problem to be given by a set of linear equations, which can directly be solved (non-iteratively). This is motivated by the formulation of the least-squares SVM … We consider the least-squares version which minimizes a quadratic penalty on the slack variables … to be given by (18)–(21) … and ε is a hyper-parameter controlling the tolerance of the approximation”, and page 6-7 section 4 “We can eliminate the constraint on the slack variables to be greater than 0 as they are redundant when their quadratic penalty is being minimized by the objective function”. Jayadeva discloses a least-squares formulation in equations (18)–(21) that minimizes a quadratic objective, including the squared norm of the input-space solution variable, while incorporating the approximation-error tolerance ε and the lower and upper bound parameters lb and ub. Jayadeva further explains that certain constraints on the slack variables may be eliminated as redundant because their quadratic penalties are already minimized by the objective function. Thus, equation (18) supplies the claimed sum-of-squares minimization, equations (19)–(21) tie that optimization to the error/tolerance and bound parameters, and a person of ordinary skill in the art would have understood the disclosed elimination of redundant constraints as suggesting an unconstrained formulation in which those constraint effects are incorporated into the objective function.)
Regarding claim 9 depends on claim 1, thus the rejection of claim 1 is incorporated.
Yang teaches or at least suggests the limitation “The method according to claim 1, further comprising: providing a further indication indicating the generated sixth set of augmented samples to a second node operating in the communications system” (Paragraph 30 “Given the information about the workload distribution of the cached subset of the current minibatch, the learners establish a scheme for … exchanged among compute nodes to assemble the minibatch, e.g., such that all data of the minibatch is loaded collectively in the compute nodes for a training step”, and Paragraph 32 “The data loading scheme may also describe data to be exchanged among compute nodes if found in the aggregated cache”. Yang discloses exchanging training data among a plurality of compute nodes and establishing a data-loading scheme identifying the data to be exchanged among the compute nodes. Thus, when applied to Jayadeva’s generated augmented training samples, Yang teaches or at least suggests providing an indication identifying the generated sixth set of augmented samples to a second node operating in the communications system.)
Regarding claim 10 depends on claim 1, thus the rejection of claim 1 is incorporated.
Jayadeva teaches or at least suggests the limitation “The method according to claim 1, further comprising: determining a first machine learning model of an event in the communications system using as input the generated sixth set of augmented samples” (Page 9 section 5.1 “A generic classifier is represented by a training function TrainClassifier() which generates a classifier model that computes predictions on test data using the method TestClassifier() … we augment the entire training set with the chosen (tuned) values of the hyperparameters and train the classifier. This is used to compute the test set prediction accuracy reported in the results”. Jayadeva discloses augmenting an entire training set using the generated samples and thereafter training a classifier using the augmented training set. The TrainClassifier() function generates a classifier model from the augmented training data. Thus, Jayadeva teaches or at least suggests determining a first machine-learning model (the classifier) using the generated sixth set of augmented samples as input.)
Regarding claim 11 depends on claim 10, thus the rejection of claim 10 is incorporated.
Jayadeva in view of Yang teaches or at least suggests the limitation “The method according to claim 10, further comprising: initiating performance of an action to manage a predicted occurrence of the event according to the determined first machine learning model” (Jayadeva discloses at Page 9 section 5.1 “A generic classifier is represented by a training function TrainClassifier() which generates a classifier model that computes predictions on test data using the method TestClassifier() … we augment the entire training set with the chosen (tuned) values of the hyperparameters and train the classifier. This is used to compute the test set prediction accuracy reported in the results”, and Yang discloses at paragraph 16 “Gradient descent is a method used to a train machine learning model such as a neural network or deep learning model. Errors in model prediction on the training dataset is used to update the model (e.g., weights and bias) to reduce the errors”. Jayadeva discloses a classifier model that computes predictions on test data. Yang discloses that errors associated with model predictions are used to initiate an update of the machine-learning model, including updating its weights and biases, to reduce the prediction errors. Thus, Jayadeva in view of Yang teaches or at least suggests initiating performance of an action, namely updating the machine-learning model based on the prediction result, to manage a predicted occurrence according to the determined first machine-learning model.)
Regarding claim 12, the applicant is further directed to the rejection of claim 1 above, because the claim recites similar limitations and processing steps.
Regarding claim 13 depends on claim 12, thus the rejection of claim 12 is incorporated. Claim 13 is rejected under the same rationale as claim 10. The applicant is further directed to the rejection of claim 10 above, because the claim recites similar limitations and processing steps.
Regarding claim 14 depends on claim 13, thus the rejection of claim 13 is incorporated. Claim 14 is rejected under the same rationale as claim 11. The applicant is further directed to the rejection of claim 11 above, because the claim recites similar limitations and processing steps.
Regarding claim 15, the applicant is further directed to the rejection of claim 1 above, because the claim recites similar limitations and processing steps.
Regarding claim 16 depends on claim 15, thus the rejection of claim 15 is incorporated. Claim 16 is rejected under the same rationale as claim 2. The applicant is further directed to the rejection of claim 2 above, because the claim recites similar limitations and processing steps.
Regarding claim 26, the applicant is further directed to the rejection of claim 1 above, because the claim recites similar limitations and processing steps.
Regarding claim 27 depends on claim 26, thus the rejection of claim 26 is incorporated. Claim 27 is rejected under the same rationale as claim 10. The applicant is further directed to the rejection of claim 10above, because the claim recites similar limitations and processing steps.
Regarding claim 29, depends on claim 1, thus the rejection of claim 1 is incorporated.
Yang teaches the limitation “A computer program product comprising a non-transitory computer readable medium storing a computer program; comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out the method according to claim 1” (paragraph 59 “The present invention may be a system, a method, and/or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention”, and paragraph 60 “The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device … A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire”. Yang discloses a computer program product comprising a non-transitory computer-readable storage medium storing program instructions that, when executed by a processor, cause the processor to perform the disclosed method operations. Thus, Yang teaches the claimed computer program product configured to carry out the method of claim 1.)
Regarding claim 31, depends on claim 12, thus the rejection of claim 12 is incorporated.
Yang teaches the limitation “A computer program product comprising a non- transitory computer readable medium storing a computer program; comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out the method according to claim 12” (paragraph 59 “The present invention may be a system, a method, and/or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention”, and paragraph 60 “The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device … A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire”. Yang discloses a computer program product comprising a non-transitory computer-readable storage medium storing program instructions that, when executed by a processor, cause the processor to perform the disclosed method operations. Thus, Yang teaches the claimed computer program product configured to carry out the method of claim 12.)
Claims 5 is rejected under 35 U.S.C. 103 as being unpatentable over Jayadeva et.al (NPL: EigenSample: A non-iterative technique for adding samples to small datasets) in view of Yang et.al (US 20200302334 A1), further in view of Last et.al (NPL: Oversampling for Imbalanced Learning Based on K-Means and SMOTE), further in view of Nock et.al (US 20170337487 A1)
Regarding claim 5 depends on claim 1, thus the rejection of claim 1 is incorporated.
Jayadeva/Yang/Last does not teach “The method according to claim 1, wherein the generating is performed by using combinatorial sampling”. However, Nock teaches or at least suggests this limitation (paragraph 8 “a computer implemented method for determining multiple training samples from multiple data samples, each of the multiple data samples comprising one or more feature values and a label that classifies that data sample”, paragraph 9 “determining each of the multiple training samples by”, paragraph 10 “randomly selecting a subset of the multiple data samples, and”, paragraph 11 “combining the feature values of the data samples of the subset based on the label of each of the data samples of the subset”, and paragraph 12 “the training samples are combinations of randomly chosen data samples”. Nock discloses generating multiple training samples from a set of data samples by, for each training sample, randomly selecting a subset of the available data samples and combining the selected samples. Nock further expressly characterizes the resulting training samples as combinations of randomly chosen data samples. Thus, Nock teaches or at least suggests generating multiple different subsets/combinations from a collection of samples by selecting and obtaining different combinations of the available samples, which corresponds to performing the generation of the third set of subsets using combinatorial sampling as claimed.)
Before the effective filing date, it would have been obvious to a person of ordinary skill in the art to incorporate the teaching of the EigenSample method to augment data and reverse projecting the augmented data to the training datasets by Jayadeva, the teaching of determining epoch and dividing training datasets into minibatches by Yang, and the teaching of determining clusters based on k-means clustering and SMOTE oversampling by Last with the teaching of determining different combination of training samples from multiple data samples by Nock. The motivation to do so is referred to in Nock’s disclosure (paragraph 12 “Since the training samples are combinations of randomly chosen data samples, the training samples can be provided to third parties without disclosing the actual training data. This is an advantage over existing methods in cases where the data is confidential and should therefore not be shared with a learner of a classifier, for example.” Nock discloses generating training samples as combinations of randomly selected data samples and further teaches that such combinations may be provided to third parties without disclosing the actual underlying training data, thereby protecting confidential training information. Yang discloses distributed learners or compute nodes that generate, exchange, and process training-data subsets. Accordingly, it would have been obvious to a person of ordinary skill in the art to apply Nock’s combination-based sample-selection technique when generating Yang’s training-data subsets so that the distributed learners could exchange and process representative training samples while reducing disclosure of the underlying original data. Such a modification would have predictably provided the confidentiality benefit expressly identified by Nock.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DUY TU DIEP whose telephone number is (703)756-1738. The examiner can normally be reached M-F 8-4:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DUY T DIEP/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123