Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This action is responsive to the following communication: Non-Provisional Application filed Jun. 3, 2024.
Claims 1-20 are pending in the case. Claims 1, 9 and 17 are independent claims.
Claim Rejections - 35 U.S.C. § 101
35 U.S.C. § 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more.
As to claim 1:
Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03.
Yes, the claim is to a process.
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
Yes, the limitation “receiving a data sample comprising one or more features and a target property” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Yes, the limitation “identifying … an unsupervised machine learning model trained to classify data samples based on the one or more features, independently of the target property, into a plurality of clusters” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Yes, the limitation “classifying the data sample based on the one or more features using the unsupervised machine learning model to compute a cluster” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Yes, the limitation “identifying, by the processor, a supervised machine learning model corresponding to the cluster; and computing a value for the target property by supplying the data sample to the supervised machine learning model” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
No, the limitation “a processor of a computer system” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §§ 2106.04(d), 2106.05(h).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
No, the limitation “a processor of a computer system” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
The additional elements, taken alone or in combination, fail to amount to significantly more than the judicial exception.
As to claim 9:
Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03.
Yes, the claim is to a machine.
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
Yes, the limitation “receiving a data sample comprising one or more features and a target property” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Yes, the limitation “identifying … an unsupervised machine learning model trained to classify data samples based on the one or more features, independently of the target property, into a plurality of clusters” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Yes, the limitation “classifying the data sample based on the one or more features using the unsupervised machine learning model to compute a cluster” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Yes, the limitation “identifying, by the processor, a supervised machine learning model corresponding to the cluster; and computing a value for the target property by supplying the data sample to the supervised machine learning model” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
No, the limitation “a memory storing instructions that, when executed by the processor, cause the processor …” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §§ 2106.04(d), 2106.05(h).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
No, the limitation “a memory storing instructions that, when executed by the processor, cause the processor …” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
The additional elements, taken alone or in combination, fail to amount to significantly more than the judicial exception.
As to claim 17:
Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03.
Yes, the claim is to a machine.
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
Yes, the limitation “ receive a data sample comprising one or more features and a target property” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Yes, the limitation “identify a first machine learning model trained to classify data samples based on the one or more features, independently of the target property, into a plurality of clusters; classify the data sample based on the one or more features using the first machine learning model to compute a cluster; identify a second machine learning model corresponding to the cluster” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Yes, the limitation “compute an intermediate value by supplying the data sample to the second machine learning model; identify a third machine learning model corresponding to the cluster; and compute a value for the target property by supplying the data sample and the intermediate value to the third machine learning model.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
No, the limitation “non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor…” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §§ 2106.04(d), 2106.05(h).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
No, the limitation “non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor…” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
The additional elements, taken alone or in combination, fail to amount to significantly more than the judicial exception.
Dependent Claims 2-8, 10-16 and 18-20 recite the same abstract idea of reading data, performing some analysis on the data, and presenting the data. They do not otherwise add any meaningful limits beyond the abstract idea.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Peter (hereinafter Peter) U.S. Patent Publication No. 2006/0047655 in view of Shazeer et al. (hereinafter Shazeer) U.S. Patent Publication No. 2023/0419079.
With respect to independent claim 1, Peter teaches a method for generating a chain of machine learning models comprising:
receiving a data sample comprising one or more features and a target property (see e.g., Abstract Para [5]-[8][12] [62]-” we used a geospatial dataset made up of 10.sup.5 points with coordinates in latitude and longitude format. The data points are shown in FIG. 4 For the ACE runs, a coarse-zoning mesh with 21 cells (22 grid points) in both the x- and y-direction was initially used, as shown in the figure. This corresponds to a total number of grid points N.sub.g=484, so N>>N.sub.g. A close look at the data in FIG. 4 suggests that the number of clusters is .about.70.”);
identifying, by a processor of a computer system, an unsupervised machine learning model trained to classify data samples based on the one or more features, independently of the target property, into a plurality of clusters (see e.g., Para [12][22]-[35][69] - “an algorithm provides a new way of clustering data in an unsupervised manner. This algorithm is fast, efficient, and robust, and is ideal for large datasets. It consists of the following steps: (1) A number N of data instances with a number n fields is linearly weighted to an n-dimensional mesh with (for example) m grid points per dimension. (2) A number of "intelligent agents" is placed randomly on the mesh. These agents move along the grid according to special rules that cause them to find grid points that have the largest weight. All clusters can be determined in this fashion and the clusters can be ranked in "strength". (3) These maxima are then used as the "centroid" of each cluster. If desired, the mesh can be gridded finer around these "centroids" to obtain finer scaling. (4) All data points within a certain specified distance of these centroids are considered to form a cluster.”” ACE is ideal for running in a massively-parallel mode, or by distributed computation. Load balancing is achieved by dividing the spatial mesh into sectors, so that each processor only acts on a certain well-defined region of space. For efficiency, each sector might contain a roughly equal number of grids on the physical mesh (unless there is an anisotropy in the data which would preferentially require more processing power in certain spatial domains). Cluster which span sectors would be handled transparently by interprocessor communication between nodes.”);
classifying the data sample based on the one or more features using the unsupervised machine learning model to compute a cluster (see e.g., Para [36]-[43] and Claim 4-“the dataset including N points and at n fields, comprising: (a) forming a n-dimensional grid; (b) "weighting" each of the "N" data instances to the grid; (c) determining at least one cluster within the data points bases on the weighting of points on the grid.”).
Peter does not expressly show identifying a supervised machine learning model corresponding to the cluster and computing a value for the target property by supplying the data sample to the supervised machine learning model. However, Shazeer teaches the above feature (see e.g. para[4][22]-[32] – “The MoE subnetwork further includes a gating subsystem configured to: select, based on the first layer output, one or more of the expert neural networks and determine a respective weight for each selected expert neural network, provide the first layer output as input to each of the selected expert neural networks, combine the expert outputs generated by the selected expert neural networks in accordance with the weights for the selected expert neural networks to generate an MoE output, and provide the MoE output as input to the second neural network layer.””the gating subsystem 110 combines the expert outputs generated by the selected expert neural networks by weighting the expert output generated by each of the selected expert neural networks (e.g., expert output 126 generated by the expert neural network 116 and expert output 128 generated by the expert neural network 120) by the weight for the selected expert neural network to generate a weighted expert output, and summing the weighted expert outputs (e.g., the weighted expert outputs 134 and 136) to generate the MoE output 132.”). Both Shazeer and Peter are directed to partitioning input datasets into subsets for data processing. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Shazeer and Peter in front of them to modify the system of Peter to include the above feature. The motivation to combine Shazeer and Peter comes from Shazeer. Shazeer discloses the motivation to improve prediction accuracy by allowing local and specialize modeling (see e.g. para [4]-[8]). This motivation for combination also applies to the remaining claims which depend on this combination.
With respect to dependent claim 2, the modified Peter teaches the one or more features are selected from a plurality of fields of a data model (see e.g. Abstract and para [5]-[8] – “A method for clustering large datasets in which a number N of data instances with a number n fields is linearly weighted to an n-dimensional mesh with (for example) m grid points per dimension, a number of "intelligent agents" is placed randomly on the mesh. “), and wherein the unsupervised machine learning model is identified from among a plurality of unsupervised machine learning models trained to classify data samples based on different combinations of fields of the data model (Peter does not expressly show identifying one model from a plurality of unsupervised machine learning models. However, based on the teachings of Shazeer (e.g., Para [27]), it would have been obvious because Shazeer teaches selecting expert subnetworks for data processing. See discussion above.).
With respect to dependent claim 3, the modified Peter teaches the supervised machine learning model comprises one or more selected from the group comprising: a linear regression model; a non-linear regression model; or a classification model (see e.g. Shazeer para [27]-[35]).
With respect to dependent claim 4, the modified Peter teaches the unsupervised machine learning model comprises a clustering model (see e.g. para [3]-[8] – “The present invention relates to unsupervised clustering of large datasets. More particularly, the present invention relates to processes for unsupervised clustering of large datasets having various types of data, including geospatial data.”).
With respect to dependent claim 5, the modified Peter teaches the unsupervised machine learning model comprises an anomaly detection model (see e.g. para [75]-[77] – “low-density noisy data can be identified (and ignored)”).
With respect to dependent claim 6, the modified Peter teaches supplying the value for the target property computed based on the supervised machine learning model to the anomaly detection model to compute a likelihood that the value for the target property is anomalous (see e.g. para [75]-[77]).
With respect to dependent claim 7, the modified Peter teaches receiving a second data sample comprising one or more features and a second target property, the second target property being different from the target property of the data sample (see e.g. para [12]-[16] – it is clear that the method can be applied with different target property); classifying the second data sample based on the one or more features of the second data sample using the unsupervised machine learning model to compute a second cluster (see e.g., Para [36]-[43]); identifying, by the processor, a second supervised machine learning model corresponding to the cluster (see e.g. Shazeer para[4][22]-[32]); and computing a second value for the second target property by supplying the second data sample to the second supervised machine learning model (see e.g. Shazeer para[4][22]-[32]).
With respect to dependent claim 8, the modified Peter teaches the supervised machine learning model computes the cluster based on a first subset of the one or more features of the data sample (see e.g. Shazeer para [4]-[8] – “select, based on the first layer output, one or more of the expert neural networks and determine a respective weight for each selected expert neural network, provide the first layer output as input to each of the selected expert neural networks, combine the expert outputs generated by the selected expert neural networks in accordance with the weights for the selected expert neural networks to generate an MoE output, and provide the MoE output as input to the second neural network layer.”), and wherein the supervised machine learning model computes the value for the target property based on a second subset of the one or more features of the data sample, the second subset being different from the first subset (see e.g. Shazeer para [4]-[8] – it is obvious that the second subset can be different from the first subset because they are not required to be same).
Claim 9 is rejected for the similar reasons discussed above with respect to claim 1.
Claim 10 is rejected for the similar reasons discussed above with respect to claim 2.
Claim 11 is rejected for the similar reasons discussed above with respect to claim 3.
Claim 12 is rejected for the similar reasons discussed above with respect to claim 4.
Claim 13 is rejected for the similar reasons discussed above with respect to claim 5.
Claim 14 is rejected for the similar reasons discussed above with respect to claim 6.
Claim 15 is rejected for the similar reasons discussed above with respect to claim 7.
Claim 16 is rejected for the similar reasons discussed above with respect to claim 8.
With respect to independent claim 17, the modified Peter teaches a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to: receive a data sample comprising one or more features and a target property (see e.g., Abstract Para [5]-[8][12] [62]-” we used a geospatial dataset made up of 10.sup.5 points with coordinates in latitude and longitude format. The data points are shown in FIG. 4 For the ACE runs, a coarse-zoning mesh with 21 cells (22 grid points) in both the x- and y-direction was initially used, as shown in the figure. This corresponds to a total number of grid points N.sub.g=484, so N>>N.sub.g. A close look at the data in FIG. 4 suggests that the number of clusters is .about.70.”); identify a first machine learning model trained to classify data samples based on the one or more features, independently of the target property, into a plurality of clusters (see e.g., Para [12][22]-[35][69] - “an algorithm provides a new way of clustering data in an unsupervised manner. This algorithm is fast, efficient, and robust, and is ideal for large datasets. It consists of the following steps: (1) A number N of data instances with a number n fields is linearly weighted to an n-dimensional mesh with (for example) m grid points per dimension. (2) A number of "intelligent agents" is placed randomly on the mesh. These agents move along the grid according to special rules that cause them to find grid points that have the largest weight. All clusters can be determined in this fashion and the clusters can be ranked in "strength". (3) These maxima are then used as the "centroid" of each cluster. If desired, the mesh can be gridded finer around these "centroids" to obtain finer scaling. (4) All data points within a certain specified distance of these centroids are considered to form a cluster.”” ACE is ideal for running in a massively-parallel mode, or by distributed computation. Load balancing is achieved by dividing the spatial mesh into sectors, so that each processor only acts on a certain well-defined region of space. For efficiency, each sector might contain a roughly equal number of grids on the physical mesh (unless there is an anisotropy in the data which would preferentially require more processing power in certain spatial domains). Cluster which span sectors would be handled transparently by interprocessor communication between nodes.”); classify the data sample based on the one or more features using the first machine learning model to compute a cluster (see e.g., Para [12][22]-[35][69]); identify a second machine learning model corresponding to the cluster (see e.g. Shazeer para[4][22]-[32] – “The MoE subnetwork further includes a gating subsystem configured to: select, based on the first layer output, one or more of the expert neural networks and determine a respective weight for each selected expert neural network, provide the first layer output as input to each of the selected expert neural networks, combine the expert outputs generated by the selected expert neural networks in accordance with the weights for the selected expert neural networks to generate an MoE output, and provide the MoE output as input to the second neural network layer.””the gating subsystem 110 combines the expert outputs generated by the selected expert neural networks by weighting the expert output generated by each of the selected expert neural networks (e.g., expert output 126 generated by the expert neural network 116 and expert output 128 generated by the expert neural network 120) by the weight for the selected expert neural network to generate a weighted expert output, and summing the weighted expert outputs (e.g., the weighted expert outputs 134 and 136) to generate the MoE output 132.”); compute an intermediate value by supplying the data sample to the second machine learning model; identify a third machine learning model corresponding to the cluster; and compute a value for the target property by supplying the data sample and the intermediate value to the third machine learning model (see e.g. para[4][22]-[32]).
With respect to dependent claim 18, the modified Peter teaches a value corresponding to the intermediate value is absent from the data sample (see e.g. Shazeer Para [5] – the weight value is generated and not from the data sample).
With respect to dependent claim 19, the modified Peter teaches receive a second data sample comprising one or more features different from the one or more features of the data sample and a second target property (see e.g. para [12]-[16] – it is clear that the method can be applied with different target property), the second target property being different from the target property of the data sample (see e.g. Shazeer para [4]-[8] ][22]-[32]); classify the second data sample based on the one or more features of the second data sample using a fourth machine learning model different from the first machine learning model to compute a second cluster (see e.g. Shazeer para [4]-[8] ][22]-[32]); identify a fifth machine learning model corresponding to the cluster; and compute a value for the second target property by supplying the second data sample to the fifth machine learning model (see e.g. Shazeer para [4]-[8]).
With respect to dependent claim 20, the modified Peter teaches receive a third data sample comprising one or more features (see e.g., Para [12][22]-[35][69]); identify the first machine learning model trained to classify data samples based on the one or more features into the plurality of clusters (see e.g., Para [12][22]-[35][69]); and detect an anomaly in the data sample based on the one or more features using the first machine learning model (see e.g. para [75]-[77] – “low-density noisy data can be identified (and ignored)”).
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain.” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)). Further, a reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill the art, including nonpreferred embodiments. Merck & Co. v. Biocraft Laboratories, 874 F.2d 804, 10 USPQ2d 1843 (Fed. Cir.), cert. denied, 493 U.S. 975 (1989). See also Upsher-Smith Labs. v. Pamlab, LLC, 412 F.3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir. 2005); Celeritas Technologies Ltd. v. Rockwell International Corp., 150 F.3d 1354, 1361, 47 USPQ2d 1516, 1522-23 (Fed. Cir. 1998).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEIYONG WENG whose telephone number is (571)270-1660. The examiner can normally be reached on Mon.-Fri. 8 am to 5 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Matthew Ell, can be reached on (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://portal.uspto.gov/external/portal. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/PEI YONG WENG/ Primary Examiner, Art Unit 2141