DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is made final.
Claims 1-20 are pending. Claims 1, 9 and 17 are independent claims.
Response to Arguments
The 35 U.S.C. 112(b) rejections of the previous office action have been withdrawn.
Applicant's arguments, filed 3/25/2026, regarding the 35 U.S.C. 101 rejections of the previous office action have been fully considered but are unpersuasive. The scope of the claims has changed, necessitating new grounds of rejection – see the updated 101 rejection below.
Part I
Applicant argues that claim 1 does not recite a judicial exception. Examiner disagrees, as a limitation like “computing… common features data among first data point and the second data point and their respective data distributions” is interpreted under broadest reasonable interpretation (BRI) to mean performing statistical calculations or possibly mental judgements of comparing distributions (i.e., a mental process).
Applicant argues that the claim requires a specific computing architecture that does not fall within an abstract idea grouping. Examiner argues that the recitation of various “modules” and APIs is an additional element, specifically linking the use of a judicial exception to a particular technological environment or field of use (see MPEP2106.05(h)). Describing that steps of the method of claim 1 are performed by interconnected modules in a cloud-based network environment does not amount to an inventive concept.
Applicant argues that the claim requires operations involving large-scale datasets, probabilistic sampling, and approximation of pairwise distances, but large-scale datasets are not claimed in a way that makes any processes mentally unperformable, sampling is a mathematical concept, and approximation of distances is a mental process or mathematical calculation.
Part II
The referenced limitations were added in the amendments and have been addressed in the updated rejection – see below.
Applicant argues that the specification identifies a technical problem and that the claims provide a solution: enabling similarity computation between heterogeneous descriptors, using data contraction, and generating explanatory elements for data mapping. However, examiner argues that this alleged solution is provided by broadly recited steps that encompass mental processes or mathematical calculation, and not by additional elements considered alone or in combination.
The referenced limitations were added in the amendments and have been addressed in the updated rejection – see below.
The referenced limitations were added in the amendments and have been addressed in the updated rejection – see below.
Part III
The referenced limitations were added in the amendments and have been addressed in the updated rejection – see below.
The referenced limitations were added in the amendments and have been addressed in the updated rejection – see below.
The referenced limitations were added in the amendments and have been addressed in the updated rejection – see below.
Applicant's arguments, filed 3/25/2026 regarding the 35 U.S.C. 103 rejections of the previous office action have been fully considered but are unpersuasive – however the scope of the claims has changed and new grounds of rejection have been applied. See the updated 103 rejection below.
Claim Objections
Claims 1, 9 and 17 are objected to because of the following informalities:
Claim 1 recites: “wherein each module being called via a corresponding application programming interface, and wherein the DCSCM is configured to…” in lines 7-9, but should read “wherein each module” or the like.
Claim 9 recites: “wherein each module being called via a corresponding application programming interface, and wherein the DCSCM is configured to…” in lines 9-11, but should read “wherein each module” or the like.
Claim 17 recites: “wherein each module being called via a corresponding application programming interface, and wherein the DCSCM is configured to…” in lines 8-10, but should read “wherein each module” or the like.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1:
Step 1: This part of the eligibility analysis evaluates whether the claim falls within any statutory category. See MPEP 2106.03. Claim 1 recites: A method for computing data contraction and estimating similarity of data points from heterogeneous data descriptors by utilizing one or more processors along with allocated memory, the method comprising… Claim 1 is directed to a method (Step 1: YES).
Step 2A prong 1: Does the claim recite a judicial exception? Claim 1 recites: and utilize the contraction as a way to evaluate data comparison and ranking (using a “contraction” to evaluate data comparison and ranking is a mental process), and thereby generating explanatory elements for data mapping and similarity comparison for heterogeneous data sets (generating explanatory elements for data mapping and similarity comparison is a mental process)… computing… common features data among first data point and the second data point by comparing the first data point and the second data point and their respective data distributions (computing “common features” data by comparing two points and their distributions is a mathematical calculation or mental process); linking… a pre-computed knowledge graph with the first data point and the second data point (linking a knowledge graph with data points is a mental process, i.e., identifying similar or related items); computing… in response to linking, knowledge-comparable features data among the first data point and the second data point based on other features, received as input, that are not common features (computing “knowledge-comparable features” data based on features of data is a mental process or mathematical calculation); computing…knowledge-comparable data based on the knowledge-comparable features data and the common features data (computing “knowledge-comparable” data based on other data is a mathematical calculation or mental process); computing… similarity of the first data point and the second data point based on the knowledge-comparable data (computing similarity between data points is a mathematical calculation); and generating… a data contraction map capturing an approximation achieved by approximating pairwise distances between the first data point and the second data point using an accuracy factor and an assigned similarity score (creating a “data contraction map” by approximating pairwise distances between points using an accuracy factor and a similarity score is a mental process, i.e., a process capable of being performed in the human mind with the aid of pen and paper). These steps can be performed mentally or are mathematical calculations (Step 2A prong 1: YES).
Step 2A prong 2: Does the claim recite additional elements? Do those additional elements, considered individually and in combination, integrate the judicial exception into a practical application? Claim 1 recites: implementing a data contraction and similarity computing module (DCSCM), by a data contraction and similarity computing device (DCSCD) executed by at least one processor, the DCSCD being hosted on a cloud based network environment, wherein DCSCM includes a receiving module, a generating module, a computing module, and a linking module wherein each module being called via a corresponding application programming interface, and wherein the DCSCM is configured to execute interactions among the modules to compute the data contraction… receiving, by calling the receiving module, a first input raw dataset and a second input raw dataset that are usable for computing common features data; generating, by calling the generating module, a first data point from the first input raw dataset and generating a second data point from the second input raw dataset… by calling the computing module… by calling the linking module… by calling the computing module… by calling the computing module… by calling the computing module… by calling the generating module… Specifying that the method is implemented on a device in a cloud-based network environment, and that the device comprises multiple modules that are each called by respective application programming interfaces (APIs) is an additional element linking the invention to a particular technological environment or field of use (i.e., cloud based networks, computing modules, including software modules, that operate using APIs) without significantly more (MPEP 2106.05(h)) or is mere instructions to apply the abstract idea on a generic computer (MPEP 2106.05(f)). Receiving datasets and generating points from the datasets are insignificant extra-solution activity of data gathering that does not add a meaningful limitation to the data contraction and similarity estimation method (MPEP 2106.05(g)). Specifying that various steps of the method are accomplished by calling modules is no more than mere instructions to implement the abstract idea which is equivalent to adding the words “apply it” to the recited judicial exception because the claim omits any details as to how the modules perform their individual tasks and only recite the idea of a solution or outcome (MPEP 2106.05(f)) (Step 2A prong 2: NO).
Step 2B: These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they link the invention to a particular technological environment or field of use without significantly more (MPEP 2106.05(h)), only amount to data gathering or outputting without significantly more (MPEP 2106.05(g)) or provide nothing more than mere instructions to implement an abstract idea on a generic computer (MPEP 2106.05(f)). These limitations, taken either alone or in combination, fail to provide an inventive concept (Step 2B: NO). Thus, the claim is not patent eligible.
Regarding claims 2-8, they recite limitations which further narrow the abstract idea by specifying more details of the mental and mathematical process that occurs (Claim 2, applying a data distribution sampling algorithm to datasets is a mathematical formula; Claim 3, describing sampled datasets that are smaller than the datasets that they were sampled from is still a mathematical formula; Claim 4, receiving the various input data types is insignificant extra-solution activity of data gathering without significantly more, and retrieving exact same features is a mental process; Claim 5, implementing a transforming algorithm is a mathematical formula; Claim 6, receiving a precomputed knowledge graph that has a tree-like data structure is still extra-solution activity of data gathering without significantly more, and specifying that the structure captures domain knowledge that corresponds to a line of business is specifying a field of use without significantly more; Claim 7, specifying that the line of business includes applications for loan approval is specifying a field of use without significantly more; Claim 8, applying a random sampling algorithm to the datasets is insignificant extra-solution activity of data gathering or selecting information, based on types of information and availability of information (see Electric Power Group, LLC v. Alstom S.A.), mapping data to seed points using a radius determined by an accuracy factor is a mathematical calculation or mental process, selecting seeds in a distance mapping is a mental process, and querying a distance between two points is a mathematical calculation or mental process).
Regarding claim 9, it is a system that implements a method similar to claim 1 and is rejected on the same grounds – see above.
Regarding claims 10-16, they recite similar limitations to claims 2-8 and are rejected on the same grounds – see above.
Regarding claim 17, it is an apparatus that implements a method similar to claim 1 and is rejected on the same grounds – see above.
Regarding claims 18-20, they recite similar limitations to claims 2-4 and are rejected on the same grounds – see above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 9 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abolhasssani et al. (US 20220358336 A1), herein Abolhasssani, in view of Villegas et al. (US 20230037339 A1), herein Villegas, Liu et al. (US 20210120206 A1), herein Liu, and Xian et al. (US 20210264244 A1), herein Xian.
Regarding claim 1, Abolhasssani teaches: A method for computing data contraction and estimating similarity of data points from heterogeneous data descriptors by utilizing one or more processors along with allocated memory (¶16, An AI-based data matching and alignment system that generates similarity mappings for a target data source from a plurality of data sources in a data corpus is disclosed), the method comprising: implementing a data contraction and similarity computing module (DCSCM), by a data contraction and similarity computing device (DCSCD) executed by at least one processor, the DCSCD being hosted on a cloud based network environment (¶52, FIG. 7 illustrates a computer system 700 that may be used to implement the AI-based data matching and alignment system 100 in accordance with the examples disclosed herein… a computer system 700 can sit on external-cloud platforms such as Amazon Web Services, AZURE® cloud or internal corporate cloud computing clusters, or organizational computing resources, etc.), wherein DCSCM includes a receiving module, a generating module, a computing module, and a linking module… and wherein the DCSCM is configured to execute interactions among the modules to compute the data contraction (¶53, The instructions or modules stored on the processor-readable medium 706 may include machine-readable instructions 774 executed by the processor(s) 702 that cause the processor(s) 702 to perform the methods and functions of the AI-based data matching and alignment system 100), and utilize the contraction as a way to evaluate data comparison and ranking, and thereby generating explanatory elements for data mapping and similarity comparison for heterogeneous data sets (Fig. 5, table describes matching resources along with a similarity score and column matches that were identified – i.e., explanatory elements); receiving, by calling the receiving module, a first input raw dataset and a second input raw dataset that are usable for computing common features data (¶16, the plurality of data sources from the data corpus are initially filtered to identify candidate data sources that are similar to the target data source); generating, by calling the generating module, a first data point from the first input raw dataset and generating a second data point from the second input raw dataset (¶16, The candidate data sources are further analyzed to identify columns from the candidate data sources that are similar to the columns of the target data source); computing, by calling the computing module, common features data among first data point and the second data point by comparing the first data point and the second data point and their respective data distributions (¶51, The relationships between the two tables can be derived by identifying the… distribution of the characters that make the attributes in these files similar); linking, by calling the linking module, a… knowledge graph with the first data point and the second data point (¶21, The AI-based data matching and alignment system can estimate matching data from different data sources based on the relationships determined through AI techniques described herein. Furthermore, the determined matches and relationships can be used to build the knowledge graph for the data from the plurality of data sources, which in turn can drive more efficient and accurate analytics by downstream applications); computing, by calling the computing module, in response to linking, knowledge-comparable features data among the first data point and the second data point based on other features, received as input, that are not common features (¶21, The AI-based data matching and alignment system can estimate matching data from different data sources based on the relationships determined through AI techniques described herein. Furthermore, the determined matches and relationships can be used to build the knowledge graph for the data from the plurality of data sources, which in turn can drive more efficient and accurate analytics by downstream applications); computing, by calling the computing module, knowledge-comparable data based on the knowledge-comparable features data and the common features data (¶51, The relationships between the two tables can be derived by identifying the patterns… of the characters that make the attributes in these files similar, the semantical and statistical features that are in common among them, and the dependencies between the files that can filter pairwise attribute comparisons); computing, by calling the computing module, similarity of the first data point and the second data point based on the knowledge-comparable data (¶39, In an example, the similarity calculation can use the feature information about the columns in the target data source 190 and the candidate data source(s)); and generating, by calling the generating module, a data contraction… capturing an approximation achieved by approximating pairwise distances between the first data point and the second data point (¶42, A distance measure is calculated at 252 between the feature matrices of each of the data sources 192, 194, . . . 198, and the target data source 190. In an example, the Mahanalobis distance technique can be used to measure the distance between the feature matrices of the data sources) using an accuracy factor (¶46, For example, the ranked list 360 shows that the column ‘power’ is similar to the columns ‘PW’ with 71% confidence and to the column ‘XYZ’ with 66% confidence. Therefore, the node representing the column ‘power’ is connected to the nodes representing the columns ‘XYZ’ and ‘PW’ the tree 350. Similarly, the column ‘Sensor 2’ is similar to the columns ‘Sensor 8’ and ‘SX23’ with 86% and 65% confidence values respectively – where confidence values are interpreted to be accuracy factors).
Abolhasssani fails to explicitly teach: wherein each module being called via a corresponding application programming interface…
However, in the same field of endeavor, Villegas teaches: wherein each module being called via a corresponding application programming interface (¶37, system 108 includes one or more devices that provide and execute one or more engines, modules, applications, or other logical components… implemented using, for example, one or more of processing devices, servers, platforms, virtual systems, cloud infrastructure, or other suitable components (or combinations of components). In addition, each engine (or other logical component) is also implemented using one or more servers, one or more platforms with corresponding application programming interfaces, cloud infrastructure, and the like)…
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the program modules using APIs in a cloud computing environment as disclosed by Villegas in the method disclosed by Abolhasssani to promote scalability and configurability (¶38, In some examples, configurable computing resources from the shared pool are rapidly provisioned, such as via virtualization, and are released with low management effort or service provider interaction. In some cases, cloud computing can provide configurable computing resources that are scalable, such as automatic scaling).
Abolhasssani in view of Villegas fails to teach: a pre-computed knowledge graph.
However, in the same field of endeavor, Liu teaches: a pre-computed knowledge graph (¶60, the assistant system 140 may pre-compute a plurality of personalized language models for a plurality of possible subjects a user may talk about. When a user requests assistance, the assistant system 140 may then swap these pre-computed language models quickly).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use a precomputed model as disclosed by Liu in the method disclosed by Abolhasssani in view of Villegas to allow for rapid and efficient configuration (¶60, As a result, the assistant system 140 may have a technical advantage of saving computational resources while efficiently determining what the user may be talking about).
Abolhasssani in view of Villegas and Liu fails to teach: a data contraction map… and an assigned similarity score…
However, in the same field of endeavor, Xian teaches: a data contraction map… and an assigned similarity score (Fig. 7A, data contraction map that approximates distances between data points – and – ¶87, the explanatory annotation system 106 can utilize a variety of approaches to determine a distance value. For example, the explanatory annotation system 106 can utilize approaches such as, but not limited to, a Euclidean distance and cosine similarities)…
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use a data contraction map along with similarity score as disclosed by Xian in the method disclosed by Abolhasssani in view of Villegas and Liu to improve efficiency and transparency in a data matching process (¶20, the explanatory annotation system can accurately, flexibly, and efficiently generate an explanatory path that provides transparency into the label determination process while also using deep learning based models for accurate column annotation).
Regarding claim 9, it is a system that implements a method similar to claim 1 and is rejected on the same grounds – see above.
Regarding claim 17, it is an apparatus that implements a method similar to claim 1 and is rejected on the same grounds – see above.
Claim(s) 2, 3, 10, 11, 18 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abolhasssani in view of Villegas, Liu and Xian as applied to claims 1, 9 and 17 above, and further in view of Shimazu (US 20210056444 A1).
Regarding claim 2, Abolhasssani in view of Villegas, Liu and Xian fails to teach: The method according to claim 1, further comprising: applying a data distribution sampling algorithm onto each of said first input raw dataset and said second input raw dataset to generate a first sampled dataset and a second sampled dataset, respectively.
However, in the same field of endeavor, Shimazu teaches: applying a data distribution sampling algorithm onto each of said first input raw dataset and said second input raw dataset to generate a first sampled dataset and a second sampled dataset, respectively (¶263, After random sampling (i.e. bootstrap), a reduced training set of data can be generated from each of the randomly sampled raw training sets of data (using e.g. the various methods described above for computing a reduced training set based on a raw training set of data).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use a data sampling algorithm to create sampled datasets as disclosed by Shimazu in the method disclosed by Abolhasssani in view of Villegas, Liu and Xian to reduce computation (¶38, thereby reducing computation complexity when processing the reduced training set of data by the classification algorithm, compared to processing the training set of data).
Regarding claim 3,Abolhasssani in view of Villegas, Liu and Xian fails to teach: The method of claim 2, wherein a size of the first sampled dataset is smaller than the first received input raw dataset, and wherein a size of the second sampled dataset is smaller than the second received input raw dataset.
However, in the same field of endeavor, Shimazu teaches: wherein a size of the first sampled dataset is smaller than the first received input raw dataset, and wherein a size of the second sampled dataset is smaller than the second received input raw dataset (¶263, After random sampling (i.e. bootstrap), a reduced training set of data can be generated from each of the randomly sampled raw training sets of data (using e.g. the various methods described above for computing a reduced training set based on a raw training set of data).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to create reduced sampled datasets as disclosed by Shimazu in the method disclosed by Abolhasssani in view of Villegas, Liu and Xian to reduce computation (¶38, thereby reducing computation complexity when processing the reduced training set of data by the classification algorithm, compared to processing the training set of data).
Regarding claims 10 and 11, they recite limitations similar to claims 2 and 3 respectively and are rejected on the same grounds – see above.
Regarding claims 18 and 19, they recite limitations similar to claims 2 and 3 respectively and are rejected on the same grounds – see above.
Claim(s) 4, 5, 12, 13 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abolhasssani in view of Villegas, Liu, Xian and Shimazu as applied to claims 3, 11 and 19 above, and further in view of Kiljanek (US 20200118691 A1).
Regarding claim 4, Abolhasssani further teaches: the method according to claim 3, wherein in computing the common features, the method further comprising: receiving as input data the following data: the first data point generated from the first input raw dataset, the second data point generated from the second input raw dataset, the first sampled data set, and the second sampled dataset; retrieving exact same features among the first data point and the second data point (¶33, Lastly, if the matching columns are of numeric type, the explanations can be generated by showing that the distance between the distributions of the two columns is minimum as compared to other non-matching, numeric columns – and – ¶34, Distribution distance between NPD_FACILITY_CODE and NPD_FACILITY_CODE_2 is 0.0 – i.e., exact same features between data points, in the case of Abolhasssani columns are data points extracted from datasets).
Abolhasssani in view of Villegas, Liu, Xian and Shimazu fails to teach: and retrieving exact same features among the first sampled data set and the second sampled dataset.
However, in the same field of endeavor, Kiljanek teaches: and retrieving exact same features among the first sampled data set and the second sampled dataset (¶60, For example, the “age” feature may have “ages” or “years” or “how old” or “how old are you?” as possible aliases. If aliases of features across different simulated patient population datasets match, these features may be normalized by renaming one or both feature names so that the features appear consistently named).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to identify matching features across datasets as disclosed by Kiljanek in the method disclosed by Abolhasssani in view of Villegas, Liu, Xian and Shimazu to enable efficient comparisons between different datasets (¶60, allowing simulated patient datasets that were originally from different simulated patient population datasets to be easily compared).
Regarding claim 5, Abolhasssani further teaches: The method according to claim 4, wherein in computing knowledge-comparable data, the method further comprising: implementing a corresponding transforming algorithm to transform corresponding received input data with respect to the common features and knowledge-comparable features sets (¶33, Lastly, if the matching columns are of numeric type, the explanations can be generated by showing that the distance between the distributions of the two columns is minimum as compared to other non-matching, numeric columns. The Kolmogorov-Smimov test may be used for determining the distance between the columns. For example, NPD_FACILITY and NPD_FACILITY_CODE_2 may be matched from Tables 1 and 2 respectively – and – ¶34, Distribution distance between NPD_FACILITY_CODE and NPD_FACILITY_CODE_2 is 0.0 – performing statistical tests on data can be interpreted as implementing a transforming algorithm).
Regarding claims 12 and 13, they recite limitations similar to claims 4 and 5 respectively and are rejected on the same grounds – see above.
Regarding claim 20, it recites limitations similar to claim 4 and is rejected on the same grounds – see above.
Claim(s) 6, 7, 14 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abolhasssani in view of Villegas, Liu and Xian as applied to claims 1 and 9 above, and further in view of Budzik (US 20210158085 A1).
Regarding claim 6, Abolhasssani further teaches: The method according to claim 1, wherein the precomputed knowledge graph is a tree-like data structure (fig. 3 demonstrates a tree-like data structure that has nodes and connections between nodes like a tree).
Abolhasssani in view of Villegas, Liu and Xian fails to teach: that captures domain knowledge corresponding to a line of business.
However, in the same field of endeavor, Budzik teaches: that captures domain knowledge corresponding to a line of business (¶15, the model purpose data is generated by domain experts (e.g., data scientists, business analysts, and the like) having specific domain knowledge related to the identified purpose).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include domain knowledge associated with a line of business as disclosed by Budzik in the method disclosed by Abolhasssani in view of Villegas, Liu and Xian to create models that can function without continued expert input (¶15, can be used to automatically generate models for “auto loan origination” purposes without further input from a data scientist).
Regarding claim 7, Abolhasssani in view of Villegas, Liu and Xian fails to teach: The method according to claim 5, wherein the line of business includes applications for loan approval.
However, in the same field of endeavor, Budzik teaches: wherein the line of business includes applications for loan approval (¶16, In some variations, the model purpose relates to consumer loan origination, and results of the model are used to determine whether to grant a consumer loan. In some variations, the model purpose relates to business loan origination, and results of the model are used to determine whether to grant a loan to a business. In other variations, the model purpose relates to loan repayment prediction, and results of the model are used to determine whether a loan already granted will be repaid).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include domain knowledge associated with applications for loan approval as disclosed by Budzik in the method disclosed by Abolhasssani in view of Villegas, Liu, Xian, Shimazu and Kiljanek to create a loan processing model that can function without continued expert input (¶15, can be used to automatically generate models for “auto loan origination” purposes without further input from a data scientist).
Regarding claim 14, it recites similar limitations to claim 6 and is rejected on the same grounds – see above.
Regarding claim 15, it recites similar limitations to claim 7 and is rejected on the same grounds – see above.
Claim(s) 8 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abolhasssani in view of Villegas, Liu and Xian as applied to claims 1 and 9 above, and further in view of Hamerly et al. (“Alternatives to the k-means algorithm that find better clusterings”, 2002), herein Hamerly, Patthak et al. (US 20160292592 A1), herein Patthak, Pitalúa García et al. (US 20210021414 A1), herein Pitalúa García, and Hegelich et al. (US 20220358282 A1), herein Hegelich.
Abolhasssani in view of Villegas, Liu and Xian fails to teach: The method according to claim 1, wherein in computing the similarity of the first data point and the second data point, the method further comprising: applying an independent and identically distributed sampling algorithm to each of said first input raw dataset and said second input raw dataset to construct seed points…
However, in the same field of endeavor, Hamerly teaches: wherein in computing the similarity of the first data point and the second data point, the method further comprising: applying an independent and identically distributed sampling algorithm to each of said first input raw dataset and said second input raw dataset to construct seed points (pg. 603, Section 5, ¶4, The Forgy method chooses k data points from the dataset at random and uses them as the initial centers. The Random Partition method assigns each data point to a random center, then computes the initial location of each center as the centroid of its assigned points – also see Fig. 3, which depicts results of both methods)…
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use an independent and identically distributed sampling algorithm (i.e., random sample) as disclosed by Hamerly in the method disclosed by Abolhasssani in view of Villegas, Liu and Xian to effectively initialize clusters (pg. 605, Section 7, ¶2, Previous work in initialization methods has concluded that the Random Partition method is good for GEM and for KM, but our experiments do not confirm this conclusion. The Forgy method of initialization (choosing random points as initial centers) works best for GEM, KM, and H1. Overall, our results suggest that the best algorithms available today are FKM, H2, and KHM, initialized by the Random Partition method).
Abolhasssani in view of Villegas, Liu, Xian and Hamerly fails to teach: implementing an intra-dataset mapping algorithm that maps an expanded set to a seed point among the constructed seed points using a radius of the accuracy factor, wherein the accuracy factor is a parameter controlling an approximation error of a distance mapping value…
However, in the same field of endeavor, Patthak teaches: implementing an intra-dataset mapping algorithm that maps an expanded set to a seed point among the constructed seed points using a radius of the accuracy factor, wherein the accuracy factor is a parameter controlling an approximation error of a distance mapping value (¶219, At this point, at 2204, a determination is made of the coverage of the sample logs that fall within the similarity radius of the centroid(s) as well as the extent of the similarity radius. In particular, if there are any sample logs at all that do not fall within the scope of one of the clusters, then the numbers of clusters must be adjusted to create a new cluster and/or the similarity radius needs to be adjusted to account for the coverage error – the radius that determines clusters influences the error, so the radius is interpreted as being an accuracy factor – also see Fig. 23)…
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to map data to centroids (i.e., seed points) using a radius of an accuracy factor that controls an approximation error as disclosed by Patthak in the method disclosed by Abolhasssani in view of Villegas, Liu, Xian and Hamerly to improve the mapping accuracy (¶222, FIG. 23 illustrates the situation where the original similarity radius 2312a was inadequate to correspond to the vectors to be clustered, e.g., where the un-clustered vectors are known to be for the exact same log type as the vectors that are actually in the cluster and hence failure to include the un-clustered vectors constitutes an error in coverage for the cluster. In some situations to correct this problem, the radius can be expanded into a modified similarity radius 2312b. This modified radius 2312b now correctly clusters all vectors for the log type without any classification errors).
Abolhasssani in view of Villegas, Liu, Xian, Hamerly and Patthak fails to teach: in an interval (0,1)…
However, in the same field of endeavor, Pitalúa García teaches: in an interval (0,1) (¶96, small allowed error rate γ∈(0,1)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use an accuracy factor in the interval (0,1) as disclosed by Pitalúa García in the method disclosed by Abolhasssani in view of Villegas, Liu, Xian, Hamerly and Patthak to represent all possible values of a rate, which can theoretically be 0%-100% or (0,1) (¶96, error rate γ∈(0,1)).
Abolhasssani in view of Villegas, Liu, Xian, Hamerly, Patthak and Pitalúa García fails to teach: implementing an inter-dataset mapping algorithm that selects, for every pair of input raw datasets, a unique pair of seeds in the distance mapping; and querying, in response to selecting, a distance between the first data point and the second data point.
However, in the same field of endeavor, Hegelich teaches: implementing an inter-dataset mapping algorithm that selects, for every pair of input raw datasets, a unique pair of seeds in the distance mapping; and querying, in response to selecting, a distance between the first data point and the second data point (¶45, In a preferred embodiment an antonym of a negatable word is determined by comparing a context of a negatable word to one or more known contexts of synsets of the negatable word stored in a synset database to determine a similarity measure, selecting a synset based on the determined similarity to the context of the negatable word – ¶48, In a preferred embodiment the similarity measure is a distance between a centroid vector of the context of the negatable word for which an antonym shall be determined in a word vector space and respective centroid vectors of the contexts of the synsets for the negatable word in the same word vector space).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to select paired seeds, or centroids, to compare a pair of datasets as disclosed by Hegelich in the method disclosed by Abolhasssani in view of Villegas, Liu, Xian, Hamerly, Patthak and Pitalúa García to improve efficiency (¶50, Based on the centroid vectors a distance between the centroid vector of the context of the negatable element and the centroid vectors of the contexts of the synsets is determined as a similarity measure where a smaller distance indicates a closer proximity or higher similarity of the respective synsets. Using a distance between the centroid vectors is a computationally efficient method of determining the similarity of two or more contexts).
Regarding claim 16, it recites similar limitations to claim 8 and is rejected on the same grounds – see above.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HARRISON CHAN YOUNG KIM whose telephone number is (571)272-0713. The examiner can normally be reached Monday - Thursday 10:00 am - 7:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HARRISON C KIM/ Examiner, Art Unit 2145
/CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145