DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1,
Step 1: Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a method/process.
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
obtaining one or more similarity matrices and one or more sets of readout values of the one or more similarity matrices from one or more dataset pairs using a data-clone detection method, each set of readout values corresponding to a respective similarity matrix; (mental evaluation/mathematical concepts, a human can look at data and using pen and paper “obtain” or math out similarity matrices)
obtaining one or more importance values for the one or more similarity matrices by processing the one or more sets of readout values using an interpretation method, each importance value corresponding to a respective similarity matrix; (mental judgement, a human can, after looking at the data, determine which values are important mentally)
obtaining one or more weighted similarity matrices by weighting each similarity matrix using the corresponding importance value; (mathematical concepts/mental evaluation, a human can continue to do the steps on a pen and paper using math)
obtaining one or more summed similarity matrices by grouping and summing the weighted similarity matrices according to one or more categories for providing an analytical result with indications of locations of the data clones in the one or more dataset pairs (mathematical concepts/mental evaluation, a human can again continue to do the math on pen and paper to calculate the values)
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application?
No additional elements are recited in the claim.
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No additional elements are recited in the claim.
Note independent claims 8 and 14 recite the same substantial subject matter as independent claim 1, only differing in embodiments. The difference in embodiments do not meaningfully change the above analysis and therefore the claims are subject to the same rejection. The additional limitations of a processor and non-transitory readable media amount to generic computer components to carry out the abstract idea and thus do not make the claims eligible.
Dependent claims 2, 9, and 15 recite generating visualizations, mental observation as a human can write and come up with these such as drawing a graph.
Dependent claims 3, 10, and 16 recite generating heatmaps, mental observation as a human can write and come up with these.
Dependent claims 4 and 17 recite the heatmaps corresponding to categories, mental observation, a human can draw the heatmaps based on any criteria such as category.
Dependent claims 5, 11, and 18 recite colors for the heat map, mental observation, a human can draw the heatmap colors based on any criteria.
Dependent claims 6, 12, and 19 recite interpretation methods, tying the abstract idea to a particular field of use, MPEP 2106.05(h).
Dependent claims 7, 13, and 20 recite similarity metrics, tying the abstract idea to a particular field of use, MPEP 2106.05(h).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 6-8, 12-14, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bilke, Alexander, and Felix Naumann. "Schema matching using duplicates." in view of Teofili, Tommaso, et al. "Effective explanations for entity resolution models." [herein Teo].
Regarding claims 1, 8, and 14, Bilke teaches “a computerized method comprising: obtaining one or more similarity matrices and one or more sets of readout values of the one or more similarity matrices from one or more dataset pairs using a data-clone detection method, each set of readout values corresponding to a respective similarity matrix” (pg. 7 §5.1
PNG
media_image1.png
1022
552
media_image1.png
Greyscale
which shows the similarity generation of pairs of data);
While Bilke generally teaches importance values, Teo more specifically teaches “obtaining one or more importance values for the one or more similarity matrices by processing the one or more sets of readout values using an interpretation method, each importance value corresponding to a respective similarity matrix” (Teo pg. 3 ¶1 “Notable examples of saliency methods are LIME [26] and SHAP [18], which were conceived for generic classification tasks on textual data and images, ignoring the semantics of the problem the classifier is used to solve.” and pg. 5 ¶ above §3 “In this context, model agnostic counterfactual explanation approaches that can be adapted to the ER task include DiCE [20], LIME-C and SHAP-C [25], which we adopt as baselines.” this shows that entity resolution, i.e. deduplicating or removing data clones as is understood in the art, uses various saliency methods to determine importance);
It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Bilke with that of Teo since a combination of known methods would yield predictable results. As shown in Teo, numerous saliency methods are known and shown to show which data is important given a particular context. Therefore by combining these techniques with Bilke, one would have a more effective system of data deduplication by showing which data is more relevant.
Bilke further teaches “obtaining one or more weighted similarity matrices by weighting each similarity matrix using the corresponding importance value” (§4.2 “In our implementation we use the cosine measure, i.e., we tokenize the tuples and compare the resulting vector rep resentations. The assignment of weights for the tokens in each tuple is crucial for the effectiveness of the cosine mea sure: A weight should represent the relative importance of a token within the tuple” and §5.2 “Given the matrix M with similarity scores, the goal of this step is to derive an overall schema matching. First, we apply a user-defined threshold to M, setting all similarity values below the threshold to zero. Applying the threshold before finding the matching accommodates the attributes in either schema that do not have a matching partner. We then use the matrix as input to the bipartite weighted matching problem, also known as the assignment problem [24]. The optimal solution to this problem is a matching with the maximal sum of similarities and can be computed in polynomial time [15]” which shows weighted based on importance, i.e. based on the data from Teo); and
“obtaining one or more summed similarity matrices by grouping and summing the weighted similarity matrices according to one or more categories for providing an analytical result with indications of locations of the data clones in the one or more dataset pairs” (previous citation, “The optimal solution to this problem is a matching with the maximal sum of similarities and can be computed in polynomial time [15]”)
Independent claims 8 and 14 recite the same substantial subject matter as independent claim 1, only differing in embodiment. The differences in embodiments, a processor and computer-readable medium are obvious variations of another and would be inherent to any computing system such as the one above.
Regarding claims 6, 12, and 19, the Bilke and Teo references have been addressed above. Teo further teaches “wherein the interpretation method is a Shapley additive explanations (Shap) method, and the one or more importance values are Shap values” (Teo pg. 5 ¶ above §3 “In this context, model agnostic counterfactual explanation approaches that can be adapted to the ER task include DiCE [20], LIME-C and SHAP-C [25], which we adopt as baselines.”)
Regarding claims 7, 13, and 20, the Bilke and Teo references have been addressed above. Bilke further teaches “wherein the one or more similarity matrices comprise one or more Jaccard indices, one or more SimHashes, one or more Levenshtein distances, one or more TextRanks, and/or one or more means and corresponding deviations” (pg. 5 §4.1 “According to the classification of Cohen et al., string comparison metrics fall into three categories: edit-distance like functions, token-based similarity measures, and hybrid similarity measures [7].” edit distance is analogous to Levenshtein distance).
Claim(s) 2-5, 9-11, and 15-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bilke in view of Teo further in view of Gu, Zuguang. "Complex heatmap visualization."
Regarding claims 2, 9, and 15, the Bilke and Teo references have been addressed above. They do not explicitly teach data visualizing. Gu however teaches “further comprising: generating one or more visualizations as the analytical result using the summed similarity matrices” (Gu abstract “Heatmap is a widely used statistical visualization method on matrix‐like data to reveal similar patterns shared by subsets of rows and columns.”)
It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Bilke and Teo with that of Gu since “Heatmap is a widely used statistical visualization method on matrix‐like data to reveal similar patterns shared by subsets of rows and columns […] ComplexHeatmap can easily establish connections between multisource information by automatically concatenating and adjusting a list of heatmaps as well as complex annotations, which makes it widely applied in data analysis in many fields, especially in bioinformatics, to find hidden structures in the data.” Gu abstract. This shows that this is a useful technique for finding and showing patterns in data.
Regarding claims 3, 10, and 16, the Bilke, Teo, and Gu references have been addressed above. Gu further teaches “wherein the one or more visualizations comprise one or more heatmaps” (Gu abstract “Heatmap is a widely used statistical visualization method on matrix‐like data to reveal similar patterns shared by subsets of rows and columns.”)
Regarding claims 4 and 17, the Bilke, Teo, and Gu references have been addressed above. Gu further teaches “wherein each of the one or more heatmaps corresponds to one of the one or more categories” (Gu pg. 3 ¶2 “The heatmap is the basic unit of complex heatmap visualization. A single heatmap is composed of the heatmap body and various heatmap components (Figure 1A). The heatmap body is a two‐dimensional arrangement of grids where each grid corresponds to a specific value in the input matrix. The heatmap components contain titles, dendrograms, labels for matrix rows and columns, and heatmap annotations. These components can be optionally put on the four sides of the heatmap body and each component is managed by a specific method that is defined for the Heatmap object. Additionally, the heatmap body can be split into rows and columns, for example, by categorical variables, into slices. Dendrograms, heatmap labels, and annotations are then reordered or split accordingly”)
Regarding claims 5, 11, and 18, the Bilke, Teo, and Gu references have been addressed above. Gu further teaches “wherein, each of the one or more heatmaps comprises colors for indicating likelihoods of the data clones in the one or more dataset pairs” (pg. 4 ¶1 “In routine data analysis procedures, the matrix for heatmap visualization is normally accompanied by hierarchical clustering, so that features with similar patterns are grouped closely and they can be easily identified from the colors on heatmap” which extends to the data above).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEVIN W FIGUEROA whose telephone number is (571)272-4623. The examiner can normally be reached Monday-Friday, 10AM-6PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
KEVIN W FIGUEROA
Primary Examiner
Art Unit 2124
/Kevin W Figueroa/ Primary Examiner, Art Unit 2124