DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 3, 7, 10, 14, 17 are objected to because of the following informalities:
"wherein applying", “wherein generating” should be "wherein the applying", “wherein the generating” [Claims 3, 7, lines 1, 1];
"to apply" should be "to said apply" [Claims 10, 17, lines 3, 2];
"to generate" should be "to said generate" [Claim 14, line 3].
Appropriate correction is required. Further, in an effort to practice compact prosecution, each of these limitations has been interpreted similarly as in the provided recommendation for each limitation, above.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
Claims 7, 14 rejected under 35 U.S.C. 112(a) as failing to comply with the written description requirement. The claims contain subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, at the time the application was filed, had possession of the claimed invention. That is, the applicant’s written description is silent with regard to combining any sort of match percentages for generating a confidence score.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP §§ 706.02(l)(1) - 706.02(l)(3) for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
Claims 1-2, 4-6, 8-9, 11-13, 15-16, 18-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-3, 6-7, 9-11, 14-15, 17-19, 21 of U.S. Patent No. 12,443,581. Although the claims at issue are not identical, they are not patentably distinct from each other because one of ordinary skill in the art would recognize that 1-3, 6-7, 9-11, 14-15, 17-19, 21 of U.S. Patent No. 12,443,581 are directed to a similar invention because they anticipate the claims in the present application.
Claim
US 19/337,154
Claim
US 12,443,581
1, 8, 15
retrieving, by a computing device, database metadata from at least two different database types without collecting row data stored in data tables of the at least two databases, wherein the database metadata comprises column names, column attribute datatypes, and column descriptions; converting, by the computing device, the database metadata from the at least two different database types to a common format; storing, by the computing device, a database metadata table comprising the converted database metadata, wherein the database metadata table comprises a first column of column names and at least one additional column comprising column attribute datatypes or column descriptions;
applying, by the computing device, one or more database metadata rules to the database metadata table to identify at least one portion of the database metadata table associated with personal information (PI);
and generating, by the computing device, a confidence score associated with the at least one portion of the database metadata table, wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata table comprise the PI
1, 9, 17
retrieving, by a computing device, first database metadata for a first plurality of data tables stored within a first database without collecting row data stored in the first plurality of data tables; retrieving, by the computing device, second database metadata for a second plurality of data tables stored within a second database without collecting row data stored in the second plurality of data tables, wherein the first database is compatible with a first database type and the second database is compatible with a second database type that is different than the first database type, wherein the first database metadata and the second database metadata each comprise two or more of: a plurality of column names, a plurality of column attribute datatypes, and a plurality of column descriptions stored within the first plurality of data tables and the second plurality of data tables, respectively;
converting, by the computing device, the first database metadata from the first database type to a common format and the second database metadata from the second database type to the common format;
based on the converting of the first database metadata from the first database type to the common format and the converting of the second database metadata to the common format, storing a database metadata table comprising both of the first database metadata and the second database metadata, wherein the database metadata table comprises two or more of: a first column comprising the plurality of column names, a second column comprising the plurality of column attribute datatypes, and a third column comprising the plurality of column descriptions stored within the first plurality of data tables and/or the second plurality of data tables;
determining, based on a request for data associated with one or more personal information (PI) elements associated with a requesting party, one or more database metadata rules, wherein the one or more database metadata rules are configured to identify one or more column names, one or more column attribute datatypes, or one or more column descriptions, within the database metadata table, associated with one or more data tables of the first plurality of data tables and/or the second plurality of data tables that comprise the data associated with the one or more PI elements;
determining, based on the one or more database metadata rules, at least one portion of the database metadata table indicative of one or more databases, of the first database and/or the second database, comprising the one or more data tables that comprise the one or more PI elements, wherein the data associated with the one or more PI elements is stored in the one or more data tables as the row data;
determining a confidence score associated with the at least one portion of the database metadata table, wherein the confidence score is indicative of a level of confidence that the one or more data tables indicated by the at least one portion of the database metadata table comprise the one or more PI elements; and
sending, based on the confidence score, a response to the request
2, 9 ,16
wherein the one or more database metadata rules comprise one or more regular expressions configured to locate character patterns indicative of the personal information (PI)
6, 14, 21
wherein the one or more database metadata rules are configured to identify one or more character patterns within the database metadata table indicative of one or more PI elements, and wherein the determining the at least one portion of the database metadata table comprises determining that one or more of a column name of the one or more column names, a column attribute datatype of the one or more column attribute datatypes, or a column description of the one or more column descriptions partially matches the one or more character patterns indicative of the one or more PI elements
3, 10, 17
ODP
1, 9, 17
ODP
4, 11, 18
wherein identifying the at least one portion comprises classifying the portion as an exact match, a partial match, or a manual match
7, 15
wherein the determining the confidence score associated with the at least one portion of the database metadata table comprises: determining, based on the one or more database metadata rules, a match percentage associated with one or more of the column name, the column attribute datatype, or the column description that partially matches the one or more character patterns indicative of the one or more PI elements; and determining, based on the one or more database metadata rules, a match percentage associated with one or more rows of the first database is associated with the at least one portion of the database metadata table
5, 12, 19
wherein the personal information (PI) comprises a full name of a requesting party, a partial name of the requesting party, a birth date of the requesting party, a birth year of the requesting party, a social security number of the requesting party, or a partial social security number of the requesting party
2, 10, 18
wherein the data associated with the one or more PI elements comprises a full name associated with the requesting party, a partial name associated with the requesting party, a birth date associated with the requesting party, a birth year associated with the requesting party, a social security number associated with the requesting party, or a partial social security number associated with the requesting party
6, 13, 20
further comprising determining, based on a jurisdictional definition of PI associated with a requesting party, the one or more database metadata rules
3, 11, 19
wherein the determining the one or more database metadata rules comprises: determining, based on the request, a jurisdictional requirement for a jurisdiction associated with one or more of the requesting party or the request; and determining, based on the jurisdictional requirement, the one or more database metadata rules, wherein the jurisdictional requirement defines PI for the jurisdiction
7, 14
ODP
1, 9, 17
ODP
Claims 3, 7, 10, 14, 17 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 9, 17 of U.S. Patent No. 12,443,581 in view of Agarwal (US 2021/0182607). Although the claims at issue are not identical, they are not patentably distinct from each other because it would be obvious to one of ordinary skill in the art that the claims in the present application are unpatentable over claims 1, 9, 17 of U.S. Patent No. 12,443,581 in view of Agarwal (US 2021/0182607).
Regarding claim 3, 10, 17, claims 1, 9, 17 of U.S. Patent No. 12,443,581 fail to disclose “wherein applying the one or more database metadata rules comprises comparing column names and/or column descriptions to terms obtained from an ontology or a thesaurus derived from a natural-language request”
However, Agarwal teaches the above limitation ([0128] Metadata feature extractor 1131 includes three classifiers for generating metadata features 1145 a: table name similarity 1231 a, column name similarity 1231 b, and column data type weighting 1231 c. Each of output classes 1030 has a predefined list of common column names and common table names for that output class. [0129] Column name similarity 1231 b is another character-based neural network classifier that identifies a particular one common column name based on the comparison of column name 1215 to the list of common table names for output class 1030 a. One feature of metadata features 1145 a corresponds to this comparison. For a given value of column name 1215, column name similarity 1231 b generates the respective feature using a prediction that the schema information matches one of the common column names on the list for output class 1030 a. Column name similarity 1231 b generates additional features of metadata features 1145 a based on additional comparisons to lists corresponding to each of the remaining output classes 1030 b-1030 n.
[0130] Each output class 1030 has a predefined list of acceptable data types).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Agarwal into the claimed invention of claims 1, 9, 17 of U.S. Patent No. 12,443,581 because the references similarly disclose the processing and/or identification of data sensitivity. Consequently, one of ordinary skill in the art would be motivated to further modify the claimed invention as in claims 1, 9, 17 of U.S. Patent No. 12,443,581 to further include the column similarities as in Agarwal in order to be able to identify similarities without directly using the sensitive data itself in order to protect it from leaking.
Regarding claim 7, 14, claims 1, 9, 17 of U.S. Patent No. 12,443,581 fail to disclose “wherein generating the confidence score comprises combining a metadata-based match percentage and a row-based match percentage for the at least one portion”
However, Agarwal teaches the above limitation ([0049] scanning process 230 may determine that a particular data object has a confidence score 245 corresponding to a 70% likelihood that a credit card number is included [0070] Risk scores may be weighted using the respective confidence scores before being compared to the threshold values [0097] For each scanned data object 115, an overall security score may be determined from one or more confidence and risk scores associated with each respective data object 115 [0119] percentage of match of data items to output class regexes [0128]-[0129] comparisons of metadata).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Agarwal into the claimed invention of claims 1, 9, 17 of U.S. Patent No. 12,443,581 because the references similarly disclose the processing and/or identification of data sensitivity. Consequently, one of ordinary skill in the art would be motivated to further modify the claimed invention as in claims 1, 9, 17 of U.S. Patent No. 12,443,581 to further include the combined confidence scores as in Agarwal in order to ensure more accurate classifications of the sensitive data.
Claims 1-2, 5-6, 8, 12, 13, 15, 19-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-2, 4, 6, 10-11, 14, 20, 22 of U.S. Patent No. 12,443,606. Although the claims at issue are not identical, they are not patentably distinct from each other because one of ordinary skill in the art would recognize that 1-2, 4, 6, 10-11, 14, 20, 22 of U.S. Patent No. 12,443,606 are directed to a similar invention because they anticipate the claims in the present application.
Claim
US 19/337,154
Claim
US 12,443,606
1, 8, 15
retrieving, by a computing device, database metadata from at least two different database types without collecting row data stored in data tables of the at least two databases, wherein the database metadata comprises column names, column attribute datatypes, and column descriptions; converting, by the computing device, the database metadata from the at least two different database types to a common format; storing, by the computing device, a database metadata table comprising the converted database metadata, wherein the database metadata table comprises a first column of column names and at least one additional column comprising column attribute datatypes or column descriptions;
applying, by the computing device, one or more database metadata rules to the database metadata table to identify at least one portion of the database metadata table associated with personal information (PI);
and generating, by the computing device, a confidence score associated with the at least one portion of the database metadata table, wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata table comprise the PI
1, 10, 19
A method comprising:
retrieving, by a computing device, first database metadata for a first plurality of data tables stored within a first database without collecting row data stored in the first plurality of data tables;
retrieving, by the computing device, second database metadata for a second plurality of data tables stored within a second database without collecting row data stored in the second plurality of data tables, wherein the first database is compatible with a first database type and the second database is compatible with a second database type that is different than the first database type, wherein the first database metadata and the second database metadata each comprise a plurality of column names from the first plurality of data tables and the second plurality of data tables, respectively;
converting, by the computing device, the first database metadata from the first database type to a common format and the second database metadata from the second database type to the common format;
based on the converting of the first database metadata from the first database type to the common format and the converting of the second database metadata to the common format, storing a database metadata table comprising both of the first database metadata and the second database metadata, wherein a first column of the database metadata table comprises the plurality of column names and a second column of the database metadata table comprises one of a plurality of column attribute datatypes or a plurality of column descriptions associated with the plurality of column names;
determining, based on a request for data associated with one or more personal information (PI) elements associated with a requesting party, one or more database metadata rules, wherein the one or more database metadata rules are configured to identify one or more column names, of the plurality of column names within the database metadata table, associated with one or more data tables that comprise the data associated with the one or more PI elements;
determining, based on the one or more database metadata rules and the one or more column names, a portion of the database metadata table indicative of one or more of the first database and/or the second database, of the first plurality of data tables stored within the first database or of the second plurality of data tables stored within the second database, comprising the one or more data tables that comprise the data associated with the one or more PI elements, wherein the data associated with the one or more PI elements is stored in the one or more data tables as row data; and
sending, based on the data associated with the one or more PI elements, a response to the request
2
wherein the one or more database metadata rules comprise one or more regular expressions configured to locate character patterns indicative of the personal information (PI)
6
wherein the one or more database metadata rules are configured to identify one or more character patterns within the database metadata table indicative of the one or more PI elements
3, 10, 17
ODP
1, 10, 19
ODP
4, 11, 18
ODP
1, 10, 19
ODP
5, 12, 19
wherein the personal information (PI) comprises a full name of a requesting party, a partial name of the requesting party, a birth date of the requesting party, a birth year of the requesting party, a social security number of the requesting party, or a partial social security number of the requesting party
4, 14, 20
wherein the data associated with the one or more PI elements comprises a full name associated with the requesting party, a partial name associated with the requesting party, a birth date associated with the requesting party, a birth year associated with the requesting party, a social security number associated with the requesting party, or a partial social security number associated with the requesting party
6, 13, 20
further comprising determining, based on a jurisdictional definition of PI associated with a requesting party, the one or more database metadata rules
2, 11, 15, 22
wherein the determining the one or more database metadata rules comprises: determining, based on the request, a jurisdictional requirement for a jurisdiction associated with one or more of the requesting party or the request; and determining, based on the jurisdictional requirement, the one or more database metadata rules, wherein the jurisdictional requirement defines PI for the jurisdiction
7, 14
ODP
1, 10, 19
ODP
Claims 3, 7, 10, 14, 17 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 10, 19 of U.S. Patent No. 12,443,606 in view of Agarwal (US 2021/0182607). Although the claims at issue are not identical, they are not patentably distinct from each other because it would be obvious to one of ordinary skill in the art that the claims in the present application are unpatentable over claims 1, 10, 19 of U.S. Patent No. 12,443,606 in view of Agarwal (US 2021/0182607).
Regarding claim 3, 10, 17, claims 1, 10, 19 of U.S. Patent No. 12,443,606 fail to disclose “wherein applying the one or more database metadata rules comprises comparing column names and/or column descriptions to terms obtained from an ontology or a thesaurus derived from a natural-language request”
However, Agarwal teaches the above limitation ([0128] Metadata feature extractor 1131 includes three classifiers for generating metadata features 1145 a: table name similarity 1231 a, column name similarity 1231 b, and column data type weighting 1231 c. Each of output classes 1030 has a predefined list of common column names and common table names for that output class. [0129] Column name similarity 1231 b is another character-based neural network classifier that identifies a particular one common column name based on the comparison of column name 1215 to the list of common table names for output class 1030 a. One feature of metadata features 1145 a corresponds to this comparison. For a given value of column name 1215, column name similarity 1231 b generates the respective feature using a prediction that the schema information matches one of the common column names on the list for output class 1030 a. Column name similarity 1231 b generates additional features of metadata features 1145 a based on additional comparisons to lists corresponding to each of the remaining output classes 1030 b-1030 n.
[0130] Each output class 1030 has a predefined list of acceptable data types).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Agarwal into the claimed invention of claims 1, 10, 19 of U.S. Patent No. 12,443,606 because the references similarly disclose the processing and/or identification of data sensitivity. Consequently, one of ordinary skill in the art would be motivated to further modify the claimed invention as in claims 1, 10, 19 of U.S. Patent No. 12,443,606 to further include the column similarities as in Agarwal in order to be able to identify similarities without directly using the sensitive data itself in order to protect it from leaking.
Regarding claim 7, 14, claims 1, 10, 19 of U.S. Patent No. 12,443,606 fail to disclose “wherein generating the confidence score comprises combining a metadata-based match percentage and a row-based match percentage for the at least one portion”
However, Agarwal teaches the above limitation ([0049] scanning process 230 may determine that a particular data object has a confidence score 245 corresponding to a 70% likelihood that a credit card number is included [0070] Risk scores may be weighted using the respective confidence scores before being compared to the threshold values [0097] For each scanned data object 115, an overall security score may be determined from one or more confidence and risk scores associated with each respective data object 115 [0119] percentage of match of data items to output class regexes [0128]-[0129] comparisons of metadata).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Agarwal into the claimed invention of claims 1, 10, 19 of U.S. Patent No. 12,443,606 because the references similarly disclose the processing and/or identification of data sensitivity. Consequently, one of ordinary skill in the art would be motivated to further modify the claimed invention as in claims 1, 10, 19 of U.S. Patent No. 12,443,606 to further include the combined confidence scores as in Agarwal in order to ensure more accurate classifications of the sensitive data.
Claims 4, 11, 18 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 10, 19 of U.S. Patent No. 12,443,606 in view of Agarwal (US 2021/0182607). Although the claims at issue are not identical, they are not patentably distinct from each other because it would be obvious to one of ordinary skill in the art that the claims in the present application are unpatentable over claims 1, 10, 19 of U.S. Patent No. 12,443,606 in view of Agarwal (US 2021/0182607).
Regarding claim 4, 11, 18, claims 1, 10, 19 of U.S. Patent No. 12,443,606 fail to disclose “wherein identifying the at least one portion comprises classifying the portion as an exact match, a partial match, or a manual match”
However, Enuka teaches the above limitations at least by ([0124] values 556 and 557 in attribute field 550 only partially match value 516 in scanned field 510 and would not be counted as a full match, [0125]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Enuka into the claimed invention of claims 1, 10, 19 of U.S. Patent No. 12,443,606 because the references similarly disclose the processing and/or identification of data sensitivity. Consequently, one of ordinary skill in the art would be motivated to further modify the claimed invention as in claims 1, 10, 19 of U.S. Patent No. 12,443,606 to further include the classifying of matches as in Enuka in order to be able to control the level of match and prevent matching errors.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Independent claims 1, 8, 15 similarly recite retrieving, by a computing device, database metadata from at least two different database types without collecting row data stored in data tables of the at least two databases, wherein the database metadata comprises column names, column attribute datatypes, and column descriptions; converting, by the computing device, the database metadata from the at least two different database types to a common format; storing, by the computing device, a database metadata table comprising the converted database metadata, wherein the database metadata table comprises a first column of column names and at least one additional column comprising column attribute datatypes or column descriptions; applying, by the computing device, one or more database metadata rules to the database metadata table to identify at least one portion of the database metadata table associated with personal information (PI); and generating, by the computing device, a confidence score associated with the at least one portion of the database metadata table, wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata table comprise the PI.
The limitations of converting, ..., the database metadata from the at least two different database types to a common format; applying, ..., one or more database metadata rules to the database metadata table to identify at least one portion of the database metadata table associated with personal information (PI); and generating,..., a confidence score associated with the at least one portion of the database metadata table, wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata table comprise the PI, as drafted, are processes that, under their broadest reasonable interpretation, cover mental processes but from the recitation of implementing them on generic computer components. That is, other than reciting “by the computing device” nothing in the claim elements preclude the steps from practically being performed in the mind. For example, but for the “by the computing device” language, the limitations pertaining to “converting”, “applying”, and “generating” in the context of this claim encompasses the user judging a conversion of database metadata to a common format, judging an application of rules to the metadata table, and judging a confidence score. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, claims 1, 8, 15 recite an abstract idea (Step 2A, Prong 1).
This judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of – retrieving, by a computing device, database metadata from at least two different database types without collecting row data stored in data tables of the at least two databases, wherein the database metadata comprises column names, column attribute datatypes, and column descriptions; by the computing device; storing, by the computing device, a database metadata table comprising the converted database metadata, wherein the database metadata table comprises a first column of column names and at least one additional column comprising column attribute datatypes or column descriptions; by the computing device. The non-transitory computer-readable media, apparatus, processors, and memory are recited at a high-level of generality (i.e., as generic computer devices performing generic computer functions) and do not meaningfully limit the claim. The additional elements pertaining to “retrieving”, and “storing” represent insignificant extra-solution activities to the judicial exception and are mere data gathering steps. Accordingly, these additional elements, individually and in combination, do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea (Step 2A, Prong 2).
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements pertaining to “retrieving”, and “storing” represent insignificant extra-solution activities that are well-understood, routine, and conventional activities previously known to the industry. That is, these limitations represent well-understood, routine, conventional activities in the fields of data processing and/or data storage and retrieval and are merely directed to the well-understood, routine, conventional activity of storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). Therefore, these limitations, both individually and in combination, fail to amount to an inventive concept because they merely append well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception, and thus, do not cause the claim to amount to significantly more than the judicial exception. (Step 2B). Accordingly, claims 1, 8, 15 are not patent eligible.
Claims 2-7, 9-14, 16-20 depend on claims 1, 8, 15 and include all the limitations of these claims. Therefore, these claims are directed to the same abstract idea and the analysis must proceed to (Step 2A, Prong 2).
Claims 2, 9, 16 similarly recite additional limitations pertaining to the metadata rules. This judicial exception is not integrated into a practical application. The additional elements represent further mental process steps of judging an application of the rules which comprise regex expressions as in the independent claims, and do not preclude this step from being performed mentally. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. This additional step is considered an abstract idea (mental process step) and does not integrate the judicial exception into a practical application.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements represent further mental process steps. Therefore, these additional limitations are not sufficient to amount to significantly more than the judicial exception. Claims 2, 9, 16 are not patent eligible.
Claims 3, 10, 17 similarly recite additional limitations pertaining to applying the one or more database metadata rules. This judicial exception is not integrated into a practical application. The additional elements further limit mental process steps as recited in the independent claims, but, do not preclude these steps from being performed in the mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. This additional step is considered an abstract idea (mental process step) and does not integrate the judicial exception into a practical application.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements represent further mental process steps. Therefore, these additional limitations are not sufficient to amount to significantly more than the judicial exception. These claims are not patent eligible.
Claims 4, 11, 18 similarly recite additional limitations pertaining to applying the one or more database metadata rules. This judicial exception is not integrated into a practical application. The additional elements further limit mental process steps as recited in the independent claims, but, do not preclude these steps from being performed in the mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. This additional step is considered an abstract idea (mental process step) and does not integrate the judicial exception into a practical application.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements represent further mental process steps. Therefore, these additional limitations are not sufficient to amount to significantly more than the judicial exception. These claims are not patent eligible.
Claims 5, 12, 19 similarly recite additional limitations pertaining to the personal information and applying of the one or more database metadata rules. This judicial exception is not integrated into a practical application. The additional elements further limit mental process steps as recited in the independent claims, but, do not preclude these steps from being performed in the mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. This additional step is considered an abstract idea (mental process step) and does not integrate the judicial exception into a practical application.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements represent further mental process steps. Therefore, these additional limitations are not sufficient to amount to significantly more than the judicial exception. These claims are not patent eligible.
Claims 6, 13, 20 similarly recite additional limitations pertaining to determining the metadata rules. This judicial exception is not integrated into a practical application. The additional elements represent further mental process steps of judging the metadata rules. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. This additional step is considered an abstract idea (mental process step) and does not integrate the judicial exception into a practical application.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements represent further mental process steps. Therefore, these additional limitations are not sufficient to amount to significantly more than the judicial exception. Claims 6, 13, 20 are not patent eligible.
Claims 7, 14 similarly recite additional limitations pertaining to generating the confidence score. This judicial exception is not integrated into a practical application. The additional elements further limit mental process steps as recited in the independent claims, but, do not preclude these steps from being performed in the mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. This additional step is considered an abstract idea (mental process step) and does not integrate the judicial exception into a practical application.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements represent further mental process steps. Therefore, these additional limitations are not sufficient to amount to significantly more than the judicial exception. These claims are not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 7-10, 14-17 are rejected under 35 U.S.C. 103 as being unpatentable over Umansky (US 2019/0138625) in view of Agarwal (US 2021/0182607).
Regarding claim 1, Umansky discloses:
A method comprising: retrieving, by a computing device, database metadata from at least two different database types..., wherein the database metadata comprises column names, column attribute datatypes, and column descriptions ([0045] data objects stored within a first and second database 110a, 110b, etc. in each stored in different security zones as shown in in Fig. 1 [0046] data may be stored in each database 110 in any number of a variety of data formats [0052] To protect sensitive information, the transmitted metadata 247 does not include the corresponding data objects 115 that are stored in the database [0107] The scan, as shown, includes multiple actions, beginning with determining metadata 1020 for a portion of database 1010, wherein metadata 1020 includes schema information 1023...“schema information” refers to metadata that describes a structure of a data item and/or a set of data items. For example, set of data items 1015 may correspond to a column of data items stored in a tabular format. In such an example, schema information includes row and column information for a particular data item in a table, a table name, a column and row headings, and other similar information. A process executing on computer system 1001 accesses the portion of database 1010 that includes set of data items 1015 and extracts metadata 1020, including schema information 1023, associated with set of data items 1015 [0116] In the illustrated example, data object 1102 is a table of data items [0117], [0127] The associated metadata is retrieved by metadata collector 1121, including table name 1202, column name 1215, and column data type 1215 a-d);
converting, by the computing device, the database metadata from the at least two different database types to a common format ([0046] Performing the risk analysis further includes converting the stored data objects 115 from a particular data format to a common data format, different from the particular data format. In some embodiments, data may be stored in database 110 using any number of a variety of data formats. In order to simplify the scanning process, computing device 101 uses conversion process 220 to convert data objects 115 from one or more particular data formats into the common data format, thereby generating converted data objects 215 [0051] Risk determination process 240 conveys metadata 247 to control process 250 that is performed by a computing device in repository 160...metadata 247 may be stored within a storage medium included in repository 160 or in a separate database such as risk analysis database 270 [0107] metadata 1020 includes schema information 1023...“schema information” refers to metadata that describes a structure of a data item and/or a set of data items. For example, set of data items 1015 may correspond to a column of data items stored in a tabular format. In such an example, schema information includes row and column information for a particular data item in a table, a table name, a column and row headings, and other similar information);
storing, by the computing device, a database metadata table comprising the converted database metadata, wherein the database metadata table comprises a first column of column names and at least one additional column comprising column attribute datatypes or column descriptions; ... the database metadata table ([0021] The data sources of data store 114 may comprise database tables interrelated via a database schema defined by metadata which is also stored in data store 114. The metadata may also include sensitivity information associated with one or more of the data sources of data store 114 as will be described below [0028]-[0029] disclose a metadata table 310 that is created during design-time by adding, modifying and deleting metadata that represents the schema of the data sources of the data stores [Fig. 3] shows the metadata table 310 which includes column names in the first column and data types in the second column extracted as well as other descriptions in additional columns).
Umansky fails to disclose “without collecting row data stored in data tables of the at least two databases; applying, by the computing device, one or more database metadata rules to the database metadata table to identify at least one portion of the database metadata table associated with personal information (PI); and generating, by the computing device, a confidence score associated with the at least one portion of the database metadata table, wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata table comprise the PI”
However, Agarwal teaches without collecting row data stored in data tables of the at least two databases ([0022] multiple databases [0023] Since the data objects being scanned may include sensitive information, the results are conveyed without conveying the actual data objects to the repository zone [0052] To protect sensitive information, the transmitted metadata 247 does not include the corresponding data objects 115 that are stored in the database);
applying, by the computing device, one or more database metadata rules to the database metadata ... to identify at least one portion of the database metadata table associated with personal information (PI) ([0057] two tables depicting respective examples of a set of security rules and a set of scan models are illustrated. As described above, security rules establish what types of data objects are permissible to be stored in a particular database and criteria for how each permitted data object is to be stored. Scan models establish criteria such as the types of data objects that will be scanned as well as a type of scan to perform on each established type. Security rules 130 includes five rules 330 a-330 e, each rule identifying a type of information that may be found in a data object, such as data objects 115 in FIGS. 1 and 2. Each rule further includes a respective criterion for storing the corresponding type of information. Scan models 235 includes four models 335 a-335 d, each model identifying a type of data object and respective criteria for scanning the corresponding data object type. Security rules 130 and scan models 235 may be applied to one or more databases within a given security zone, such as security zone 105 a or 105 b. [0060] If the scan of the encrypted file does not include decrypting the file, then the credit card data may be detected based on hints that the encrypted file includes credit card data, such as a particular data pattern that is indicative of a 16-digit number [0107] The scan, as shown, includes multiple actions, beginning with determining metadata 1020 for a portion of database 1010, wherein metadata 1020 includes schema information 1023. Metadata 1020 includes information about a data item, excluding actual data included in the data item. For example, metadata 1020 may include information such as a timestamp for when the file was last edited, a size of the data item, identification of a user that created the data item, and the like. Schema information 1023 is a subset of metadata 1020 that includes information regarding a structure for how a data item is stored in database 1010. As used herein, “schema information” refers to metadata that describes a structure of a data item and/or a set of data items. For example, set of data items 1015 may correspond to a column of data items stored in a tabular format. In such an example, schema information includes row and column information for a particular data item in a table, a table name, a column and row headings, and other similar information. A process executing on computer system 1001 accesses the portion of database 1010 that includes set of data items 1015 and extracts metadata 1020, including schema information 1023, associated with set of data items 1015. Based on the extracted metadata, a set of features are generated. Each feature has a value that is indicative of a correlation between the extracted metadata and a respective one of the output classes 1030. As used herein, a “correlation” refers to a degree to which a characteristic of a data item matches a typical characteristic of a particular output class. For example, a feature may have a value between “0” and “1,” with “0” representing no match between the extracted metadata and “1” representing a very close match. Additional details regarding the metadata extraction process will be disclosed below in reference to FIG. 12);
and generating, by the computing device, a confidence score associated with the at least one portion of the database metadata ..., wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata ... comprise the PI ([0033] To perform the first and second scans, computer system 103, utilizing the respective computing devices 101 a and 101 b, uses the one or more criteria to determine a confidence score for a particular one of data object 115 a-115 f. This confidence score indicates a level of confidence that the particular data object matches the particular classification [0052] To protect sensitive information, the transmitted metadata 247 does not include the corresponding data objects 115 that are stored in the database. Metadata 247, however, includes confidence score 245 and risk score 243.).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Agarwal into the teaching of Umansky because the references similarly disclose the processing and/or identification of data sensitivity. Consequently, one of ordinary skill in the art would be motivated to further modify the system as in Umansky to further include the scanning, applying of rules and confidence scores as in Agarwal in order to “improve a speed of execution of the risk analysis scan and/or to improve an accuracy of the scan results” (Agarwal, [0046]).
As per claim 2, claim 1 is incorporated, Agarwal further discloses:
wherein the one or more database metadata rules comprise one or more regular expressions configured to locate character patterns indicative of the personal information (PI) ([0033] the particular data object is encrypted, and to determine the confidence score for an encrypted data object, computer system 103 determines the confidence score without performing a decryption operation. For example, computer system 103 may not have access to a decryption key for an encrypted data object. In such a case, computer system 103 may evaluate the encrypted data object by looking for particular patterns in the encrypted data that may be indicative of particular data types such as credit card numbers or email addresses. The confidence score may typically be lower for encrypted data than for unencrypted data [0060] If the scan of the encrypted file does not include decrypting the file, then the credit card data may be detected based on hints that the encrypted file includes credit card data, such as a particular data pattern that is indicative of a 16-digit number).
As per claim 3, claim 1 is incorporated, Agarwal further discloses:
wherein applying the one or more database metadata rules comprises comparing column names and/or column descriptions to terms obtained from an ontology or a thesaurus derived from a natural-language request ([0128] Metadata feature extractor 1131 includes three classifiers for generating metadata features 1145 a: table name similarity 1231 a, column name similarity 1231 b, and column data type weighting 1231 c. Each of output classes 1030 has a predefined list of common column names and common table names for that output class. [0129] Column name similarity 1231 b is another character-based neural network classifier that identifies a particular one common column name based on the comparison of column name 1215 to the list of common table names for output class 1030 a. One feature of metadata features 1145 a corresponds to this comparison. For a given value of column name 1215, column name similarity 1231 b generates the respective feature using a prediction that the schema information matches one of the common column names on the list for output class 1030 a. Column name similarity 1231 b generates additional features of metadata features 1145 a based on additional comparisons to lists corresponding to each of the remaining output classes 1030 b-1030 n.
[0130] Each output class 1030 has a predefined list of acceptable data types).
As per claim 7, claim 1 is incorporated, Agarwal further discloses:
wherein generating the confidence score comprises combining a metadata-based match percentage and a row-based match percentage for the at least one portion ([0049] scanning process 230 may determine that a particular data object has a confidence score 245 corresponding to a 70% likelihood that a credit card number is included [0070] Risk scores may be weighted using the respective confidence scores before being compared to the threshold values [0097] For each scanned data object 115, an overall security score may be determined from one or more confidence and risk scores associated with each respective data object 115 [0119] percentage of match of data items to output class regexes [0128]-[0129] comparisons of metadata).
Regarding claim 8, Umansky discloses:
One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to: retrieve database metadata from at least two different database types ..., wherein the database metadata comprises column names, column attribute datatypes, and column descriptions ([0045] data objects stored within a first and second database 110a, 110b, etc. in each stored in different security zones as shown in in Fig. 1 [0046] data may be stored in each database 110 in any number of a variety of data formats [0052] To protect sensitive information, the transmitted metadata 247 does not include the corresponding data objects 115 that are stored in the database [0107] The scan, as shown, includes multiple actions, beginning with determining metadata 1020 for a portion of database 1010, wherein metadata 1020 includes schema information 1023...“schema information” refers to metadata that describes a structure of a data item and/or a set of data items. For example, set of data items 1015 may correspond to a column of data items stored in a tabular format. In such an example, schema information includes row and column information for a particular data item in a table, a table name, a column and row headings, and other similar information. A process executing on computer system 1001 accesses the portion of database 1010 that includes set of data items 1015 and extracts metadata 1020, including schema information 1023, associated with set of data items 1015 [0116] In the illustrated example, data object 1102 is a table of data items [0117], [0127] The associated metadata is retrieved by metadata collector 1121, including table name 1202, column name 1215, and column data type 1215 a-d);
convert the database metadata from the at least two different database types to a common format ([0046] Performing the risk analysis further includes converting the stored data objects 115 from a particular data format to a common data format, different from the particular data format. In some embodiments, data may be stored in database 110 using any number of a variety of data formats. In order to simplify the scanning process, computing device 101 uses conversion process 220 to convert data objects 115 from one or more particular data formats into the common data format, thereby generating converted data objects 215 [0051] Risk determination process 240 conveys metadata 247 to control process 250 that is performed by a computing device in repository 160...metadata 247 may be stored within a storage medium included in repository 160 or in a separate database such as risk analysis database 270 [0107] metadata 1020 includes schema information 1023...“schema information” refers to metadata that describes a structure of a data item and/or a set of data items. For example, set of data items 1015 may correspond to a column of data items stored in a tabular format. In such an example, schema information includes row and column information for a particular data item in a table, a table name, a column and row headings, and other similar information);
store a database metadata table comprising the converted database metadata, wherein the database metadata table comprises a first column of column names and at least one additional column comprising column attribute datatypes or column descriptions; the database metadata table ([0021] The data sources of data store 114 may comprise database tables interrelated via a database schema defined by metadata which is also stored in data store 114. The metadata may also include sensitivity information associated with one or more of the data sources of data store 114 as will be described below [0028]-[0029] disclose a metadata table 310 that is created during design-time by adding, modifying and deleting metadata that represents the schema of the data sources of the data stores [Fig. 3] shows the metadata table 310 which includes column names in the first column and data types in the second column extracted as well as other descriptions in additional columns).
Umansky fails to disclose “without collecting row data stored in data tables of the at least two databases; apply one or more database metadata rules to the database metadata table to identify at least one portion of the database metadata table associated with personal information (PI); and generate a confidence score associated with the at least one portion of the database metadata table, wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata table comprise the PI”
However, Agarwal teaches without collecting row data stored in data tables of the at least two databases ([0022] multiple databases [0023] Since the data objects being scanned may include sensitive information, the results are conveyed without conveying the actual data objects to the repository zone [0052] To protect sensitive information, the transmitted metadata 247 does not include the corresponding data objects 115 that are stored in the database);
apply one or more database metadata rules to the database metadata ... to identify at least one portion of the database metadata table associated with personal information (PI) ([0057] two tables depicting respective examples of a set of security rules and a set of scan models are illustrated. As described above, security rules establish what types of data objects are permissible to be stored in a particular database and criteria for how each permitted data object is to be stored. Scan models establish criteria such as the types of data objects that will be scanned as well as a type of scan to perform on each established type. Security rules 130 includes five rules 330 a-330 e, each rule identifying a type of information that may be found in a data object, such as data objects 115 in FIGS. 1 and 2. Each rule further includes a respective criterion for storing the corresponding type of information. Scan models 235 includes four models 335 a-335 d, each model identifying a type of data object and respective criteria for scanning the corresponding data object type. Security rules 130 and scan models 235 may be applied to one or more databases within a given security zone, such as security zone 105 a or 105 b. [0060] If the scan of the encrypted file does not include decrypting the file, then the credit card data may be detected based on hints that the encrypted file includes credit card data, such as a particular data pattern that is indicative of a 16-digit number [0107] The scan, as shown, includes multiple actions, beginning with determining metadata 1020 for a portion of database 1010, wherein metadata 1020 includes schema information 1023. Metadata 1020 includes information about a data item, excluding actual data included in the data item. For example, metadata 1020 may include information such as a timestamp for when the file was last edited, a size of the data item, identification of a user that created the data item, and the like. Schema information 1023 is a subset of metadata 1020 that includes information regarding a structure for how a data item is stored in database 1010. As used herein, “schema information” refers to metadata that describes a structure of a data item and/or a set of data items. For example, set of data items 1015 may correspond to a column of data items stored in a tabular format. In such an example, schema information includes row and column information for a particular data item in a table, a table name, a column and row headings, and other similar information. A process executing on computer system 1001 accesses the portion of database 1010 that includes set of data items 1015 and extracts metadata 1020, including schema information 1023, associated with set of data items 1015. Based on the extracted metadata, a set of features are generated. Each feature has a value that is indicative of a correlation between the extracted metadata and a respective one of the output classes 1030. As used herein, a “correlation” refers to a degree to which a characteristic of a data item matches a typical characteristic of a particular output class. For example, a feature may have a value between “0” and “1,” with “0” representing no match between the extracted metadata and “1” representing a very close match. Additional details regarding the metadata extraction process will be disclosed below in reference to FIG. 12);
and generate a confidence score associated with the at least one portion of the database metadata..., wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata ... comprise the PI ([0033] To perform the first and second scans, computer system 103, utilizing the respective computing devices 101 a and 101 b, uses the one or more criteria to determine a confidence score for a particular one of data object 115 a-115 f. This confidence score indicates a level of confidence that the particular data object matches the particular classification [0052] To protect sensitive information, the transmitted metadata 247 does not include the corresponding data objects 115 that are stored in the database. Metadata 247, however, includes confidence score 245 and risk score 243.).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Agarwal into the teaching of Umansky because the references similarly disclose the processing and/or identification of data sensitivity. Consequently, one of ordinary skill in the art would be motivated to further modify the system as in Umansky to further include the scanning, applying of rules and confidence scores as in Agarwal in order to “improve a speed of execution of the risk analysis scan and/or to improve an accuracy of the scan results” (Agarwal, [0046]).
Regarding claim 15, Umansky discloses:
An apparatus, comprising: one or more processors; and a memory storing processor-executable instructions that, when executed by the one or more processors, cause the apparatus to: retrieve database metadata from at least two different database types..., wherein the database metadata comprises column names, column attribute datatypes, and column descriptions ([0045] data objects stored within a first and second database 110a, 110b, etc. in each stored in different security zones as shown in in Fig. 1 [0046] data may be stored in each database 110 in any number of a variety of data formats [0052] To protect sensitive information, the transmitted metadata 247 does not include the corresponding data objects 115 that are stored in the database [0107] The scan, as shown, includes multiple actions, beginning with determining metadata 1020 for a portion of database 1010, wherein metadata 1020 includes schema information 1023...“schema information” refers to metadata that describes a structure of a data item and/or a set of data items. For example, set of data items 1015 may correspond to a column of data items stored in a tabular format. In such an example, schema information includes row and column information for a particular data item in a table, a table name, a column and row headings, and other similar information. A process executing on computer system 1001 accesses the portion of database 1010 that includes set of data items 1015 and extracts metadata 1020, including schema information 1023, associated with set of data items 1015 [0116] In the illustrated example, data object 1102 is a table of data items [0117], [0127] The associated metadata is retrieved by metadata collector 1121, including table name 1202, column name 1215, and column data type 1215 a-d);
convert the database metadata from the at least two different database types to a common format ([0046] Performing the risk analysis further includes converting the stored data objects 115 from a particular data format to a common data format, different from the particular data format. In some embodiments, data may be stored in database 110 using any number of a variety of data formats. In order to simplify the scanning process, computing device 101 uses conversion process 220 to convert data objects 115 from one or more particular data formats into the common data format, thereby generating converted data objects 215 [0051] Risk determination process 240 conveys metadata 247 to control process 250 that is performed by a computing device in repository 160...metadata 247 may be stored within a storage medium included in repository 160 or in a separate database such as risk analysis database 270 [0107] metadata 1020 includes schema information 1023...“schema information” refers to metadata that describes a structure of a data item and/or a set of data items. For example, set of data items 1015 may correspond to a column of data items stored in a tabular format. In such an example, schema information includes row and column information for a particular data item in a table, a table name, a column and row headings, and other similar information);
store a database metadata table comprising the converted database metadata, wherein the database metadata table comprises a first column of column names and at least one additional column comprising column attribute datatypes or column descriptions; ... the database metadata table ([0021] The data sources of data store 114 may comprise database tables interrelated via a database schema defined by metadata which is also stored in data store 114. The metadata may also include sensitivity information associated with one or more of the data sources of data store 114 as will be described below [0028]-[0029] disclose a metadata table 310 that is created during design-time by adding, modifying and deleting metadata that represents the schema of the data sources of the data stores [Fig. 3] shows the metadata table 310 which includes column names in the first column and data types in the second column extracted as well as other descriptions in additional columns).
Umansky fails to disclose “without collecting row data stored in data tables of the at least two databases; apply one or more database metadata rules to the database metadata table to identify at least one portion of the database metadata table associated with personal information (PI); and generate a confidence score associated with the at least one portion of the database metadata table, wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata table comprise the PI”
However, Agarwal teaches without collecting row data stored in data tables of the at least two databases ([0022] multiple databases [0023] Since the data objects being scanned may include sensitive information, the results are conveyed without conveying the actual data objects to the repository zone [0052] To protect sensitive information, the transmitted metadata 247 does not include the corresponding data objects 115 that are stored in the database);
Apply one or more database metadata rules to the database metadata ... to identify at least one portion of the database metadata table associated with personal information (PI) ([0057] two tables depicting respective examples of a set of security rules and a set of scan models are illustrated. As described above, security rules establish what types of data objects are permissible to be stored in a particular database and criteria for how each permitted data object is to be stored. Scan models establish criteria such as the types of data objects that will be scanned as well as a type of scan to perform on each established type. Security rules 130 includes five rules 330 a-330 e, each rule identifying a type of information that may be found in a data object, such as data objects 115 in FIGS. 1 and 2. Each rule further includes a respective criterion for storing the corresponding type of information. Scan models 235 includes four models 335 a-335 d, each model identifying a type of data object and respective criteria for scanning the corresponding data object type. Security rules 130 and scan models 235 may be applied to one or more databases within a given security zone, such as security zone 105 a or 105 b. [0060] If the scan of the encrypted file does not include decrypting the file, then the credit card data may be detected based on hints that the encrypted file includes credit card data, such as a particular data pattern that is indicative of a 16-digit number [0107] The scan, as shown, includes multiple actions, beginning with determining metadata 1020 for a portion of database 1010, wherein metadata 1020 includes schema information 1023. Metadata 1020 includes information about a data item, excluding actual data included in the data item. For example, metadata 1020 may include information such as a timestamp for when the file was last edited, a size of the data item, identification of a user that created the data item, and the like. Schema information 1023 is a subset of metadata 1020 that includes information regarding a structure for how a data item is stored in database 1010. As used herein, “schema information” refers to metadata that describes a structure of a data item and/or a set of data items. For example, set of data items 1015 may correspond to a column of data items stored in a tabular format. In such an example, schema information includes row and column information for a particular data item in a table, a table name, a column and row headings, and other similar information. A process executing on computer system 1001 accesses the portion of database 1010 that includes set of data items 1015 and extracts metadata 1020, including schema information 1023, associated with set of data items 1015. Based on the extracted metadata, a set of features are generated. Each feature has a value that is indicative of a correlation between the extracted metadata and a respective one of the output classes 1030. As used herein, a “correlation” refers to a degree to which a characteristic of a data item matches a typical characteristic of a particular output class. For example, a feature may have a value between “0” and “1,” with “0” representing no match between the extracted metadata and “1” representing a very close match. Additional details regarding the metadata extraction process will be disclosed below in reference to FIG. 12);
and generate a confidence score associated with the at least one portion of the database metadata ..., wherein the confidence score is indicative of a level of confidence that one or more data tables indicated by the at least one portion of the database metadata ... comprise the PI ([0033] To perform the first and second scans, computer system 103, utilizing the respective computing devices 101 a and 101 b, uses the one or more criteria to determine a confidence score for a particular one of data object 115 a-115 f. This confidence score indicates a level of confidence that the particular data object matches the particular classification [0052] To protect sensitive information, the transmitted metadata 247 does not include the corresponding data objects 115 that are stored in the database. Metadata 247, however, includes confidence score 245 and risk score 243.).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Agarwal into the teaching of Umansky because the references similarly disclose the processing and/or identification of data sensitivity. Consequently, one of ordinary skill in the art would be motivated to further modify the system as in Umansky to further include the scanning, applying of rules and confidence scores as in Agarwal in order to “improve a speed of execution of the risk analysis scan and/or to improve an accuracy of the scan results” (Agarwal, [0046]).
Claims 9-10, 14, 16-17 recite similar claim limitations as the method of claims 2-3, 7, except that they set forth the claimed invention as one or more non-transitory computer-readable media or an apparatus and, as such, they are rejected for the same reasons as applied hereinabove.
Claims 4, 11, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Umansky (US 2019/0138625) in view of Agarwal (US 2021/0182607) and further in view of Enuka (US 2020/0050966).
As per claim 4, claim 1 is incorporated, Umansky, Agarwal fail to disclose “wherein identifying the at least one portion comprises classifying the portion as an exact match, a partial match, or a manual match”
However, Enuka teaches the above limitations at least by ([0124] values 556 and 557 in attribute field 550 only partially match value 516 in scanned field 510 and would not be counted as a full match, [0125]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Enuka into the teaching of Umansky, Agarwal because the references similarly disclose the processing and/or identification of data sensitivity. Consequently, one of ordinary skill in the art would be motivated to further modify the system as in the combination of references to further include the classifying of matches as in Enuka to be able to control the level of match and prevent matching errors.
Claims 11, 18 recite similar claim limitations as the method of claim 4, except that they set forth the claimed invention as one or more non-transitory computer-readable media or an apparatus and, as such, they are rejected for the same reasons as applied hereinabove.
Claims 5-6, 12-13, 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Umansky (US 2019/0138625) in view of Agarwal (US 2021/0182607) and further in view of Barday (US 2019/0179799).
As per claim 5, claim 1 is incorporated, Umansky, Agarwal fail to disclose “wherein the personal information (PI) comprises a full name of a requesting party, a partial name of the requesting party, a birth date of the requesting party, a birth year of the requesting party, a social security number of the requesting party, or a partial social security number of the requesting party”
However, Barday teaches the above limitations at least by ([0177] “the system is configured to use intelligent identity scanning (e.g., as described above) to identify the requestor's personal data and related information that is to be used to fulfill the request” [0297] “FIG. 42 depicts an exemplary data subject access request form that the system may substantially automatically generate, complete and/or submit for the data subject on the data subject's behalf. As shown in this figure, the system may complete information such as, for example: (1) what type of requestor the data subject is (e.g., employee, customer, etc.); (2) what the request involves (e.g., deleting data, etc.); (3) the requestor's first name; (4) the requestor's last name; (5) the requestor's email address; (6) the requestor's telephone number; (7) the requestor's home address; and/or (8) one or more details associated with the request”).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Barday into the teaching of Agarwal, Umansky because the references similarly disclose data retrieval involving PII. Consequently, one of ordinary skill in the art would be motivated to further modify the system as in the combination of references to further include the processing of requests for private data based on the requestors personal data as in Barday so that the system can “prevent interception of the data or unwarranted access to the additional information” (Barday, [0222]).
As per claim 6, claim 1 is incorporated, Umansky, Agarwal fail to disclose “further comprising determining, based on a jurisdictional definition of PI associated with a requesting party, the one or more database metadata rules”
However, Barday teaches the above limitations at least by ([0314] “when executing the DSAR Processing via Local Storage Node Module 4700, the system begins, at Step 4710, by receiving a data subject access request associated with a data subject.” [0316] “Continuing to Step 4720, the system is configured to identify a suitable local storage node based at least in part on the request and/or the data subject. For example, the system may be configured to identify the suitable local storage node (e.g., suitable one or more local storage nodes) based at least in part on: (1) a jurisdiction in which the data subject resides; (2) a country in which the data subject resides; (3) a jurisdiction from which the data subject made the request” [0318] “the system may identify the suitable storage node based at least in part on one or more residency laws about storage, one or more country-based business rules, one or more rules defined by one or more privacy administrators, etc”).).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the teaching of Barday into the teaching of Agarwal, Umansky because the references similarly disclose data retrieval involving PII. Consequently, one of ordinary skill in the art would be motivated to further modify the system as in the combination of references to further include the processing of requests for private data based on the jurisdiction of the subject that made the request as in Barday so that the system ensures that any legal requirements, such as those mandates by state, local, or federal government, are not violated by providing private data to unauthorized parties.
Claims 12-13, 19-20 recite similar claim limitations as the method of claims 5-6, except that they set forth the claimed invention as one or more non-transitory computer-readable media or an apparatus and, as such, they are rejected for the same reasons as applied hereinabove.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM P BARTLETT whose telephone number is (469)295-9085. The examiner can normally be reached on M-Th 11:30-8:30, F 11-3.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sherief Badawi can be reached on 571-272-9782. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WILLIAM P BARTLETT/
Primary Examiner, Art Unit 2169