DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
1.This action is responsive to Response to Election/Restriction filed on June 22, 2026. Applicant elects Group I, claims 1-8 and 13-20. Claims 9-12 are hereby cancelled accordingly without prejudice. Therefore, claims 1-8 and 13-20 are pending and addressed below.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Double Patenting
2. The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the "right to exclude" granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory obviousness-type double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Omum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement.
Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b).
Claims 1-8 and 13-19 are rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1-8, 9-13 and 15 of US Patent number US 12118107 B2. The conflicting claims are not identical, they are not patentably distinct from each other because the current application contains claims that are broader in scope than the claims of the patent number US 12118107 B2 and are anticipated by the claims 1-8, 9-13 and 15.
This is a Non-provisional double patenting rejection.
Claims Comparison Table
US Application 18/884,841
US Patent 12118107 B2
1. A method for detecting sensitive information, comprising: determining a first likelihood that a record contains at least a given type of sensitive information using a first detection technique comprising: providing one or more inputs to a machine learning model based on the record; and receiving the first likelihood as an output from the machine learning model based on the one or more inputs; determining a second likelihood that the record contains at least the given type of sensitive information using a second detection technique comprising a search of the record; applying a policy to determine whether the record contains sensitive information based on the first likelihood and the second likelihood; performing one or more actions based on whether the record contains sensitive information; and revising the policy based on a ground truth label indicating whether the record contains sensitive information.
1. A method for detecting sensitive information, comprising: determining a first likelihood that a record contains at least a given type of sensitive information using a first detection technique comprising: providing one or more inputs to a machine learning model based on the record; and receiving the first likelihood as an output from the machine learning model based on the one or more inputs; determining a second likelihood that the record contains at least the given type of sensitive information using a second detection technique comprising a search of the record; applying a policy to determine whether the record contains sensitive information based on the first likelihood and the second likelihood, wherein the policy states that the second detection technique is used when the first likelihood determined using the first detection technique is above a threshold likelihood; performing one or more actions based on whether the record contains sensitive information; receiving a ground truth label indicating whether the record contains sensitive information; and revising the policy based on the ground truth label.
2. The method of claim 1, wherein the search of the record comprises: converting text in the record into a plurality of tokens; searching the plurality of tokens for a first set of search terms; detecting a first term from the first set of search terms in the plurality of tokens; and searching the plurality of tokens for a second set of search terms within a given proximity of the first term.
2. The method of claim 1, wherein the search of the record comprises: converting text in the record into a plurality of tokens; searching the plurality of tokens for a first set of search terms; detecting a first term from the first set of search terms in the plurality of tokens; and searching the plurality of tokens for a second set of search terms within a given proximity of the first term.
3. The method of claim 2, further comprising: generating the first set of search terms by excluding, from a list of first names, one or more exceptional first names and one or more first names longer than a given number of characters; and generating the second set of search terms by excluding, from a list of last names, one or more exceptional last names and one or more last names longer than the given number of characters.
3. The method of claim 2, further comprising: generating the first set of search terms by excluding, from a list of first names, one or more exceptional first names and one or more first names longer than a given number of characters; and generating the second set of search terms by excluding, from a list of last names, one or more exceptional last names and one or more last names longer than the given number of characters.
4. The method of claim 2, wherein: the first set of search terms comprises years in a given range; and the second set of search terms comprises months.
4. The method of claim 2, wherein: the first set of search terms comprises years in a given range; and the second set of search terms comprises months.
5. The method of claim 1, further comprising: selecting the second detection technique from a plurality of detection techniques based on the first likelihood.
5. The method of claim 1, further comprising selecting the second detection technique from a plurality of detection techniques based on the first likelihood.
6. The method of claim 1, wherein providing the one or more inputs to the machine learning model based on the record comprises providing, as the one or more inputs to the machine learning model, one or more of: tokens determined from the record; or characters from the record.
6. The method of claim 1, wherein providing the one or more inputs to the machine learning model based on the record comprises providing, as the one or more inputs to the machine learning model, one or more of: tokens determined from the record; or characters from the record.
7. The method of claim 1, wherein performing the one or more actions based on whether the record contains sensitive information comprises one or more of: removing one or more sensitive information items from the record; encrypting one or more sensitive information items in the record; generating a notification indicating whether the record contains sensitive information; or modifying code relating to generation of the record.
7. The method of claim 1, wherein performing the one or more actions based on whether the record contains sensitive information comprises one or more of: removing one or more sensitive information items from the record; encrypting one or more sensitive information items in the record; generating a notification indicating whether the record contains sensitive information; or modifying code relating to generation of the record.
8. The method of claim 1, wherein revising the policy based on the ground truth label comprises adjusting an order of performing the first detection technique and the second detection technique from a first order in which the second detection technique is used based on the first likelihood being above a threshold likelihood to a second order in which the first detection technique is used based on the second likelihood being above the threshold likelihood.
15. The method of claim 1, wherein revising the policy based on the ground truth label comprises adjusting an order of performing the first detection technique and the second detection technique from a first order in which the second detection technique is used based on the first likelihood being above the threshold likelihood to a second order in which the first detection technique is used based on the second likelihood being above the threshold likelihood.
13. A system for detecting sensitive information, comprising: one or more processors; and a memory comprising instructions that, when executed by the one or more processors, cause the system to: determine a first likelihood that a record contains at least a given type of sensitive information using a first detection technique comprising: providing one or more inputs to a machine learning model based on the record; and receiving the first likelihood as an output from the machine learning model based on the one or more inputs; determine a second likelihood that the record contains at least the given type of sensitive information using a second detection technique comprising a search of the record; apply a policy to determine whether the record contains sensitive information based on the first likelihood and the second likelihood; perform one or more actions based on whether the record contains sensitive information; and revise the policy based on a ground truth label indicating whether the record contains sensitive information.
8. A system for detecting sensitive information, comprising: one or more processors; and a memory comprising instructions that, when executed by the one or more processors, cause the system to: determine a first likelihood that a record contains at least a given type of sensitive information using a first detection technique comprising: providing one or more inputs to a machine learning model based on the record; and receiving the first likelihood as an output from the machine learning model based on the one or more inputs; determine a second likelihood that the record contains at least the given type of sensitive information using a second detection technique comprising a search of the record; apply a policy to determine whether the record contains sensitive information based on the first likelihood and the second likelihood, wherein the policy states that the second detection technique is used when the first likelihood determined using the first detection technique is above a threshold likelihood; perform one or more actions based on whether the record contains sensitive information; receive a ground truth label indicating whether the record contains sensitive information; and revise the policy based on the ground truth label.
14. The system of claim 13, wherein the search of the record comprises: converting text in the record into a plurality of tokens; searching the plurality of tokens for a first set of search terms; detecting a first term from the first set of search terms in the plurality of tokens; and searching the plurality of tokens for a second set of search terms within a given proximity of the first term.
9. The system of claim 8, wherein the search of the record comprises: converting text in the record into a plurality of tokens; searching the plurality of tokens for a first set of search terms; detecting a first term from the first set of search terms in the plurality of tokens; and searching the plurality of tokens for a second set of search terms within a given proximity of the first term.
15. The system of claim 14, wherein the instructions, when executed by the one or more processors, further cause the system to: generate the first set of search terms by excluding, from a list of first names, one or more exceptional first names and one or more first names longer than a given number of characters; and generate the second set of search terms by excluding, from a list of last names, one or more exceptional last names and one or more last names longer than the given number of characters.
10. The system of claim 9, wherein the instructions, when executed by the one or more processors, further cause the system to: generate the first set of search terms by excluding, from a list of first names, one or more exceptional first names and one or more first names longer than a given number of characters; and generate the second set of search terms by excluding, from a list of last names, one or more exceptional last names and one or more last names longer than the given number of characters.
16. The system of claim 14, wherein: the first set of search terms comprises years in a given range; and the second set of search terms comprises months.
11. The system of claim 9, wherein: the first set of search terms comprises years in a given range; and the second set of search terms comprises months.
17. The system of claim 13, wherein the instructions, when executed by the one or more processors, further cause the system to select the second detection technique from a plurality of detection techniques based on the first likelihood.
12. The system of claim 8, wherein the instructions, when executed by the one or more processors, further cause the system to select the second detection technique from a plurality of detection techniques based on the first likelihood.
18. The system of claim 13, wherein providing the one or more inputs to the machine learning model based on the record comprises providing, as the one or more inputs to the machine learning model, one or more of: tokens determined from the record; or characters from the record.
13. The system of claim 8, wherein providing the one or more inputs to the machine learning model based on the record comprises providing, as the one or more inputs to the machine learning model, one or more of: tokens determined from the record; or characters from the record.
19. The system of claim 13, wherein revising the policy based on the ground truth label comprises adjusting an order of performing the first detection technique and the second detection technique from a first order in which the second detection technique is used based on the first likelihood being above a threshold likelihood to a second order in which the first detection technique is used based on the second likelihood being above the threshold likelihood.
15. The method of claim 1, wherein revising the policy based on the ground truth label comprises adjusting an order of performing the first detection technique and the second detection technique from a first order in which the second detection technique is used based on the first likelihood being above the threshold likelihood to a second order in which the first detection technique is used based on the second likelihood being above the threshold likelihood.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 5-8, 13 and 17-20 are rejected under 35 U.S.C 103 as being unpatentable over Agarwal, US pat. No 20210182607 (IDS Submitted 09/13/2024) in view of Jaiswal, US 8862522 B1.
1. Agarwal discloses a method for detecting sensitive information, (See Agarwal, [0025-0026]; Architecture for Detecting Sensitive Database Information) comprising: determining a first likelihood that a record contains at least a given type of sensitive information using a first detection technique comprising: (See Agarwal, [0049]; [0049] Results of a given risk analysis include a confidence score 245 that indicates a probability that a corresponding data object is a particular type of data object. For example, if a particular one of data objects 115 is encrypted, the scanning process may be performed without decrypting the data object. Accordingly, the scanning process may determine a probability that the particular data object includes sensitive information such as email addresses or credit card information.)
providing one or more inputs to a machine learning model based on the record; (See Agarwal, [0111] [0120] After features 1145, associated with set of data items 1015 have been generated by each of metadata feature extractor 1131, data profile feature extractor 1133, and rules engine feature extractor 1137, computer system 1001 predicts whether set of data items 1015 corresponds to a particular one of output classes 1030. This predicting includes sending features 1145 that are based on the metadata, the data profile, and the comparison of the regexes to neural network classifier 1005. Neural network classifier 1005 generates outputs 1035a-1035n (collectively outputs 1035), based on the received features 1145. Each one of outputs 1035 corresponds to one of output classes 1030. The respective values of outputs 1035 are used to predict if one of output classes 1030 corresponds to set of data items 1015)) and receiving the first likelihood as an output from the machine learning model based on the one or more inputs; (See Agarwal, [0111], [0120] After features 1145, associated with set of data items 1015 have been generated by each of metadata feature extractor 1131, data profile feature extractor 1133, and rules engine feature extractor 1137, computer system 1001 predicts whether set of data items 1015 corresponds to a particular one of output classes 1030. This predicting includes sending features 1145 that are based on the metadata, the data profile, and the comparison of the regexes to neural network classifier 1005. Neural network classifier 1005 generates outputs 1035a-1035n (collectively outputs 1035), based on the received features 1145. Each one of outputs 1035 corresponds to one of output classes 1030. The respective values of outputs 1035 are used to predict if one of output classes 1030 corresponds to set of data items 1015))
determining a second likelihood that the record contains at least the given type of sensitive
information using a second detection technique comprising a search of the record; (See Agarwal, [0111], [0119-0120], [0147]; [0119] As shown, computer system 1001 performs data collector 1127 by comparing a plurality of regular expressions to set of data items 1015. Each of the plurality of regular expressions (also referred to herein as “regexes”) includes a character pattern that is associated with a respective one of output classes 1030. For example, a regex for detecting an email address may search for a pattern of one or more alphanumeric characters followed by the “@” symbol, then followed by more alphanumeric characters, then followed by a suffix string such as “.com.” Additional regexes for email addresses may look for other suffix strings, for example, “.net,” “.org,” “.gov,” etc. Values indicating a degree of matching between a particular data item (e.g., data item 1015b) to a particular regex are sent to rules engine feature extractor 1137. computer system 1001 predicts whether set of data items 1015 corresponds to a particular one of output classes 1030. This predicting includes sending features 1145 that are based on the metadata, the data profile, and the comparison of the regexes to neural network classifier)
Agarwal does not appear to explicitly disclose applying a policy to determine whether the record contains sensitive information based on the first likelihood and the second likelihood; performing one or more actions based on whether the record contains sensitive information; and revising the policy based on a ground truth label indicating whether the record contains sensitive information.
However, Jaiswal discloses applying a policy to determine whether the record contains sensitive information based on the first likelihood and the second likelihood; (See Jaiswal, Fig 2 SECTION; (31) The DLP policy 250 may include one or more profiles 255, 260, 265. Each profile may be used to identify sensitive data. In one embodiment, the DLP policy 250 includes a described content matching (DCM) profile 255. DCM profile 255 defines one or more key words and/or regular expressions to be searched for. For example, DCM profile 255 may define a social security number using a regular expression. Using DCM profile 255, DLP agent 205 determines whether any information included in scanned data match the key words and/or regular expressions. If a match is found, then it may be determined that the data includes sensitive information. See Fig 2 SECTION OF SPEC: (34) One example of a policy violation detector is a machine learning module 225. The ML module 225 includes a ML engine 230 that takes as inputs a MLD profile 265 and unclassified data (e.g., a file 235), and outputs a classification for the data. The ML engine 230 processes the input data using the classification model 275 and the feature set 280. Therefore, the ML module 225 can use the MLD profile 265 to distinguish between sensitive data and non-sensitive data, See Fig 2 SECTION OF SPEC: 30) The DLP agent 205 may include one or more policy violation detectors, each of which may process different DLP policies 250 and/or different profiles 255, 260, 265 within a DLP policy 250 to identify and secure sensitive data. DLP policy 250 may include criteria that may indicate an increased risk of data loss. DLP policy 250 is violated if one or more criteria included in the DLP policy 250 are satisfied. Examples of criteria include user status (e.g., whether the user has access privileges to a file), file location (e.g., whether a file to be copied is stored in a confidential database), file contents (e.g., whether a file includes sensitive information), time (e.g., whether an operation is requested during normal business hours), data loss vector, application attempting the operation, and so on.) performing one or more actions based on whether the record contains sensitive information; (See Jaiswal, See Fig 11 section of spec: (100) Referring to FIG. 11, at block 1105 processing logic receives a request to perform an operation on a document. At block 1110, a ML module analyzes the document using a MLD profile to classify the document. At block 1125, processing logic determines whether the document was classified as sensitive or non-sensitive. If the document was classified as sensitive, the method continues to block 1130, and an action specified by a DLP response rule is performed, and an incident report is generated. This may include preventing the operation, generating an incident response report, etc. If the document was classified as non-sensitive, the method proceeds to block 1135, and the operation is performed.) and revising the policy based on a ground truth label indicating whether the record contains sensitive information. (See Jaiswal, section fig 3 of spec; (45) In one embodiment, MLD profile trainer 325 performs incremental training for an existing MLD profile. With incremental training, MLD profile trainer 325 may add new positive data and/or negative data to the training data set based on incident reports that have been generated since the MLD profile was last trained. The MLD profile trainer 325 may then retrain the MLD profile 365 using the updated training data set (e.g., recompute a feature set 375 and/or a classification model 380). In one embodiment, the MLD profile trainer 325 performs a full retraining of the MLD profile 365 using all of the previous contents of the training data set as well as the newly added content. In another embodiment, the MLD profile trainer 325 performs a partial retraining using only the recently added content. In still another embodiment, incremental training is used to generate an entirely new MLD profile. The new MLD profile may be based on just the new positive and/or negative data or based on a subset of the original positive and/or negative data along with the new positive and/or negative data. For example, MLD profile trainer 325 may generate a new MLD profile with the originally used positive examples of sensitive data and new negative examples of sensitive data to generate the new MLD profile. See section Fig 6 of Spec: (71) Once a threshold number of positive documents and negative documents have been added to the training data set (e.g., 20 documents of each type, 50 documents of each type, etc.), a train profile operation becomes available. In one embodiment, a "train profile" button 630 becomes active when the threshold number of positive documents and negative documents have been added. A user may select the "train profile" button 630 to train the MLD profile (e.g., to generate a feature set and a classification model for the MLD profile).)
Agarwal and Jaiswal are analogous art because they are from the same field of endeavor which is Data Loss Prevention (DLL). It would have been obvious to a person of ordinary skill in the art at the time the invention was made to modify the invention of Agarwal with the teaching of Jaiswal to include the ground truth because it would have allowed the computing device to analyze the modified training data set in response to determining that at least the threshold number of documents have been received. (See Jaiswal, section Summary)
5.The combination of Agarwal and Jaiswal discloses the method of claim 1, further comprising: selecting the second detection technique from a plurality of detection techniques based on the first likelihood. (See Jaiswal, Fig 2 SECTION; (31) The DLP policy 250 may include one or more profiles 255, 260, 265. Each profile may be used to identify sensitive data. In one embodiment, the DLP policy 250 includes a described content matching (DCM) profile 255. DCM profile 255 defines one or more key words and/or regular expressions to be searched for. For example, DCM profile 255 may define a social security number using a regular expression. Using DCM profile 255, DLP agent 205 determines whether any information included in scanned data match the key words and/or regular expressions. If a match is found, then it may be determined that the data includes sensitive information. See Fig 2 SECTION OF SPEC: (34) One example of a policy violation detector is a machine learning module 225. The ML module 225 includes a ML engine 230 that takes as inputs a MLD profile 265 and unclassified data (e.g., a file 235), and outputs a classification for the data. The ML engine 230 processes the input data using the classification model 275 and the feature set 280. Therefore, the ML module 225 can use the MLD profile 265 to distinguish between sensitive data and non-sensitive data, See Fig 2 SECTION OF SPEC: 30) The DLP agent 205 may include one or more policy violation detectors, each of which may process different DLP policies 250 and/or different profiles 255, 260, 265 within a DLP policy 250 to identify and secure sensitive data. DLP policy 250 may include criteria that may indicate an increased risk of data loss. DLP policy 250 is violated if one or more criteria included in the DLP policy 250 are satisfied. Examples of criteria include user status (e.g., whether the user has access privileges to a file), file location (e.g., whether a file to be copied is stored in a confidential database), file contents (e.g., whether a file includes sensitive information), time (e.g., whether an operation is requested during normal business hours), data loss vector, application attempting the operation, and so on.) Agarwal and Jaiswal are analogous art because they are from the same field of endeavor which is Data Loss Prevention (DLL). It would have been obvious to a person of ordinary skill in the art at the time the invention was made to modify the invention of Agarwal with the teaching of Jaiswal to include the ground truth because it would have allowed the computing device to analyze the modified training data set in response to determining that at least the threshold number of documents have been received. (See Jaiswal, section Summary)
6. The combination of Agarwal and Jaiswal discloses the method of Claim 1, wherein providing the one or more inputs to the machine learning model based on the record comprises providing, as the one or more inputs to the machine learning model, one or more of:
tokens determined from the record; or characters from the record. (See Agarwal, [0111], [0119-0120], [0147]; [0119] As shown, computer system 1001 performs data collector 1127 by comparing a plurality of regular expressions to set of data items 1015. Each of the plurality of regular expressions (also referred to herein as “regexes”) includes a character pattern that is associated with a respective one of output classes 1030. For example, a regex for detecting an email address may search for a pattern of one or more alphanumeric characters followed by the “@” symbol, then followed by more alphanumeric characters, then followed by a suffix string such as “.com.” Additional regexes for email addresses may look for other suffix strings, for example, “.net,” “.org,” “.gov,” etc. Values indicating a degree of matching between a particular data item (e.g., data item 1015b) to a particular regex are sent to rules engine feature extractor 1137. computer system 1001 predicts whether set of data items 1015 corresponds to a particular one of output classes 1030. This predicting includes sending features 1145 that are based on the metadata, the data profile, and the comparison of the regexes to neural network classifier) Agarwal and Jaiswal are analogous art because they are from the same field of endeavor which is Data Loss Prevention (DLL). It would have been obvious to a person of ordinary skill in the art at the time the invention was made to modify the invention of Agarwal with the teaching of Jaiswal to include the ground truth because it would have allowed the computing device to analyze the modified training data set in response to determining that at least the threshold number of documents have been received. (See Jaiswal, section Summary)
7. The combination of Agarwal and Jaiswal discloses the method of Claim 1, wherein performing the one or more actions based on whether the record contains sensitive information comprises one or more of:
removing one or more sensitive information items from the record; encrypting one or more sensitive information items in the record; generating a notification indicating whether the record contains sensitive information; or modifying code relating to generation of the record. (See Agarwal, [0033], [0043], [0060; encrypt of data)
8. The combination of Agarwal and Jaiswal discloses the method of Claim 1, wherein revising the policy based on the ground truth label comprises adjusting an order of performing the first detection technique and the second detection technique from a first order in which the second detection technique is used based on the first likelihood being above a threshold likelihood to a second order in which the first detection technique is used based on the second likelihood being above the threshold likelihood. (See Jaiswal, Fig 3 section of spec: (53) Once the feature extractor 335 generates the feature set 375 and the model generator 330 generates the classification model 380, a MLD profile 365 is complete. The MLD profile 365 may include the feature set 375, the classification model 380 and/or the training data set 370. The MLD profile 365 may also include user defined settings. In one embodiment, the user defined settings include a sensitivity threshold (also referred to as a confidence level threshold). The sensitivity threshold may be set to, for example, 75%, 90%, etc. When an ML engine uses the MLD profile 365 to classify a document as sensitive or not sensitive, the ML engine may assign a confidence value to the classification. If the confidence value for the document is 100%, then it is more likely that the decision that the document is sensitive (or not sensitive) is accurate than if the confidence value is 50%, for example. If the confidence value is less than the sensitivity threshold, then an incident may not be generated even though a document was classified as a sensitive document. This feature can help a user to further control and reduce false positives and/or false negatives. If an ML engine is trying to classify a document of a type that the training has never seen, it has a very low confidence of the document being positive and/or negative. The sensitivity threshold can be used to reduce occurrences of false positive in such cases. In one embodiment, the MLD profile trainer 325 automatically selects a sensitivity threshold for the MLD profile 365 based on the training.)
Agarwal and Jaiswal are analogous art because they are from the same field of endeavor which is Data Loss Prevention (DLL). It would have been obvious to a person of ordinary skill in the art at the time the invention was made to modify the invention of Agarwal with the teaching of Jaiswal to include the ground truth because it would have allowed the computing device to analyze the modified training data set in response to determining that at least the threshold number of documents have been received. (See Jaiswal, section Summary)
13. As to claim 13, the claim is rejected under the same rationale as claim 1. See the rejection of claim 1 above.
one or more processors; (See Agarwal, [0100- 0101]) and a memory comprising instructions that, when executed by the one or more processors, (See Agarwal, [0100- 0101]) cause the system to:
17. As to claim 17, the claim is rejected under the same rationale as claim 5. See the rejection of claim 5 above.
18. As to claim 18, the claim is rejected under the same rationale as claim 6. See the rejection of claim 6 above.
19. As to claim 19, the claim is rejected under the same rationale as claim 8. See the rejection of claim 8 above.
20. The system of Claim 13, wherein the policy states that the first detection technique is used when the second likelihood determined using the second detection technique is above a threshold likelihood. (See Jaiswal, Fig 3 section of spec: (53) Once the feature extractor 335 generates the feature set 375 and the model generator 330 generates the classification model 380, a MLD profile 365 is complete. The MLD profile 365 may include the feature set 375, the classification model 380 and/or the training data set 370. The MLD profile 365 may also include user defined settings. In one embodiment, the user defined settings include a sensitivity threshold (also referred to as a confidence level threshold). The sensitivity threshold may be set to, for example, 75%, 90%, etc. When an ML engine uses the MLD profile 365 to classify a document as sensitive or not sensitive, the ML engine may assign a confidence value to the classification. If the confidence value for the document is 100%, then it is more likely that the decision that the document is sensitive (or not sensitive) is accurate than if the confidence value is 50%, for example. If the confidence value is less than the sensitivity threshold, then an incident may not be generated even though a document was classified as a sensitive document. This feature can help a user to further control and reduce false positives and/or false negatives. If an ML engine is trying to classify a document of a type that the training has never seen, it has a very low confidence of the document being positive and/or negative. The sensitivity threshold can be used to reduce occurrences of false positive in such cases. In one embodiment, the MLD profile trainer 325 automatically selects a sensitivity threshold for the MLD profile 365 based on the training.)
Claims 2, 14 are rejected under 35 U.S.C 103 as being unpatentable over Agarwal, US pat. No 20210182607 (IDS Submitted (IDS Submitted 09/13/2024)) in view of Jaiswal, US 8862522 B1 in further view of Hendrey, US pat. No 10198530 (IDS Submitted 09/13/2024).
2.The combination of Agarwal and Jaiswal does not appear to explicitly disclose the method of Claim 1, wherein the search of the record comprises:
converting text in the record into a plurality of tokens;
searching the plurality of tokens for a first set of search terms;
detecting a first term from the first set of search terms in the plurality of tokens; and
searching the plurality of tokens for a second set of search terms within a given proximity
of the first term.
However, Hendrey discloses converting text in the record into a plurality of tokens; (See Hendrey, fig 1A) searching the plurality of tokens for a first set of search terms; (See Hendrey, fig 1A )
detecting a first term from the first set of search terms in the plurality of tokens; (See Hendrey, fig 7A and col 2 to col 3) and searching the plurality of tokens for a second set of search terms within a given proximity of the first term. (See Hendrey, fig 7A and col 2 to col 3)
Agarwal and Jaiswal and Hendrey are analogous art because they are from the same field of endeavor which is data classification. It would have been obvious to a person ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Agarwal and Jaiswal with the teaching of Hendrey to include the token because it would have allowed for performing information retrieval involving large amounts of documents. (Hendrey, Col 1, lines 15-20)
14. As to claim 14, the claim is rejected under the same rationale as claim 2. See the rejection of claim 2 above.
Claims 4 and 16 are rejected under 35 U.S.C 103 as being unpatentable over Agarwal, US pat. No 20210182607 ( (IDS Submitted 09/13/2024)) in view of Jaiswal, US 8862522 B1 in further view of Wang, US 11321529 B2.
4. The combination of Agarwal and Jaiswal does not appear to explicitly disclose the method of Claim 2, wherein: the first set of search terms comprises years in a given range; and the second set of search terms comprises months.
However, Wang discloses the first set of search terms comprises years in a given range; (See Wang, col 4, lines 3-15, 64-66; (46) A validation operation 312 validates labeling of all tokens as dates, including the tokens that are labeled as part of date ranges. For example, the validation operation 3121 may compare the start date of the range (here <16 February 2016>) and the end date of the range (here <29 May 2016>) to ensure that the end date is higher than the start date. (38) For example, the date range chucking module 118 interprets a pair of dates “1/4/2018˜2/3/2018” as MDY˜MDY.
(39) Subsequently, a validation module 120 validates labeling of all tokens as dates, including the tokens that are labeled as part of date ranges. For example, one rule used by the validation module 120 may be that the start dates are earlier than the end date in a date range. In one implementation, the validation module 120 may be customized for addition of other validation rules thereto, such as a rule that requires a date range to be within one year, etc. The validation module 120 generates the date range output 150 with substantially high level of precision and recall rate. The date extractor system 100 may also be applied to languages other than English by changing the corresponding patterns and vocabulary. ) and the second set of search terms comprises months. (See Wang, col 4, lines 3-15, 64-66; (46) A validation operation 312 validates labeling of all tokens as dates, including the tokens that are labeled as part of date ranges. For example, the validation operation 3121 may compare the start date of the range (here <16 February 2016>) and the end date of the range (here <29 May 2016>) to ensure that the end date is higher than the start date. )
Agarwal and Jaiswal and Wang are analogous art because they are from the same field of endeavor which is data classification. It would have been obvious to a person ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Agarwal and Jaiswal with the teaching of Wang to include the date range because it would have allowed to enumerate pattern with better precision.
16. As to claim 16, the claim is rejected under the same rationale as claim 4. See the rejection of claim 4 above.
Allowable Subject Matter
Claims 3 and 15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
BORUP, US 20180293400 A1, title “SYSTEM TO PREVENT EXPORT OF SENSITIVE DATA.”
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSNEL JEUDY whose telephone number is (571)270-7476. The examiner can normally be reached M-F 10:00-8:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Arani T Taghi can be reached at (571)272-3787. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Date: 08/18/2026
/JOSNEL JEUDY/ Primary Examiner, Art Unit 2438