DETAILED CORRESPONDENCE
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Acknowledgments
This Action is in response to the patent application filed on February 22, 2024. Claims 1-20 are currently pending and have been fully examined.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
Claims 1-10 are directed to a method (process.) Claims 11-20 are directed to a non-transitory computer-readable medium (product). Therefore, these claims fall within the four statutory categories of invention.
Claims 1-20 are directed to the abstract idea of classifying sensitive data, as explained in detail below. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Analysis
In the following analysis, bolded text indicates abstract idea and the rest of the text indicates additional elements. Independent claims 1 and 11, recite:
receiving an input comprising data in any of a plurality of formats;
processing the input to determine whether or not the data includes sensitive data;
and
responsive to the input including sensitive data, performing steps of:
processing the input to classify the input into a category of a plurality of categories; and
providing an indication of the category of the plurality of categories.
In addition, claim 11 recites:
a non-transitory computer-readable medium comprising instructions that, when executed, cause one or more processors...
Specifically claims 1 and 11 recite the abstract idea of classifying sensitive data. Therefore the claims recite a fundamental economic principle or practice grouped within the “certain methods of organizing human activity” grouping of abstract ideas in prong one of step 2A of the Alice/Mayo test (See MPEP 2106) because the claims involve a series of steps for receiving data, determining whether the data includes sensitive data and classifying the data into a category. Accordingly, the claims recite an abstract idea (See pages 7, 10, Alice Corporation Pty. Ltd. v. CLS Bank International, et al., US Supreme Court, No. 13-298, June 19, 2014; MPEP 2106).
This judicial exception is not integrated into a practical application because, when analyzed under prong two of step 2A of the Alice/Mayo test, the additional elements of a non-transitory computer-readable medium, and one or more processors, merely use one or more computers as tool to perform the abstract idea. The use of a non-transitory computer-readable medium, and one or more processors does not integrate the abstract idea into a practical application because it requires no more than one or more computing devices performing functions that correspond to acts required to carry out the abstract idea. The additional elements do not involve improvements to the functioning of a computer, or to any other technology or technical field (MPEP 2106.05(a)), and the claims do not apply or use the abstract idea in some other meaningful way beyond generally linking the use of the abstract idea to a particular technological environment, such that the claim as a whole is more than a drafting effort designed to monopolize the exception (MPEP 2106.05(e) and Vanda Memo). Therefore, the claims do not, for example, purport to improve the functioning of a computer. Nor do they effect an improvement in any other technology or technical field. Accordingly, the additional elements do not impose any meaningful limits on practicing the abstract idea, and the claims are directed to an abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when analyzed under step 2B of the Alice/Mayo test (See MPEP 2106), the additional elements of a non-transitory computer-readable medium and one or more processors, amount to no more than using computing devices or processors to automate and/or implement the abstract idea.
As discussed above, taking the claim elements separately, these additional elements perform the steps or functions that correspond to the actions required to perform the abstract idea. Viewed as a whole, the combination of elements recited in the claims merely recite the abstract idea.
Dependent claims 2 and 12, recite:
responsive to the input including non-sensitive data, providing an indication the data is non-sensitive, thereby either allowing the data in transit or not marking the data at rest,
which further describes the abstract idea. Ther claims do not include any additional elements to be analyzed under prong two of step 2A and step 2B of the Alice/Mayo test (See MPEP 2106.)
Dependent claims 3 and 13, recite:
responsive to the input including sensitive data, providing an indication of the category of the plurality of categories and a sub-category associated with the category,
which further describes the abstract idea. Ther claims do not include any additional elements to be analyzed under prong two of step 2A and step 2B of the Alice/Mayo test (See MPEP 2106.)
Dependent claims 4 and 14, recite:
wherein the plurality of formats include text formats, image formats, audio formats, video formats, source code, and a combination thereof,
which further describes the abstract idea. Ther claims do not include any additional elements to be analyzed under prong two of step 2A and step 2B of the Alice/Mayo test (See MPEP 2106.)
Dependent claims 5 and 15, recite:
wherein the processing the input to determine whether or not the data includes sensitive data utilizes (1) a Large Language Model (LLM) and embeddings and (2) a machine learning model configured for classification.
The claims further describe the abstract idea recited in claims 1 and 11 and recite utilizing a Large Language Model and a machine learning model that are mathematical concepts. The claims do not include any additional elements to be analyzed under prong two of step 2A and step 2B of the Alice/Mayo test (See MPEP 2106.)
Dependent claims 6 and 16, recite:
wherein the processing the input to classify the input into the category utilizes (1) a Large Language Model (LLM) and (2) a zero-shot classifier.
The claims further describe the abstract idea recited in claims 1 and 11 and recite utilizing a Large Language Model and a zero-shot classifier that are mathematical concepts. The claims do not include any additional elements to be analyzed under prong two of step 2A and step 2B of the Alice/Mayo test (See MPEP 2106.)
Dependent claims 7 and 17, recite:
wherein the processing the input to determine whether or not the data includes sensitive data and the processing the input to classify the input into the category both utilize one or more machine learning models that were trained based on a set of training documents with labels.
The claims further describe the abstract idea recited in claims 1 and 11 and recite utilizing trained machine learning models that are mathematical concepts. The claims do not include any additional elements to be analyzed under prong two of step 2A and step 2B of the Alice/Mayo test (See MPEP 2106.)
Dependent claims 8 and 18, recite:
prior to training the one or more machine learning models with the set of training documents with labels, identifying any mislabeled documents therein by performing Optical Character Recognition (OCR) and checking if associated keywords are present.
The judicial exception is not integrated into a practical application because, when analyzed under prong two of step 2A of the Alice/Mayo test (See MPEP 2106), the additional element of performing Optical Character Recognition, merely use one or more computers as tool to perform the abstract idea. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when analyzed under step 2B of the Alice/Mayo test (See MPEP 2106), the additional element amount to no more than using computing devices or processors to automate and/or implement the abstract idea.
Dependent claims 9 and 19, recite:
prior to training the one or more machine learning models with the set of training documents with labels, filtering out an images in the set of training documents with labels based on file size,
which further describes the abstract idea. Ther claims do not include any additional elements to be analyzed under prong two of step 2A and step 2B of the Alice/Mayo test (See MPEP 2106.)
Dependent claims 10 and 20, recite:
prior to training the one or more machine learning models with the set of training documents with labels, grouping images in the set of training documents with subtle differences based on a hash of the images,
The claims further describe the abstract idea recited in claims 7 and 17 and recite utilizing image hashes for grouping images, that is a mathematical concept. The claims do not include any additional elements to be analyzed under prong two of step 2A and step 2B of the Alice/Mayo test (See MPEP 2106.)
Claim Rejections - 35 USC § 102
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-5, 7-8, 11-15 and 17-18 are rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by Zhang (US Patent Publication No. 2024/0386897.)
With respect to claims 1 and 11, Zhang, which like the present Application is directed to classifying sensitive data using machine learning, teaches:
receiving an input comprising data in any of a plurality of formats; (receiving multi-modal input: [0019]-[0020]
processing the input to determine whether or not the data includes sensitive data; (AI model classification system determines sensitive or non-sensitive data: [0027])
responsive to the input including sensitive data, performing steps of:
processing the input to classify the input into a category of a plurality of categories; (AI model classifies sensitive-data: [0027]-[0028])
providing an indication of the category of the plurality of categories. (output the classification: [0028], [0032])
In addition, with respect to claim 11, Zhang teaches:
a non-transitory computer-readable medium comprising instructions executed, cause one or more processors… (system includes storage device or medium capable of storing data and/or computer-readable instructions and one or more physical processors: [0071])
With respect to claims 2 and 12, Zhang teaches the limitations of claims 1 and 11.
responsive to the input including non-sensitive data, providing an indication the data is non-sensitive, thereby either allowing the data in transit or not marking the data at rest. (determine non-sensitive data: [0027], [0076])
The examiner notes that the claim recitation “thereby either allowing the data in transit or not marking the data at rest,” indicates an intended result of indicating non-sensitive data and therefore does not further limit the scope of the claim, as no specific function has been positively recited.
With respect to claims 3 and 13, Zhang teaches the limitations of claims 1 and 11.
Moreover, Zhang teaches:
responsive to the input including sensitive data, providing an indication of the category of the plurality of categories and a sub-category associated with the category. (the system uses hierarchical clustering of input data from a first layer to subclusters (i.e., sub-categories): [0062]-[0063])
With respect to claims 4 and 14, Zhang teaches the limitations of claims 1 and 11.
Moreover, Zhang teaches:
wherein the plurality of formats include text formats, image formats, audio formats, video formats, source code, and a combination thereof. (data can be of various formats video, audio, image, text: [0019])
With respect to claims 5 and 15, Zhang teaches the limitations of claims 1 and 11.
Moreover, Zhang teaches:
wherein the processing the input to determine whether or not the data includes sensitive data utilizes (1) a Large Language Model (LLM) and embeddings and (2) a machine learning model configured for classification. (the analysis system uses reinforcement learning (i.e., machine learning): [0038], [0040], the system may have multiple modular components such as natural language understanding (i.e. LLM): [0047])
With respect to claims 7 and 17, Zhang teaches the limitations of claims 1 and 11.
Moreover, Zhang teaches:
wherein the processing the input to determine whether or not the data includes sensitive data and the processing the input to classify the input into the category both utilize one or more machine learning models that were trained based on a set of training documents with labels. (train the model using labeled feature input: [0073]-[0074])
The examiner notes that the claim recitation: “…one or more machine learning models that were trained based on…” indicates not positively recited language and therefore does not further limit the scope of the claim, because the “training” function is not positively recited.
With respect to claims 8 and 18, Zhang teaches the limitations of claims 7 and 17.
prior to training the one or more machine learning models with the set of training documents with labels, identifying any mislabeled documents therein by performing Optical Character Recognition (OCR) and checking if associated keywords are present. (preprocessing the input using OCR: [0021], [0024], [0046])
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, as applied to claims 1 and 11 above, in view of Kuo ( US Patent Publication No. 2025/0384650.)
With respect to claims 6 and 16, Zhang teaches the limitations of claims 1 and 11.
Zhang does not explicitly teach; however, Kuo, which like the present Application is directed to classifying data using language models teaches:
wherein the processing the input to classify the input into the category utilizes (1) a Large Language Model (LLM) and (2) a zero-shot classifier. (language models using zero-shot classification: [0044]-[0047]
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the zero-shot classification as taught by Kuo, into the system of Zhang for multi-modal classification, in order to detect data categories that may not be present in training data. (Kuo: [0047])
Claims 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, as applied to claims 7 and 17 above, in view of Daniels ( US Patent No. 11,010,640.)
With respect to claims 9 and 19, Zhang teaches the limitations of claims 7 and 17.
Zhang does not explicitly teach; however, Daniels, which like the present Application is directed to processing training data sets teaches:
prior to training the one or more machine learning models with the set of training documents with labels, filtering out an images in the set of training documents with labels based on file size. (filter labeled training data based on image size: (Col. 11 l. 55-Col. 12 l. 59, claim 11)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the filtering of training data, as taught by Daniels, into the system of Zhang for multi-modal classification, in order to improve quality of training data. (Daniels: Abstract, Col. 4 ll. 5-15)
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, as applied to claims 7 and 17 above, in view of Liu ( US Patent Publication No. 2022/0067083.)
With respect to claims 10 and 20, Zhang teaches the limitations of claims 7 and 17.
Zhang does not explicitly teach; however, Liu, which like the present Application is directed to processing image data for training machine learning models, teaches:
prior to training the one or more machine learning models with the set of training documents with labels, grouping images in the set of training documents with subtle differences based on a hash of the images. (grouping images based on hash ranges: [0028]-[0030] prior to training a machine learning model: [0044]-[0046], claim 20)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the filtering and sorting of images, as taught by Liu, into the system of Zhang for multi-modal classification, in order to increase accuracy of the classification process . (Liu: Abstract, [0020])
Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Cannon (US 2022/0164472 ) teaches classification of sensitive data, Goodsitt (US 2021/0097343 ) teaches training machine learning models for classifying sensitive data, Pushkin (US 11,861,039 ) teaches a hierarchical model for identifying sensitive content.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIMA ASGARI whose telephone number is (571)272-2037. The examiner can normally be reached M-F 9am-6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Patrick McAtee can be reached at (571)272-7575. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SIMA ASGARI/Examiner, Art Unit 3698
/PATRICK MCATEE/Supervisory Patent Examiner, Art Unit 3698