DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the amendment filed on 06/29/2026. Claims 1-26 are pending in the case. All claims are examined and rejected accordingly.
Applicant Response
3. In Applicant’s response dated 06/29/2026, Applicant argued against all objections
Response to Arguments
4. Applicant’s arguments have been considered and are persuasive. The rejections under 35 U.S.C. 103 is withdrawn because the prior art applied (Segev) is also not applicable prior art to the instant application at least because the applications were owned by the same entity at least as of the effective filing date of the instant application and a second non-Final office action is issued as shown below.
Claim Rejections - 35 USC § 112
5. The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
6. Claim 1-26 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites: a first prompt indicate “.. characteristics of an entity” and … wherein the fine-tuned language model is queried using at least one second prompt indicating the plurality of second classifications and data indicating a set of characteristics of an entity,
It is unclear whether the second entity is the same as the first or a different one. For purpose of examination, Examiner will interpret the entities are the same and suggest : “… a first prompt indicate “.. characteristics of an entity” and … wherein the fine-tuned language model is queried using at least one second prompt indicating the plurality of second classifications and data indicating the set of characteristics of an entity. All the independent claims 18, 15 and 21 are rejected under the same rationale.
Claim 1 recites: “… labeling training data including the plurality of second classifications based on the sensitivity output by the language model for each of the plurality of second classifications in order to create a labeled training data set …”. The language model lacks antecedent bases and It is not clear weather the language model is referring to the pre-tuned language model with no output and the fine tuned language model with the sensitivity output. For purpose of examination, Examiner will interpret : “… based on the sensitivity output by the fine-tuned language model for each of the plurality of second classifications in order to create a labeled training data set …”. All the independent claims 18, 15 and 21 are rejected under the same rationale.
Claim 3, 10, 16 and 22
Claim 3 recites : “applying a cybersecurity policy based on the detected sensitivity for the third resource, wherein the cybersecurity policy defines at least one of: at least one permissible condition, and at least one forbidden condition.”
Examiner notes that Nested “at least one of … and …” is ambiguous and it can not be determined weather the policy must define one or the other or both. For purpose of examination, Examiner will interpret : “…at least one of: at least one permissible condition, or at least one forbidden condition.” The independent claims 10, 16 and 22 are rejected under the same rationale.
Claim Rejections - 35 USC § 101
7. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
8. Claims 1-26 are rejected under 35 U.S.C. 101 because the claimed invention is directed towards an abstract idea, without significantly more.
Step 1
According to the first part of the analysis, in the instant case, claim is directed to a computer implemented method, which is a process and falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Regarding Claim 1
At step 2A, prong 1, Does the claim recite a judicial exception?
Claim 1 recites the steps of:
fine-tuning a language model by iteratively applying the language model to a plurality of first prompts and adjusting weights of the language model in order to produce a fine-tuned language model (This step relies training a model which involve mathematical operations which falls into the “Mathematical concepts grouping of abstract ideas.),
wherein the plurality of first prompts indicate a plurality of first classifications for a plurality of first resources and characteristics of an entity (This step relies organizing/ evaluating information which can be performed by mentally which falls into the “Mathematical process grouping of abstract ideas.),
querying the fine-tuned language model with respect to a plurality of second classifications of a plurality of second resources in order to obtain a set of language model outputs, … (This step relies training a model which involve evaluation of a data by applying rules to input data to obtain outputs which falls into the “Mathematical process grouping of abstract ideas.),
labeling training data including the plurality of second classifications based on the sensitivity output by the language model for each of the plurality of second classifications in order to create a labeled training data set (This step relies on organizing information which falls into the “mental process” grouping of abstract ideas.);
training a sensitivity detection machine learning model via supervised machine learning using the labeled training data set in order to create a trained sensitivity detection machine learning model, wherein the trained sensitivity detection machine learning model is configured to output sensitivities (This step relies training a model which involve mathematical operations which falls into the “Mathematical concepts grouping of abstract ideas.),
The claim involved data manipulation, evaluation, and classification, which falls in to the mathematical concepts and mathematical process grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
Step 2A prong 2: Does the claim recite additional elements? Do those additional elements, individually and in combination, integrate the judicial exception into a practical application?
Further, the claim does not recite any additional element which could integrate this abstract idea into a practical application, because the additional elements recited of consist of:
“… a fine-tuned language model (generic ML training), … prompting/ querying (generic data input output),
… a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system …” (claim 8), “processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to…”, (claim 21) (Generic computer components on which to implement the math abstract idea (see MPEP 2106.05(f));
The additional elements are recited at a high level of generality and do not amount to significantly more than the abstract idea (MPEP 2106.05(f)). The claim use a computer to perform a math and does not improve the function of the computer or other technology. Accordingly, the claim does not integrate the abstract idea into practical application.
Thus, the claim is directed towards the abstract idea.
Step 2B: Do the additional elements, considered individually and in combination, amount to significantly more than the judicial exception?
No, as shown above with respect to integration of the abstract idea into a practical application, the additional element of “… a fine-tuned language model (generic ML training), … prompting/ querying (generic data input output),
… a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system …” (claim 8), “processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to…”, (claim 21) (Generic computer components on which to implement the math abstract idea (see MPEP 2106.05(f));
The additional elements, alone and in combination, fail to integrate the abstract idea into a practical application or add “significantly more.” Thus, the claims are not patent eligible. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. Neither can insignificant extra-solution activity. All of these additional elements as generically claimed are thus considered well-understood, routine, and conventional. Therefore, these limitations, taken alone or in combination, do not integrate the abstract idea into a practical application or recite significantly more that the abstract idea.
Thus, these independent claims are not patent eligible.
The dependent claims respectively recite a judicial exception in limitations of: “detecting a sensitivity for a third resource by applying the trained sensitivity detection machine learning model to inputs indicating at least a third classification of the third resource.” (claims 2/9), “applying a cybersecurity policy based on the detected sensitivity for the third resource, wherein the cybersecurity policy defines at least one of: at least one permissible condition, and at least one forbidden condition.” (claims 3/10/16/22), “detecting a policy violation based on the application of the cybersecurity policy; and performing at least one remediation action based on the detected policy violation.” (claims 4/11/17/23), “classifying the third resource to determine the third classification by applying a classifier machine learning model to data of the third resource.” (claims 5/12/18/24), “wherein fine-tuning the language model further includes, at each iteration: comparing outputs generated by the language model at the iteration to a set of reference sensitivities for the plurality of first resources, wherein the weights of the language model are adjusted based on the comparison.” (claims 6/13/19/25), “wherein each first prompt of the plurality of first prompts further indicates a predetermined set of sensitivity levels, wherein the fine-tuned language model is trained to output sensitivity levels among the predetermined set of sensitivity levels.” (claims 7/14/20/26).
These additional limitations (in claims 2-7, 9-14, 16-20 and 22-26) also constitute concepts performed Mathematical concept or mathematical operation groupings of abstract ideas.
This judicial exception is not integrated into a practical application. Additional elements “computer readable medium comprising: computer program code (in claims 2-7, 9-14, 16-20 and 22-26), all amount to no more than adding insignificant extra-solution activity/specifications related to data gathering, data input, or data transmittal. These additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The dependent claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of non-transitory computer readable medium comprising: computer program code are again insignificant extra-solution activity steps that cannot provide an inventive concept. All of these additional elements as generically claimed are considered well-understood, routine, and conventional.
Therefore, these limitations, taken alone or in combination, do not integrate the abstract idea into a practical application or recite significantly more that the abstract idea. Thus, all of the dependent claims are also not patent eligible.
Examiner Comments
7. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 103
9. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
10. Claims 1-26 are rejected under 35 U.S.C. 103 as being unpatentable over Smith (Pub. No. US 20240160900 A1, Pub. Date 2024-05-16) in view of Williamson (Pub. No. US 20200410116 A1, 2020-12-31).
Smith teaches a method for sensitivity detection training (see Smith: Fig.1(a), [0009], “semi-automatically generate labels for data based on implementation of a clustering or language model prompting technique”), comprising:
fine-tuning a language model by iteratively applying the language model to a plurality of first prompts (see Smith: Fig.8, [0242]-[0248], “Foundation (Large Language) Model Fine-tuning: Fine-tune modern foundation models such as GPT-3.5, GPT-4, RoBERTa, CLIP, and more, by programmatically creating large, task- and domain-specific training data sets; [0243] Foundation Model Warm Start: Use foundation models with zero- and few-shot learning to auto-label training data for the purpose of training high quality supervised ML models; and [0244] Foundation Model Prompt Builder: Develop, evaluate, and combine prompts to tune and correct the output of foundation models to precisely label datasets for the purpose of training high quality supervised ML models.”), and adjusting weights of the language model in order to produce a fine-tuned language model (see Smith: Fig.3, [0087], “it is checked whether more models are to be trained and, if so, execution continues with S340; otherwise, execution terminates. When multiple models are to be used, training may continue until each such model has been trained.”), wherein the plurality of first prompts indicate a plurality of first classifications for a plurality of first resources and characteristics of an entity (see Smith: Fig.8, [0271], “prompts may be considered a type of labeling function. They fit into the disclosed data-centric workflow, which supports combining, tuning, and integrating their outputs using theoretically grounded modeling and (in some cases) human-in-the-loop error analysis techniques. With the combination of the Prompt Builder feature and functionality, and the data-centric AI development loop, users can identify specific Foundation Model errors, correct them, and refine their output by pruning, modifying, and developing new model prompts.”)
querying the fine-tuned language model with respect to a plurality of second classifications of a plurality of second resources in order to obtain a set of language model outputs (see Smith: Fig.8, [0260], “User-defined lookup mapping, where the user specifies (possibly multiple) raw model outputs that should be mapped to a given label; [0263] User-defined custom code mapping, where the user defines code that takes the raw model output as input and outputs the label it should map to (this may be implemented with an interactive code editor”), wherein the fine-tuned language model is queried using at least one second prompt indicating the plurality of second classifications and data indicating a set of characteristics of an entity (see Smith: Fig.8, [0265], “With the Foundation Model Prompt Builder feature/functionality, a user can more efficiently query foundation models with specific questions or prompts to extract domain-relevant knowledge. For example, to use the canonical example of a spam classifier, one might auto-label some types of spam with Warm Start, then use the Prompt Builder to create a more targeted prompt labeling function (LF) asking “Is this email asking for my password?” In doing so, a user conveys domain knowledge about a particular type of spam (phishing)..”)
Smith does not teach the system wherein:
labeling training data including the plurality of second classifications based on the sensitivity output by the language model for each of the plurality of second classifications in order to create a labeled training data set
training a sensitivity detection machine learning model via supervised machine learning using the labeled training data set in order to create a trained sensitivity detection machine learning model, wherein the trained sensitivity detection machine learning model is configured to output sensitivities.
However, Williamson teaches the system wherein:
labeling training data including the plurality of second classifications based on the sensitivity output by the language model for each of the plurality of second classifications in order to create a labeled training data set (see Williamson: Fig.3, [0081], “The data context tuner 308 automatically selects contextual data for a data portion, which may then be used by the contextual analyzer 210 to determine if that data portion is sensitive data. The data context tuner 308 may receive multiple sets of labeled training data. Each set includes data portions that are labeled as sensitive data and those that are labeled as not sensitive data. Each set of training data additionally includes different combinations of contextual data”)
training a sensitivity detection machine learning model via supervised machine learning using the labeled training data set in order to create a trained sensitivity detection machine learning model (see Williamson: Fig.2, [0080], “Various machine learning models may be used, depending on effectiveness, such as a support vector machine, a naive Bayes classifier, or a neural network. Multiple machine learning models may be trained on the training data and the one that produces the best accuracy for the selected training data may be selected. The trained machine learning model is tested against a validation data set, and if the accuracy exceeds a certain percentage, then the machine learning model is selected by the machine learning trainer 306 for use with the deep learning classifier 212.”), wherein the trained sensitivity detection machine learning model is configured to output sensitivities (see Williamson: Fig.2, [0086], “For each of the set of classification rules, the sensitive data scanner 104 applies 406 the classification rule to the data portion to obtain an output representative of whether the data portion is sensitive.”)
Because both Smith and Williamson are in the same/similar field of endeavor of machine learning based classification systems, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify Smith’s language-based classification system to include the sensitivity detection model and the labeling training data and training a sensitivity detection machine learning model as taught by Williamson. Smith teaches applying a language model to input data to generate classification outputs but does not explicitly disclose a model for detecting a sensitive data and adjusting model parameter through refinement. Williamson teaches classifying data as sensitive using a machine leaning based classifier and further teaches refining models using iterative training process. One would have been motivated to make such a combination in order to improve the tedious procedure of manually discovering sensitive data in electronic systems to scale and compete with the pace of data growth and distribution. (see Williamson: [0003])
Regarding Claim 2,
Smith and Williamson and teaches all the limitations of claim 1. Smith further teaches a system comprising:
detecting a sensitivity for a third resource by applying the trained sensitivity detection machine learning model to inputs indicating at least a third classification of the third resource (see Williamson: Fig.3, [0081], “The data context tuner 308 automatically selects contextual data for a data portion, which may then be used by the contextual analyzer 210 to determine if that data portion is sensitive data. The data context tuner 308 may receive multiple sets of labeled training data. Each set includes data portions that are labeled as sensitive data and those that are labeled as not sensitive data. Each set of training data additionally includes different combinations of contextual data”)
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify Smith’s language-based classification system to include the sensitivity detection model and the labeling training data and training a sensitivity detection machine learning model as taught by Williamson. One would have been motivated to make such a combination in order to improve the tedious procedure of manually discovering sensitive data in electronic systems to scale and compete with the pace of data growth and distribution. (see Williamson: [0003])
Regarding Claim 3,
Smith and Williamson and teaches all the limitations of claim 2. Williamson further teaches a system comprising:
applying a cybersecurity policy based on the detected sensitivity for the third resource, wherein the cybersecurity policy defines at least one of: at least one permissible condition, and at least one forbidden condition (see Williamson: Fig.1, [0038], “ the data classifier 108 further determines the security level of data portions determined to be sensitive (or likely to be sensitive and having a confidence value beyond a certain threshold). The security level indicates how well the sensitive data portion is currently protected by various security features. The data classifier 108 may be able to detect the number and type of security feature(s) applied to the data portion.”)
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify Smith’s language-based classification system to include the sensitivity detection model and the labeling training data and training a sensitivity detection machine learning model as taught by Williamson. One would have been motivated to make such a combination in order to improve the tedious procedure of manually discovering sensitive data in electronic systems to scale and compete with the pace of data growth and distribution. (see: [0003])
Regarding Claim 4,
Smith and Williamson and teaches all the limitations of claim 3. Williamson further comprising:
detecting a policy based on the application of the cybersecurity policy (see Williamson: Fig.2, [0082], “data that is potentially sensitive in an organization can be detected continuously over time, with the detection becoming more accurate over time. The sensitive data that is found is reported to the user in a user interface that indicates the sources, types, confidence value, locations, and other characteristics of the detected sensitive data. The system may also automatically apply additional security features to data that is detected to be sensitive.”); and
performing at least one remediation action based on the detected policy violation (see Williamson: Fig.5, [0091], “an exemplary UI of a primary reporting page indicating where sensitive data is in an organization's data sources, according to an embodiment. In the exemplary UI, the data classification reporting module 114 presents a count of the detected sensitive data in the sensitive data types view 502. The counts presented by the data classification reporting module 114 may represent each detection of a data portions that are sensitive data, or may represent detections of sensitive data types per subsection of the input data source.”)
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify Smith’s language-based classification system to include the sensitivity detection model and the labeling training data and training a sensitivity detection machine learning model as taught by Williamson. One would have been motivated to make such a combination in order to improve the tedious procedure of manually discovering sensitive data in electronic systems to scale and compete with the pace of data growth and distribution. (see: [0003])
Regarding Claim 5,
Smith and Williamson and teaches all the limitations of claim 1. Smith further teaches the system comprising:
classifying the third resource to determine the third classification by applying a classifier machine learning model to data of the third resource (see Smith: Fig.8, [0187]-[0189], For Each Group or Cluster, Train a Classifier to Classify a Datapoint as Either Inside or Outside the Cluster (module 411); [0188] Store Each Trained Classifier and Associate Classifier with Cluster or Group's Identifier (module 412); [0189] For Each New Datapoint, Use Datapoint as Input to Each Classifier to Determine Most Likely Cluster or Group to Which Datapoint Should be Assigned (module 413);”)
Regarding Claim 6,
Smith and Williamson and teaches all the limitations of claim 1. Smith further teaches the method wherein:
fine-tuning the language model (see Smith: Fig.8 [0248], “When adapting foundation models to be performant in production on complex predictive tasks, fine-tuning on custom-labeled training data is critical from initial self-supervised training to final target task fine-tuning.”), further includes, at each iteration:
comparing outputs generated by the language model at the iteration to a set of reference sensitivities for the plurality of first resources, wherein the weights of the language model are adjusted based on the comparison (see SEGEV: Fig. 8, [0238], “Training of a network is performed using a “labelled” dataset of inputs in an assortment of representative input patterns (or datasets) that are associated with their intended output response. Training uses general-purpose methods to iteratively determine the weights for intermediate and final feature neurons. In terms of a computational model, each neuron calculates the dot product of inputs and weights, adds a bias, and applies a non-linear trigger or activation function (for example, using a sigmoid response function)..”)
Regarding Claim 7,
Smith and Williamson and teaches all the limitations of claim 1. Williamson further teaches the method wherein:
each first prompt of the plurality of first prompts further indicates a predetermined set of sensitivity levels, wherein the fine-tuned language model is trained to output sensitivity levels among the predetermined set of sensitivity levels (see Williamson: Fig.1, [0038], “data classifier 108 further determines the security level of data portions determined to be sensitive (or likely to be sensitive and having a confidence value beyond a certain threshold). The security level indicates how well the sensitive data portion is currently protected by various security features. The data classifier 108 may be able to detect the number and type of security feature(s) applied to the data portion. The security level may be determined to be higher based on the number of security features applied to the data portion, as well as the strength of the security feature applied to the data portion. In one embodiment, the data classifier 108 has received the various encryption keys, tokenization mapping, de-obfuscation protocol, and other information needed to read data for which one or more security features have been applied.”)
it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify Smith’s language-based classification system to include the sensitivity detection model and the labeling training data and training a sensitivity detection machine learning model as taught by Williamson. Smith teaches applying a language model to input data to generate classification outputs but does not explicitly disclose a model for detecting a sensitive data and adjusting model parameter through refinement. One would have been motivated to make such a combination in order to improve the tedious procedure of manually discovering sensitive data in electronic systems to scale and compete with the pace of data growth and distribution. (see Williamson: [0003])
Regarding independent Claim 8,
Claim 8 is directed to a system claim and has similar/same claim limitation as Claim 1 and is rejected under the same rationale.
Regarding Claims 9-14,
Claims 9-14 are directed to a system claim and have similar/same claim limitation as Claims 2-7 respectively and are rejected under the same rationale.
Regarding independent Claim 15,
Claim 15 is directed to a method claim and has similar/same claim limitation as Claim 1 and is rejected under the same rationale.
Regarding Claims 16-20,
Claims 16-20 are directed to a method claim and have similar/same claim limitation as Claims 3-7 respectively and are rejected under the same rationale.
Regarding independent Claim 21,
Claim 21 is directed to a system claim and has similar/same claim limitation as Claim 1/15 and is rejected under the same rationale.
Regarding Claims 22-26,
Claims 22-26 are directed to a method claim and have similar/same claim limitation as Claims 3-7 respectively and are rejected under the same rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
PGPUB
NUMBER:
INVENTOR-INFORMATION:
TITLE / DESCRIPTION
US 20220027680 A1
Friedland; Gerald
Title: METHODS AND SYSTEMS FOR FACILITATING CLASSIFICATION OF LABELLED DATA
Description: Generally, the present disclosure relates to the field of data processing. More specifically, the present disclosure relates to methods and systems for facilitating classification of labelled data..
US 20230153462 A1
Tutuianu; Aurelian
Title: EFFICIENT STATISTICAL TECHNIQUES FOR DETECTING SENSITIVE DATA
Description: The present disclosure relates to methods and apparatus for efficiently detecting the presence of sensitive data within large data sets using statistical techniques. The target data sets for which sensitive data presence analysis is to be conducted may, for example, be stored at database services or other storage-related services of a provider network or cloud computing environment.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZELALEM W SHALU whose telephone number is (571)272-3003. The examiner can normally be reached M- F 0800am- 0500pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Zelalem Shalu/Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145