Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The disclosure is objected to because of the following informalities: The abstract because the page contains a label “Figure 8”.
Appropriate correction is required.
Drawings
The drawings Figure 7A and 7C are objected because they contain blurry and illegible characters.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 7, 15, 17 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention
Claim 7 and 17 recite “inside classification engine”, it is not clear if this is the same or different engine from the classification engine recited in the parent claims.
Claim 15 depends from claim 14 which is a “method claim”. Claim 15 recites “The system of claim 14” even though claim 14 is a method claim with no antecedent “system”.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 9, 11, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Padiyar et al. (US Pub. No. 2025/0061214) in view of Sahu et al. (US Pat. No. 12,499,267).
Claim 1. Padiyar discloses A system for detecting and preventing sensitive data sharing by users based on Application Programming Interface (API) call analysis, the system comprising: a receiver engine to receive one or more API calls between a client device and an application, wherein the API calls include data transmitted during user interactions with the application; (Padiyar Par. (0042) "API request and API responses may be intercepted by tracing agents at step 415. Intercepting API requests and responses may include tracing agent code inserted within the micro service to generate a copy of an incoming API request and store a copy of an outgoing API response before the request is processed and/or before the response is transmitted.") (Par. (0026) “In operation, a user 112 may initiate a request through client device 110 to network service 103. API gateway 120 receives the API request, and process the request by calling on one of micro-services 121, 122, 123 or 124 to process the request.”)) (Par. (0004) “A user model may be generated from the API request and response data, and may identify a user's typical API access points, geographical location, sensitive information requests, and other user data.”))
Padiyar does not explicitly teach a classification engine to: classify the data included in the API calls into one or more categories of information by employing a customized Large-Language Model (LLM), and identify sensitive data from the classified data based on pre-defined criteria.
However, Sahu teaches; a classification engine to: classify the data included in the API calls into one or more categories of information by employing a customized Large-Language Model (LLM), and identify sensitive data from the classified data based on pre-defined criteria; (Sahu Par. (7) "A detection module includes a semantic rules engine and an ensemble of artificial intelligence models configured to perform context-based classification…An identification module receives detected data attributes output by the detection module and invokes identification markers to generate sensitive data identification information.")
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar with Sahu because the present disclosure relates to a multi-layered, multi-pathed system, method, and service for using a cognoscible computing engine to systematically parse incoming data in any form, detect the relevant information, identify and classify the information, confirm if the information is accurate, then tag, flag and take appropriate information as deemed to be relevant as per the configuration. [Sahu, para. 6]
Padiyar in view of Sahu further discloses a report and response engine to generate a report of the identified sensitive data and communicate the generated report to an administrator; (Padiyar Par. (0033) "Alert generation module 230 may be triggered to generate an alert, set flags, and generate and transmit notifications to a user, administrator, customer, or some other party.") and a solution engine to execute one or more actions based on the generated report to prevent the sharing of the identified sensitive data with the application. (Padiyar Par. (0059) "At step 640, at least a portion of a request or response is blocked. Blocking a portion of a response may include modifying a portion of the response. The modification can include, for example, replacing sensitive user data with a token, hash, or other value. In some instances, modification includes suppressing the sensitive user data by scrambling the data or removing the data.").
Claim 9. Padiyar in view of Sahu discloses the system of claim 1, wherein the solution engine is further configured to perform actions including at least one of: blocking, filtering, and altering the data before it is transmitted to the application. (Padiyar Par. (0059) "Blocking a portion of a response may include modifying a portion of the response. The modification can include, for example, replacing sensitive user data with a token, hash, or other value. In some instances, modification includes suppressing the sensitive user data by scrambling the data or removing the data.")
Claims 11 and 19 are similar to claims 1 and 9, therefore they are rejected for the same reasons.
Claim 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Padiyar in view of Sahu, and further in view of Padgett et al. (US Pat No.2024/0160902).
Claim 2. Padiyar in view of Sahu discloses the features of claim 1. Padiyar in view of Sahu does not explicitly teach the system of claim 1, wherein the customized Large-Language Model (LLM) is pre-trained on publicly available data and fine-tuned using proprietary data specific to an organization.
However, Padgett teaches; The system of claim 1, wherein the customized Large-Language Model (LLM) is pre-trained on publicly available data and fine-tuned using proprietary data specific to an organization. (Padgett Par. (0046) “a ML model for generating natural language that has been trained generically on publicly-available text corpuses may be, e.g., fine-tuned by further training using… training data samples”)) (Par. (0073) “ the training may include a two-stage training session in which a large language model is first trained on a non-specific training set.”)) (Par. (0074) “The second training stage may include a more specific or fine-grained tuning of the generally-trained model using a more specific training data set.”)) (Par. (0108) “The training data set may include items protected by one or more forms of intellectual property and/or PII, and unavailable for free and unrestricted use, or available only under the terms of one or more user licenses.”)) (Par. (0074) “For example, the more specific training data set may be tailored to the intended purpose or objective of the LLM, such as a data set relating to corporate or brand logos in the case of creating an LLM for suggesting new logos, or the data set may be a training data set of computer code snippets or excerpts in the case of creating an LLM for generating suggested computer code.”))
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu with Padghett because It would be advantageous to provide for systems and methods that improve generative AI models in a manner that avoids or reduces “problematic” outputs, i.e. outputs that are too similar to pre-existing content, without necessarily compromising the quality of the training data set. In one aspect, the present application provides for a similarity-assessment layer that determines a similarity measure for the output of a generative AI model. The similarity measure is based on a determination of the similarity of the output to items in a repository of pre-existing content. The determination of similarity may be based on a distance metric that evaluates the similarity of two items (text, images, multimedia, or other content). Examples may include an L1 norm, L2 norm, Chebyshev distance, Euclidean distance, cosine distance, or others. [Padgett, para.0037]
Claim 12 is similar to claim 2, therefore it is rejected for the same reasons
Claim 3, 7, 13 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Padiyar in view of Sahu, and further in view of Padgett (US Pat No.2024/0160902) and Crume et al. (US Pat. No. 12,430,464).
Claim 3. Padiyar in view of Sahu in view of Padgett discloses the system of claim 2, Padiyar in view of Sahu in view of Padgett do not explicitly teach wherein the customized LLM is fine-tuned to recognize and classify sensitive data.
However, Crume teaches wherein the customized LLM is fine-tuned to recognize and classify sensitive data (Crume Par. (50) "data types specifying personally identifiable information (PII), e.g., credit card numbers, social security numbers, and the like, data types indicative of intellectual property, or any other data types determined to be of a sensitive nature.") (Par. (82) "The pre-trained large language model (LLM) is fine-tuned to perform DLP sensitive data identification (step 620). The fine tuning may involve retraining the pre-trained LLM using labeled DLP training data specifying an input text and the types of sensitive information present in the input text")
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu with Crume because the present application relates generally to an improved data processing apparatus and method and more specifically to an improved computing tool and improved computing tool operations/functionality for data leakage protection using generative large language models. [Crume, para. 1]
Crume does not teach, LLM is fine-tuned to recognize…. proprietary information specific to the organization, including at least one of: trade secrets, business strategies and plans, financial information, customer and vendor lists, product formulas and recipes, new technologies and inventions, software and databases, internal correspondence and communications, marketing tactics and materials, negotiation strategies and pricing models, and employee information and HR records.
However, Padgett teaches, LLM is fine-tuned to recognize…. proprietary information specific to the organization, including at least one of: trade secrets, business strategies and plans, financial information, customer and vendor lists, product formulas and recipes, new technologies and inventions, software and databases, internal correspondence and communications, marketing tactics and materials, negotiation strategies and pricing models, and employee information and HR records (Padgett (Par. (0074) “For example, the more specific training data set may be tailored to the intended purpose or objective of the LLM, such as a data set relating to corporate or brand logos in the case of creating an LLM for suggesting new logos, or the data set may be a training data set of computer code snippets or excerpts in the case of creating an LLM for generating suggested computer code.”))
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu with Padghett because It would be advantageous to provide for systems and methods that improve generative AI models in a manner that avoids or reduces “problematic” outputs, i.e. outputs that are too similar to pre-existing content, without necessarily compromising the quality of the training data set. In one aspect, the present application provides for a similarity-assessment layer that determines a similarity measure for the output of a generative AI model. The similarity measure is based on a determination of the similarity of the output to items in a repository of pre-existing content. The determination of similarity may be based on a distance metric that evaluates the similarity of two items (text, images, multimedia, or other content). Examples may include an L1 norm, L2 norm, Chebyshev distance, Euclidean distance, cosine distance, or others. [Padgett, para.0037].
Claim 13 is similar to claim 3, therefore it is rejected for the same reasons
Claim 7. Padiyar in view of Sahu discloses the system of claim 1, Padiyar in view of Sahu does not explicitly disclose wherein the classification engine is further configured to include a customized LLM model training services for both the cloud and on-premise deployments by fine tuning one or more selected foundational LLM models with at least one of: customer specific and proprietary data, wherein the customized LLM model is further utilized inside classification engine for sensitive data detection.
However, Crume teaches wherein the classification engine is further configured to include a customized LLM model training services for both the cloud and on-premise deployments by [fine tuning one or more selected foundational LLM models with at least one of: customer specific and proprietary data,] wherein the customized LLM model is further utilized inside classification engine for sensitive data detection. (Crume Par. (64) "LLMs 320 can be employed in on-premises or private cloud environments, minimizing the exposure of sensitive data to external systems or networks. This ensures greater control and privacy over the information being processed.") (Par. (73) "a pre-trained generative LLM 412, i.e., a LLM with generative AI computer models such as LLMs 320 and generative AI computer models 330 in FIG. 3, is provided and a model fine-tuning operation 414 is performed based on a labeled dataset for DLP 416. That is, the pre-trained generative LLM 412 is further trained based on the labeled dataset for DLP 416 to fine tune the pre-trained generative LLM 412 for use in DLP operations.") (Par. (82) "The pre-trained large language model (LLM) is fine-tuned to perform DLP sensitive data identification (step 620). The fine tuning may involve retraining the pre-trained LLM using labeled DLP training data specifying an input text and the types of sensitive information present in the input text") (Par. (42) "The illustrative embodiments enhance the collection engine's 220 data discovery and classification logic by implementing large language models (LLMs)…to perform automatic detection of sensitive data in various contexts")
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu with Crume because it relates generally to an improved data processing apparatus and method and more specifically to an improved computing tool and improved computing tool operations/functionality for data leakage protection using generative large language models [Crume, para. 1].
Crume does not teach, fine tuning one or more selected foundational LLM models with at least one of: customer specific and proprietary data.
However, Padgett teaches, fine tuning one or more selected foundational LLM models with at least one of: customer specific and proprietary data. (Padgett Par. (0057) “a ML model for generating natural language that has been trained generically on publically-available text corpuses may be, e.g., fine-tuned by further training using… training data samples”)) (Par. (0073) “ the training may include a two-stage training session in which a large language model is first trained on a non-specific training set.”)) (Par. (0074) “The second training stage may include a more specific or fine-grained tuning of the generally-trained model using a more specific training data set.”)) (Par. (0108) “The training data set may include items protected by one or more forms of intellectual property and/or PII, and unavailable for free and unrestricted use, or available only under the terms of one or more user licenses.”))
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu and Crume with Padghett because It would be advantageous to provide for systems and methods that improve generative AI models in a manner that avoids or reduces “problematic” outputs, i.e. outputs that are too similar to pre-existing content, without necessarily compromising the quality of the training data set. In one aspect, the present application provides for a similarity-assessment layer that determines a similarity measure for the output of a generative AI model. The similarity measure is based on a determination of the similarity of the output to items in a repository of pre-existing content. The determination of similarity may be based on a distance metric that evaluates the similarity of two items (text, images, multimedia, or other content). Examples may include an L1 norm, L2 norm, Chebyshev distance, Euclidean distance, cosine distance, or others. [Padgett, para.0037]
Claim 17 is similar to claim 7, therefore it is rejected for the same reasons
Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Padiyar in view of Sahu, and further in view of Crume et al. (US Pat. No. 12,430,464).
Claim 4. Padiyar in view of Sahu discloses the system of claim 1, Padiyar in view of Sahu does not teach: wherein the sensitive data identified by the classification engine includes Personal Identifiable Information (PII).
However, Crume teaches: wherein the sensitive data identified by the classification engine includes Personal Identifiable Information (PII). (Crume Par. (50) "Data discovery tools scan network repositories such as databases, servers, and storage devices, e.g., data sources 280-284 of FIG. 2, for data types that match predefined policies, such as data types specifying personally identifiable information (PII), e.g., credit card numbers, social security numbers, ")
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu with Crume because it relates generally to an improved data processing apparatus and method and more specifically to an improved computing tool and improved computing tool operations/functionality for data leakage protection using generative large language models. [Crume, para. 1]
Claim 14 is similar to claim 4, therefore it is rejected for the same reasons.
Claims 6, 8, 10, 16, 18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Padiyar in view of Sahu, and further in view of Manmohan (US Pat. No. 9,256,727).
Claim 6. Padiyar in view of Sahu discloses the features of claim 1. Padiyar in view of Sahu does not explicitly teach The system of claim 1, wherein the classification engine is further configured to analyze fragments of information from multiple API calls to detect potential sensitive data leaks through data aggregation.
However, Manmohan teaches; wherein the classification engine is further configured to analyze fragments of information from multiple API calls to detect potential sensitive data leaks through data aggregation. (Manmohan Par. (8) "by identifying and accumulating multiple full and/or partial DLP policy violations, the disclosed systems and methods may detect drip data leaks that may have gone undetected using traditional DLP systems… performing a security action only when the aggregate of multiple DLP policy violations has exceeded a threshold.") (Par. (36) "If the user distributes the credit card information of 25 customers via a removable device and later distributes the credit card information of an additional 26 customers via a messaging account, detection module 106 may determine that the user has partially violated the DLP policy twice. Since the user distributed the credit card information of more than 50 customers in total, detection module 106 may record the two partial DLP policy violations as a full DLP policy violation.") (Par. (26) "monitoring module 104 may identify one or more computing devices and/or web accounts (e.g., email accounts, social networking accounts, etc.) that belong to a user. Monitoring module 104 may then form an association between the multiple data-distribution channels in order to detect data leaks committed by the user from multiple access points.")
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu with Manmohan because the instant disclosure generally relates to systems and methods for both (1) detecting data leaks (by, e.g., tracking and accumulating both full and partial DLP policy violations) and (2) providing feedback to DLP administrators to restructure relaxed policies. [Manmohan, para. 4]
Claim 8. Padiyar in view of Sahu discloses The system of claim 1, Padiyar in view of Sahu does not teach: the solution engine is further configured to block the client device from accessing at least one of: the application and the network upon detection of sensitive data within the API calls.
Manmohan teaches: the solution engine is further configured to block the client device from accessing at least one of: the application and the network upon detection of sensitive data within the API calls (Manmohan Par. (45) "Security module 110 may also disable the entity's access to one or more data-distribution channels and/or sensitive data, such as by disabling the entity's access to a computing device and/or the Internet.") (Par. (43) “security module 110 may perform a security action in response to determining that DLP policy violations 210 exceed the predetermined threshold.”))
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu with Manmohan because the instant disclosure generally relates to systems and methods for both (1) detecting data leaks (by, e.g., tracking and accumulating both full and partial DLP policy violations) and (2) providing feedback to DLP administrators to restructure relaxed policies. [Manmohan, para. 4]
Claim 10. Padiyar in view of Sahu discloses The system of claim 1, Padiyar in view of Sahu does not teach wherein the administrator corresponds to authorized personnel within the organization including at least one of: a designated security manager and a system administrator
Manmohan teaches: wherein the administrator corresponds to authorized personnel within the organization including at least one of: a designated security manager and a system administrator. (Manmohan Par. (45) "security module 110 may notify an administrator that the entity's DLP policy violations cumulatively exceed the predetermined threshold.")
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu with Manmohan because the instant disclosure generally relates to systems and methods for both (1) detecting data leaks (by, e.g., tracking and accumulating both full and partial DLP policy violations) and (2) providing feedback to DLP administrators to restructure relaxed policies. [Manmohan, para. 4]
Claims 16, 18, and 20 are similar to claims 6, 8, and 10, respectively, therefore they are rejected for the same reasons.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Padiyar in view of Sahu and Crume, and further in view of Gupta et al. (US Pub. No. 2024/0289492), and Zhang et al (US Pat. No. 11,475,158).
Claim 5. Padiyar in view of Sahu and Crume discloses the features of claim 4. Padiyar in view of Sahu and Crume does not teach the system of claim 4, wherein the PII is further categorized into: a common PII that is identified based on a global set of rules and regular expressions stored in the classification engine; an industry-specific PII that is identified using industry-specific rules and a machine learning model trained on industry-specific data; and a customer-specific PII that is identified using customer-defined rules and a machine learning model trained on customer-specific data.
However, Gupta teaches; wherein the PII is further categorized into: a common PII that is identified based on a global set of rules and regular expressions stored in the classification engine; an industry-specific PII that is identified using industry-specific rules and a machine learning model trained on industry-specific data; (Gupta Par. (0006) "The first layer builds a machine learning model…A second layer may utilize a rule-based pattern-matching algorithm that finds matches of existing and available PII data") (Par. (0007) "The disclosed PII sensitivity detection framework can be adapted...for a particular organization (e.g., adapts to rules and regulations that define PII)")
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu in view of Crume with Gupta because currently, the process of identifying PII is manual (e.g., review by a human expert). This process is undesirable because it is prone to errors and misclassifications due to the reviewer's subjective skills and understanding. Manual reviews are also undesirable because they are not scalable. [Gupta, para. 3]
Gupta does not teach; a customer-specific PII that is identified using customer-defined rules and a machine learning model trained on customer-specific data.
Zhang teaches; a customer-specific PII that is identified using customer-defined rules and a machine learning model trained on customer-specific data. (Zhang Par. (17) "Customers further develop disclosed custom classifiers using their own sensitive training data, such as medical/design images, human resources (HR) documents, etc.") (Par. (48) "using the received organization-specific examples to generate a customer-specific DL stack classifier")
(Examiner's note: DL stands for deep learning, which is interpreted as the ML model.)
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify Padiyar in view of Sahu in view of Crume in view of Gupta with Zhang because the technology disclosed relates generally to security for network delivered services. In particular it relates to building a customized deep learning (DL) stack classifier to detect organization sensitive data in images, referred to as image-borne organization sensitive documents, under the organization's control, for protecting against loss of the image-borne organization sensitive documents without the organization sharing the image-borne organization-sensitive documents, even with the security services provider. [Zhang, para. 21]
Claim 15 is similar to claim 5, therefore it is rejected for the same reasons.
Pertinent prior art
Pertinent prior art made of record, however not relied upon, includes:
Jaiswal et al (US Patent No. 8688601).
ConclusionAny inquiry concerning this communication or earlier communications from the examiner should be directed to ABDULLAHI MOHAMED ABDULLAHI whose telephone number is (571)272-9615. The examiner can normally be reached 7:30 - 5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ali Shayanfar can be reached at (571) 270-1050. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ABDULLAHI MOHAMED ABDULLAHI/Examiner, Art Unit 2434
/NOURA ZOUBAIR/Primary Examiner, Art Unit 2434