DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
2. This is in response to the applicant response filed on 06/24/2026. In the applicant’s response, claims 1and 18 were amended. Accordingly, claims 1-19 are pending and being examined. Claims 1 and 18 are independent form.
Claim Rejections - 35 USC § 102
3. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
4. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
5. Claims 1-4, 7-11, 13-15, and 18 are rejected under 35 U.S.C. 102(a)(1)/102(a)(2) as being anticipated by Harkous et al (WO2023/224672, hereinafter “Harkous”).
Regarding claim 1, Harkous discloses a computer-implemented method for autonomously classifying uncategorised data items within an environment (the method using a computing system is described that classifies each text from a plurality of texts into one or more categories; see Abstract and fig.1), the method comprising:
obtaining, from at least one data source within the environment, a plurality of uncategorised data items (see the “feedback texts” of fig.6A, the “input” of fig.3A, and para.58, lines 4-8: “Thus, FIG. 3 A illustrates machine-learned model 300 performing inference. For example, the input data received by machine-learned model 300 may be feedback texts, and the output data provided by machine- learned [ML] model 300 may be issue clusters and names, theme clusters and names, quality measures, emotion classifications, junk classifications, etc.”);
generating, using a machine learning, ML, model, at least one embedding vector for each uncategorised data item, where the at least one embedding vector represents content of each uncategorised data item (ibid. Further see para.59, lines 1-3: “The input data may include one or more features that are associated with an instance or an example. In some implementations, one or more features associated with the instance or example may be organized into a feature vector.”);
clustering, using the at least one embedding vector generated for each uncategorised data item, the plurality of uncategorised data items into a plurality of clusters (see the “output” of fig.3A, and para.58, lines 4-8: “Thus, FIG. 3 A illustrates machine-learned model 300 performing inference. For example, the input data received by machine-learned model 300 may be feedback texts, and the output data provided by machine- learned model 300 may be issue clusters and names, theme clusters and names, quality measures, emotion classifications, junk classifications, etc. Also see the “issue grouping 632” of fig.6A and para.184), where each cluster contains a subset of the plurality of uncategorised data items that are more similar to each other than to the uncategorised data items in other clusters (see para.42: “these issues may be aggregated across the whole corpus and grouped into themes based on semantic similarity. Each group of issues and the associated feedback texts may constitute a theme.”)
generating, using a large language model, LLM, at least one classification label for each cluster wherein the at least one classification label is specific to content of the subset of the plurality of uncategorised data items in the cluster (see fig.3A and para.62: the “machine-learned model 300 may perform various types of natural language processing (NLP) based on the input data. For example, machine-learned model 300 may summarize, translate, or organize the input data. The natural language processing (NLP) model may use recurrent neural networks (RNNs) and/or transformer models (self-attention models), such as GPT-3, BERT, and T5.”; also see para.27, wherein a feedback text may label/classify as “privacy” or “not privacy”); and
applying, to each uncategorised data item in each cluster, the at least one classification label generated for the cluster, thereby generating a labelled data item (see the “output” of fig.6A, and para.58, lines 4-8: “Thus, FIG. 3 A illustrates machine-learned model 300 performing inference. For example, the input data received by machine-learned model 300 may be feedback texts, and the output data provided by machine- learned model 300 may be issue clusters and names, theme clusters and names, quality measures, emotion classifications, junk classifications, etc.);
retrieving, for each of generated labelled data item, at least one data management policy corresponding to the at least one classification label of the labelled data item; and
using the at least one data management policy to control an action performed with respect to the generated labelled data item (For an input text of “I don’t’ know how to delete my account”, the classifier model 443A labels (i.e., classifies) it as “privacy”; accordingly, the model 433B generates the issue of “account deletion” and the model 433C generates the corresponding account management action, namely “account deactivation” corresponding to the classification label. See fig.4 and para.146.).
Regarding claim 2, Harkous discloses the method of claim 1 wherein obtaining a plurality of uncategorised data items comprises obtaining any one or more of: an email, a document, a file, a text file, a folder, an image, a video, an audio file, a diagram, a geographical map, a medical image, a medical data file, and a portable document format file (see the “feedback texts” in fig.6A, and the “feedback texts storage 103” in fig.1).
Regarding claim 3, Harkous discloses the method of claim 1 wherein clustering the plurality of uncategorised data items comprises using any one of: a data clustering algorithm, a k-means clustering algorithm, and a density-based spatial clustering algorithm (see para.68: “In some implementations in which machine-learned model 300 performs clustering, machine-learned model 300 may be trained using unsupervised learning techniques.”).
Regarding claim 4, Harkous discloses the method of claim 1 wherein, when a single embedding vector is generated for each uncategorised data item, clustering the plurality of uncategorised data items comprises clustering each embedding vector in embedding space, and thereby clustering the plurality of uncategorised data items into a plurality of clusters (see para.33: “Theme title generation module 112 may take the issues and group them in clusters representing high-level themes. Each cluster may contain a set of related fine-grained issues. Theme title generation module 112 may include a generative model that assigns a concise title for each theme (e.g., “Sharing Concerns” or “Data Deletion”). This eliminates the manual work required to interpret clusters.”).
Regarding claim 7, Harkous discloses the method of claim 1 wherein generating, using a large language model, at least one classification label comprises: analysing the uncategorised data items in each cluster to determine at least topic representative of content of the subset of the plurality of uncategorised data items in the cluster (ibid.).
Regarding claim 8, Harkous discloses the method of claim 7 further comprising specifying a maximum number of topics to be generated for the plurality of uncategorised data items (see para.33: “Theme title generation module 112 may take the issues and group them in clusters representing high-level themes. Each cluster may contain a set of related fine-grained issues. Theme title generation module 112 may include a generative model that assigns a concise title for each theme (e.g., “Sharing Concerns” or “Data Deletion”). This eliminates the manual work required to interpret clusters.”).
Regarding claim 9, Harkous discloses the method of claim 7 wherein generating, using a large language model, at least one classification label comprises: inputting the at least one topic for each cluster into the large language model, LLM; and obtaining for each topic, from the LLM, at least one classification label and a description of the topic (see fig.3A and para.62: the “machine-learned model 300 may perform various types of natural language processing (NLP) based on the input data”; see para.58, lines 6-8: “the output data provided by machine- learned model 300 may be issue clusters and names, theme clusters and names, quality measures, emotion classifications, junk classifications, etc )
Regarding claim 10, Harkous discloses the method of claim 1 wherein generating, using a large language model, at least one classification label comprises: selecting a sample of uncategorised data items from the cluster (see 113 of fig.1 para.24: “The feedback texts may be filtered using junk classifier module 113 to remove junk feedback texts before processing.”); inputting the sample of uncategorised data items into the large language model, LLM together with at least one prompt to instruct the LLM to output at least one classification label; and obtaining, from the LLM, at least one classification label for the input sample of uncategorised data items (see 107, 105, 11, and 112 of fig.1).
Regarding claim 11, Harkous discloses the method of claim 10 further comprising: inputting, into the LLM, a maximum number of classification labels to be generated by the LLM (see fig.1 and par.22—par.33).
Regarding claim 13, Harkous discloses the method of claim 1 further comprising: storing, in a database, the generated embedding vectors and associated classification label (ibid.).
Regarding claim 14, Harkous discloses the method of claim 13 further comprising: obtaining a new uncategorised data item; generating at least one embedding vector for the new uncategorised data item; comparing the generated at least one embedding vector to the database of stored embedding vectors; selecting, responsive to the comparing, at least one stored embedding vector that is most similar to the generated at least one embedding vector for the new uncategorised data item; and applying to the new uncategorised data item, at least one classification label corresponding to the selected at least one stored embedding vector, thereby generating a new labelled data item (see the issue generation module 611 and theme generation module 612 shown by fig.6A, see para.160, para.164-para.166, and para.184).
Regarding claim 15, Harkous discloses the method of claim 14 further comprising: outputting information explaining how the at least one classification label of the new labelled data item is determined (see para.184: “Theme title generation module 612 may include two stages: issue grouping 632 and title generator 634. After obtaining the issues, these issues may be grouped into themes.”).
Regarding claim 18, claim 18 is an inherent variation of claim 1, thus it is interpreted and rejected for the reasons set forth in the rejection of claim 1.
Claim Rejections - 35 USC § 103
6. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
7. Claims 5-6, and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Harkous and in view of Gebow et al (US20250384382, hereinafter “Gebow”).
Regarding claim 5, Harkous does not explicitly disclose: prior to generating at least one embedding vector, dividing the uncategorised data item into two or more segments; wherein generating the at least one embedding vector comprises generating an embedding vector for each of the two or more segments. However, in the same field of endeavor, that is, in the field of performing natural language processing (NLP) based on documents, Gebow teaches, identifying and segmenting the different regions within the document into text blocks, tables, and/or images. See para.104. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was made to incorporate the teachings of Gebow into the teachings of Harkous and identify and segment the different regions within the input document into text blocks taught by Gebow. Suggestion or motivation for doing so would have been to preprocess text image regions and recognize individual characters and words as taught by Gebow, see para.105. Therefore, the combination of Harkous and Gebow suggests or teaches all the limitations of the claim.
Regarding claim 6, the combination of Harkous and Gebow discloses the method of claim 5 further comprising: calculating an average embedding vector for each uncategorised data item by averaging the embedding vector generated for each segment of the data item (Harkous, see para.100, some or all of the input data may be normalized by subtracting the mean across a given dimension’s feature values from each feature value and then dividing by the standard deviation or another metric); wherein clustering the plurality of uncategorised data items comprises clustering the average embedding vectors in embedding space, and thereby clustering the plurality of uncategorised data items into a plurality of clusters (Harkous, see para.105: the output data may include various types of classification data, the output data may include clustering data, anomaly detection data, recommendation data, or any of the other forms of output data discussed above.).
Regarding claim 12, Harkous does not disclose inputting, into the LLM, at least one further prompt to ensure the at least one classification label complies with predefined responsible AI guidelines. However, in the same field of endeavor, Gebow teaches, the large language models (LLM) may incorporate proprietary guidelines, as AI-enabled large language models generally have access to published federal and international regulations but may not have access to regulatory guidelines such as ISO standards without licensure. Finally, the invention enforces homogenized criteria for reporting thresholds, ensuring consistent and impartial results across diverse quality assessments. See pra.16. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was made to incorporate the teachings of Gebow into the teachings of Harkous and incorporate proprietary guidelines, as AI-enabled large language models taught by Gebow. Suggestion or motivation for doing so would have been to have access to regulatory guidelines as taught by Gebow, see para.16. Therefore, the combination of Harkous and Gebow suggests or teaches all the limitations of the claim.
8. Claims 16-17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Harkous and in view of Fan et al (CN116227473, hereinafter “Fan”). A machine translated English version (i.e., CN116227473-Eng) of the document CN116227473 is provided by the examiner with this office action.
Regarding claim 16, Harkous does not explicitly disclose, wherein when none of the stored embedding vectors are similar to the generated at least one embedding vector for the new uncategorised data item, the method comprises: storing, in a second database, the new uncategorised data item. However, adding a new cluster for representing newly discovered data is well-known and widely used in the field of clustering. As evidence, the same field of endeavor, Fan teaches, after the increment supplementary label word and newly discovered label word, for the new label word, if the new label word of the word combination is a new word combination (not comprising the preset service label word in the combination of words), not only to increment the type of the service label word. See pg.12, lines 6-12. In other words, Fan teaches, as long as a new label word (i.e., a new data item) is found by the system, a corresponding type of label words (i.e., a new cluster) is increased by the clustering method. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was made to incorporate the teachings of Fan into the teachings of Harkous and increase a new cluster for describing the new text (i.e., “newly discovered label word”) found in the input data. Suggestion or motivation for doing so would have been to improve the automatic digging efficiency and accuracy of the label synonym as taught by Fan, see pg.14, lines 14-18. Therefore, the combination of Harkous and Fan suggests or teaches all the limitations.
Regarding claim 17, 19, the combination of Harkous and Fan discloses the method of claim 16 wherein when the second database contains a predefined threshold number of new uncategorised data items, the method further comprises clustering, using the generated at least one embedding vector for each new uncategorised data item, the new uncategorised data items into a plurality of clusters, where each cluster contains a subset of the new uncategorised data items that are more similar to each other than to the new uncategorised data items in other clusters (Fan, see pg.12, lines 14-18: “the topic core word for clustering the label type number can be synchronously increased with the type of the newly discovered label word, so as to automatically increment updating the synonym word stock of the label word according to the method of the embodiment”).
Response to Arguments
9. Applicant’s arguments, with respects to claims 1 and 18, filed on 6/24/2026, have been fully considered but they are not persuasive.
On page 8 of applicant’s response, applicant argues:
Harkous fails to disclose either "retrieving, for each generated labelled data item, a least one data management policy corresponding to the at least one classification label of the labelled data item;" or "using the at least one data management policy to control an action performed with respect to the generated labelled data item." (emphasis added). Accordingly, Harkous fails to anticipate claims 1 or 18, and the § 102 rejections thereof are improper and must be withdrawn.
The examiner respectfully disagrees with the applicant’s argument. It is because Harkous, see paragraph [0146], clearly states “model 433A corresponding to classifier module 107 of FIG. 1 inputs a feedback text of “I don’t know how to delete my account” and classifies it as related to privacy. Model 433B corresponding to issue generation module 111 of FIG. 1 generates an issue for the feedback text of “I don’t know how to delete my account” as being “account deletion”. Model 433C corresponding to theme title generation module 112 of FIG. 1. generates a theme using input issue titles such as “account deletion” and “account deactivation” to produce a theme title of “account management.”” In other words, the classification model 433A in Harkous retrieves (extracts) the management policy as being “privacy” from an input text of “I don’t know how to delete my account”, and then the Model 433C in Harkous performs the action of “account deactivation” based on the policy. Thus, the argument is unpersuasive.
Conclusion
10. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
11. Any inquiry concerning this communication or earlier communications from the examiner should be directed to RUIPING LI whose telephone number is (571)270-3376. The examiner can normally be reached 8:30am--5:30pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, HENOK SHIFERAW can be reached on (571)272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit https://patentcenter.uspto.gov; https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center, and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RUIPING LI/Primary Examiner, Ph.D., Art Unit 2676