DETAILED ACTION
The action is in response to the original filing on June 28, 2023 and the Remarks and Amendments filed on July 8, 2026. Claims 1-20 are pending and have been considered below. Claims 1, 8, and 15 are independent claims. Claims 1, 3, 5, 8, 10, 12, 15, 17, and 19 are amended. Claims 7 and 14 are canceled.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
In review of Applicant’s amendments, filed July 8, 2026, the objections to claim 1 and the specification made in the previous office action have been withdrawn.
The rejections of claims 3, 5, 7, 10, 12, 14, 17 and 19 under 112(b) set forth in the previous office action are withdrawn in view of the amendments and cancellations.
Response to Arguments
Applicant’s arguments and amendments, filed July 8, 2026 regarding the rejections from the previous office actions made under 35 U.S.C. 101 have been fully considered but are unpersuasive.
On page 11 of the Remarks, Applicant asserts that “the claims integrate the alleged judicial exception into a practical application by using the information provided by the alleged judicial exception – the parameters – and applying the parameters from the fine tuning shot that achieved a highest unsupervised class separability metric to the large language model. Thus, the claim improves the operation of the large language model by using the parameters determined to be best for the large language model.” Applicant cites as reasoning MPEP 2106.04(d) and 2106.05(e) and refers to how “Example 46 provided with the October 2019 PEG” shows how a meaningful limitation “can employ the information provided by the judicial exception,” equating it to the limitation reciting the “fine tuning shot that achieved a highest unsupervised class separability metric…” Examiner respectfully disagrees. MPEP 2106.05(e) describes how “Diamond v. Diehr provides an example of a claim that recited meaningful limitations beyond generally linking the use of the judicial exception to a particular technological environment. 450 U.S. 175, 209 USPQ 1 (1981). In Diehr, the claim was directed to the use of the Arrhenius equation (an abstract idea or law of nature) in an automated process for operating a rubber-molding press. 450 U.S. at 177-78, 209 USPQ at 4. The Court evaluated additional elements… and found them to be meaningful because they sufficiently limited the use of the mathematical equation to the practical application of molding rubber products. 450 U.S. at 184, 187, 209 USPQ at 7, 8.” Similarly, Appendix 1 to the October 2019 Update: Subject Matter Eligibility Life Sciences & Data Processing Examples states how claim 2 of Example 46 is eligible because “Limitation (d) specifies that the monitoring component automatically sends a control signal to the feed dispenser to dispense a therapeutically effective amount of supplemental salt and minerals mixed with the feed when the analysis results for the animal indicate that the animal is exhibiting an aberrant behavioral pattern indicative of grass tetany. Thus, limitation (d) does not merely link the judicial exceptions to a technical field, but instead adds a meaningful limitation in that it can employ the information provided by the judicial exception (the mental analysis of whether the animal is exhibiting an aberrant behavioral pattern indicative of grass tetany) to operate the feed dispenser.”
However, MPEP 2106.05(e) also describes how “the claims in Alice Corp. v. CLS Bank International did not meaningfully limit the abstract idea of mitigating settlement risk. 573 U.S. 208, 110 USPQ2d 1976 (2014). In particular, the Court concluded that the additional elements… recited in the system claims did not meaningfully limit the abstract idea because they merely linked the use of the abstract idea to a particular technological environment (i.e., "implementation via computers").” Similarly, the additional element of “applying the parameters” does not meaningfully limit the abstract idea of claim 1, as it merely links the use of the abstract idea to a particular technological environment, which in this case, is unsupervised training of a large language model, instead of a practical application like in Diehr as described in MPEP 2106.05(e), since it describes the technological environment that the claim is used for, but not how it leads to a practical application of the abstract idea (MPEP 2106.05(h) “the additional element in Flook regarding the catalytic chemical conversion of hydrocarbons was not sufficient to make the claim eligible, because it was merely an incidental or token addition to the claim that did not alter or affect how the process steps of calculating the alarm limit value were performed… employing generic computer functions to execute an abstract idea, even when limiting the use of the idea to one particular environment, does not add significantly more, similar to how limiting the abstract idea in Flook to petrochemical and oil-refining industries was insufficient”). It is unclear how a meaningful limitation is added when applying the model parameters (provided by a judicial exception) from a fine-tuning shot which scored a highest unsupervised class separability metric (derived from a judicial exception), unlike the explicit dispensing of “a therapeutically effective amount of supplemental salt and minerals” as a result of the abstract idea as described in claim 2, limitation (d) in Example 46.
One page 12 of the Remarks, Applicant asserts that “amended claim 1 similarly recites a specific technical mechanism, namely the use the unsupervised class separability metric to select parameters for the large language model. Thus, these elements together recite a meaningful way of using the alleged judicial exception beyond generally linking the use of the judicial exception to a particular technological environment.” Applicant cites as reasoning “Ex parte Desjardins, Appeal 2024-000567 (PTAB Sept. 26, 2025, Appeals Review Panel Decision) (recognizing eligibility where claimed user-interface functionality constituted a specific improvement in computer operation).” Examiner respectfully disagrees. As explained above, the elements together do not add a meaningful limitation beyond generally linking the judicial exception to the particular technological environment of using unsupervised training on a large language model (MPEP 2106.05(e) and 2106.05(h)).
Furthermore, unlike the Applicant’s cited reasoning, the amended claim fails to explain how use of “the unsupervised class separability metric to select parameters” constitutes a specific improvement in computer operation. MPEP 2106.05(a) states that “If it is asserted that the invention improves upon conventional functioning of a computer, or upon conventional technology or technological processes, a technical explanation as to how to implement the invention should be present in the specification. That is, the disclosure must provide sufficient details such that one of ordinary skill in the art would recognize the claimed invention as providing an improvement. The specification need not explicitly set forth the improvement, but it must describe the invention such that the improvement would be apparent to one of ordinary skill in the art. Conversely, if the specification explicitly sets forth an improvement but in a conclusory manner (i.e., a bare assertion of an improvement without the detail necessary to be apparent to a person of ordinary skill in the art), the examiner should not determine the claim improves technology. An indication that the claimed invention provides an improvement can include a discussion in the specification that identifies a technical problem and explains the details of an unconventional technical solution expressed in the claim, or identifies technical improvements realized by the claim over the prior art.” The present application’s specification describes at ¶38 how “a metric, referred to herein as ‘the unsupervised class separability metric,’ may leverage topological characteristics of data manifolds to quantify class separability of the data without requiring labels for the data. The unsupervised class separability metric may also be used… for data grouping, detecting dataset duplicates, and assessment of model’s generalizability to a dataset that the model has not trained on” and at ¶39 how “The unsupervised class separability metric may be generated by computing the log-likelihood of the H0 points conditioned on the trained GMM_H0. A data manifold with good class separability tends to generate H0 points that tend to group into distinct clusters, while a data manifold with bad class separability tends to generate H0 points that are more spread and has less tendency to group into distinct clusters. Thus, the computed conditional log-likelihood tends to increase when there is good class separability, and tends to decrease when we have bad class separability.” However, the disclosure neither explicitly sets forth the technical problem solved nor describes the invention in the technical detail required for one of ordinary skill in the art to find it apparent that the claimed invention results in an improvement to the technology or functioning of a computer.
In consideration of these conclusions, the previous rejections under 35 U.S.C. 101 still stand for claims 1-20.
Applicant’s amendments and arguments filed July 8, 2026 regarding the rejections from the previous office action made under 35 U.S.C. 103 have been fully considered but are not persuasive.
On page 13 of the Remarks, Applicant asserts that “the Office Action has failed to establish a prima facie case of obviousness.” Examiner respectfully disagrees. Applicant cites as reasoning “Id.; In re Icon Health & Fitness, Inc., 496 F.3d 1374, 1380 (Fed. Cir. 2007)” regarding how “there must still be some ‘reason that would have prompted’ a person of ordinary skill in the art to combine the elements in the specific way that he or she did.” On pages 13-14, applicant also cites “KSR, 127 S. Ct. at 1740-41” regarding how “modification of a prior art reference may be obvious only if there exists a reason that would have prompted a person of ordinary skill in the art to make the change.” Liu’s own disclosure cites a need for models which are robust against attacks where “a carefully-constructed or “adversarial” sentence or image may “fool” a model into outputting a clearly incorrect classification for that sentence or image, even when the correct classification is readily apparent to a human user” and proposes “a mechanism for virtual adversarial pretraining of one or more mapping layers of a model. After pretraining using the disclosed techniques, the pretrained mapping layers can be tuned with a task-specific layer to perform a specific task using supervised learning” as a solution for the problem within “a natural language processing context” (Liu, ¶¶20-22). As stated in the previous office action, a teaching, suggestion, or motivation to combine the references Liu and Jakubowski is that Jakubowski discloses a method that can better determine the meaning of polysemous words within text data using degree zero persistent homology (Jakubowski, Page 110, Col. 2, Section 5, ¶1 “we challenge the manifold hypothesis for static word vector embeddings and experimentally show that it is more accurate and helpful to view the space of word embeddings as a pinched manifold. We introduce a topological measure of polysemy that correlates well with the number of meanings of a word according to the gold standard of the SemEval-2010 task on Word Sense Induction & Disambiguation. We also produce a surprisingly simple, but topologically motivated solution to the task itself that achieves highly competitive results”). One of ordinary skill in the art may combine the “topological polysemy” method of Jakubowski with Liu’s training methodology to further enhance a model’s ability to better learn the meanings of words or tokens within textual data, especially if the words have multiple meanings (MPEP § 2143(I)(G) “implicit motivation to combine exists not only when a suggestion may be gleaned from the prior art as a whole, but when the ‘improvement’ is technology-independent and the combination of references results in a product or process that is more desirable”).
Furthermore, while Liu discloses using “unsupervised learning” for its pretraining task, Liu still requires supervised, labeled data for its downstream fine-tuning tasks (Liu, ¶21 “the pretrained mapping layers can be tuned with a task-specific layer to perform a specific task using supervised learning,” one of ordinary skill in the art would also recognize that a labeled “dev set” is implicit when training based on accuracy: ¶54 “continuing with tuning iterations until a stopping condition is reached, e.g., the model converges, achieves a threshold accuracy on a test data set,” ¶74 “the most accurate task-specific model was picked based on its performance on the dev set”). A reason that would prompt one of ordinary skill in the art to combine the references Liu and Cunningham is that Cunningham discloses a training method using a “log-likelihood” metric which is compatible with unlabeled or unsupervised learning (Cunningham, ¶47 “The EM algorithm [12] is based on distance computation. It can be seen as a generalization of clustering based on computing a mixture of probability distributions. It works by successively improving the solution found so far. The algorithm stops when the quality of the current solution becomes stable. The quality of the current solution is measured by a statistical quantity called log-likelihood (llh),” Fig. 2B, 228, ¶111 “Block 228 of FIG. 2B performs the EM algorithm with different numbers of clusters keeping track of log-likelihood and the total number of parameters. Akaike's Information Criteria combines these two parameters, wherein the highest AIC is the best model,” one of ordinary skill in the art would recognize that a “log-likelihood” metric which generalizes “clustering based on computing a mixture of probability distributions” of data points can operate even if the data points are unlabeled). Liu’s own disclosure implies that solutions are needed when faced with a lack of training data (Liu, ¶18 “There are many machine learning tasks for which there is a relative lack of training data”), hence one of ordinary skill in the art would be motivated to combine Liu’s accuracy-based fine-tuning shots with the log-likelihood metric of Cunningham. Additionally, Jakubowski further cites how its cluster-scoring method can be improved in future studies (Jakubowski, Page 106, Col. 1, ¶1 “The Wasserstein distance provides a notion of distance between two such persistence diagrams, and hence a measure of similarity between different point clouds and their underlying spaces,” Page 111, Col. 1-2 “Our method of taking the Wasserstein norm of a persistence diagram is rather crude”), which may prompt one of ordinary skill in the art to seek the known “log-likelihood” scoring method as disclosed in Cunningham to cure the self-identified deficiency of Jakubowski.
On page 15 of the Remarks, Applicant asserts that “Liu does not disclose ‘fine tuning, by the fine-tuning computer program, the model parameters in response to a specified threshold or maximum number of fine tuning shots not being met.’” Examiner respectfully disagrees. While Liu teaches “fine tuning, by the fine-tuning computer program, the model parameters in response to a specified threshold…” and fails to explicitly teach the element wherein “the stopping conditions are based on a number of fine tuning shots being executed,” under the claim’s broadest reasonable interpretation, the claim limitation is given its plain meaning in light of the specification, which suggests that the items “specified threshold or a maximum number of fine tuning shots not being met” are interpreted to be disjunctive (Fig. 2 – 245 and Present Specification, ¶55 “If the unsupervised class separability metric Li does not meet the specified threshold, or if a threshold is not available due to lack of a set of labeled data, in step 245, the fine-tuning computer program may determine if the maximum number of fine-tuning shots have been performed”). This suggests that fine tuning is in response to either the specified threshold, the number of fine-tuning shots, or both, but does not necessarily require both (MPEP § 2111.01(I) “words of the claim must be given their plain meaning, unless such meaning is inconsistent with the specification”).
On page 15 of the Remarks, Applicant asserts that the proposed combination of Liu, Jakubowski, and Cunningham fails to “disclose ‘applying, by the fine-tuning computer program, parameters from the fine tuning shot that achieved a highest unsupervised class separability metric to the large language model.’” Examiner respectfully disagrees. While Liu only teaches this limitation when applied to a highest accuracy metric using supervised fine-tuning instead of “a highest unsupervised class separability metric” (see claim 1’s rejection below), one of ordinary skill in the art may combine the references Liu and Cunningham by substituting the accuracy-based metric of Liu’s fine tuning with the log-likelihood metric, or the “unsupervised class separability metric from a log-likelihood” of Cunningham’s training in order to allow Liu’s fine-tuning to operate in an unsupervised manner when faced with a lack of labeled training data (see the reasoning as explained above). Cunningham cures the deficiencies of Liu, and the combination of Liu, Jakubowski, and Cunningham discloses all the elements of amended independent claim 1.
In consideration of these conclusions, the previous rejections under 35 U.S.C. 103 still stand for independent claims 1, 8, and 15, and their associated dependent claims 2-6, 9-13, and 15-20, respectively.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1:
Step 1 – Claim 1 is directed to a method: A method for learning…
Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mathematical concepts (see MPEP 2106.04(a)(2)(I)):
learning from unlabeled data and fine-tuning of models using persistent homology… To learn from unlabeled data and fine-tune models “using persistent homology” involves performing algebraic topology and linear algebra to analyze the geometric shape and structure of complex data, which are mathematical concepts. Furthermore, to perform “fine-tuning of models” involves performing an algorithm, for example gradient descent, to adjust the weights and biases of the model according to a calculated loss, which is a mathematical concept. Hence “learning from unlabeled data and fine-tuning of models using persistent homology” is a mathematical concept.
performing… text embedding on the dataset… To perform “text embedding” on a dataset involves performing linear algebra to convert text into high-dimensional vector representations, which is a mathematical concept. Hence “performing… text embedding on the dataset” is a mathematical concept.
generating… a persistent diagram of 0-Homology Group of an embedding manifold for the text embedding, wherein the embedding manifold is from the embedding part of the model… To generate “a persistent diagram of 0-Homology group of an embedding manifold for the text embedding” involves performing algebraic topology and linear algebra to track the stability of connected components in an “embedding manifold for the text embedding,” which is a mathematical concept. Furthermore, an embedding manifold that “is from the embedding part of the model” involves calculating a lower-dimensional space for data generated by “the embedding part of the model” that preserves the geometric structure of original higher dimensional data, which is a mathematical concept. Hence, “generating… a persistent diagram of 0-Homology Group of an embedding manifold for the text embedding, wherein the embedding manifold is from the embedding part of the model” is a mathematical concept.
fitting… a probabilistic model comprising mixing components computed using a model selection criterion… To fit “a probabilistic model comprising mixing components” involves performing an algorithm, for example the Expectation-Maximization algorithm, to compute the probability that data points fall within the various “mixing components” of a probabilistic model, which is a mathematical concept. Furthermore, “computed using a model selection criterion” involves calculating a value which measures the best model to select for a task, which is a mathematical concept. Hence, “fitting… a probabilistic model comprising mixing components computed using a model selection criterion” is a mathematical concept.
generating… an unsupervised class separability metric from a log-likelihood of 0-homology group points conditioned on the probabilistic model… To generate “an unsupervised class separability metric from a log-likelihood of 0-homology group points” is to calculate a log-likelihood, which involves using a likelihood function and a log transformation to calculate how well a model explains the “0-homology group points,” which is a mathematical concept. Furthermore, “conditioned on the probabilistic model” involves using the probabilistic model to assign data points to clusters or components based on calculated probabilities, which is a mathematical concept. Hence “generating… an unsupervised class separability metric from a log-likelihood of 0-homology group points conditioned on the probabilistic model” is a mathematical concept.
determining… whether the unsupervised class separability metric meets a prespecified threshold or not… To determine “whether the unsupervised class separability metric meets a prespecified threshold” is to calculate whether a metric is greater or less than a threshold value, which is a mathematical concept. Hence, “determining… whether the unsupervised class separability metric meets a prespecified threshold or not” is a mathematical concept.
fine tuning… the model parameters in response to a specified threshold or maximum number of fine tuning shots not being met… To perform “fine tuning” on “the model parameters” involves performing an algorithm, for example gradient descent, to modify the weights and biases of a model according to a calculated loss, which is a mathematical concept. Furthermore, to perform fine tuning “in response to a specified threshold or maximum number of fine tuning shots” is to calculate whether or not a threshold value or a maximum number of fine tuning shots are exceeded, which is a mathematical concept. Hence, “fine tuning… the model parameters in response to a specified threshold or maximum number of fine tuning shots not being met” is a mathematical concept.
Step 2A, Prong 2 – The following limitations are additional elements that fail to integrate the abstract idea into a practical application:
by a fine-tuning computer program… a computer program used as a mere tool to apply (e.g., performing, generating, fitting, determining, or fine tuning) an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)).
receiving… a dataset from a dataset database… “receiving” a dataset amounts to insignificant extra-solution activity of data gathering that does not add a meaningful limitation to the “method for learning” (see MPEP 2106.05(g)).
receiving… model parameters of an embedding part of a large language model… “receiving” model parameters amounts to insignificant extra-solution activity of data gathering that does not add a meaningful limitation to the “method for learning” (see MPEP 2106.05(g)).
applying… parameters from the fine tuning shot that achieved a highest unsupervised class separability metric to the large language model… “applying” model parameters from a particular fine tuning shot amounts to insignificant extra-solution activity of data outputting that does not add a meaningful limitation to “the method for learning” (see MPEP 2106.05(g)).
Step 2B – These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they provide nothing more than mere instructions to implement an abstract idea on a generic computer (MPEP 2106.05(f)) or only amount to data gathering or outputting (MPEP 2106.05(g)) without significantly more. These limitations, taken either alone or in combination, fail to provide an inventive concept. Thus, the claim is not patent eligible.
Claims 2-6 recite limitations which further narrow the abstract ideas of claim 1 by specifying more details of the mathematical concepts that occur:
Regarding claim 2, specifying wherein the dataset comprises logs that are generated by switches and/or routers in a network of electronic devices in this manner does not overcome the rejection of claim 1 as modifying the dataset does not make the abstract ideas of claim 1 to not be mathematical concepts.
Regarding claim 3, this claim further limits the abstract idea of claim 1 to be based on a mathematical concept: wherein the probabilistic model is trained on H0 points using an Expectation Maximization method. To use an Expectation Maximization method is a mathematical concept.
Regarding claim 4, this claim further limits the abstract idea of claim 1 to be based on a mathematical concept: performing… preprocessing of raw data in the dataset, wherein the preprocessing comprises text-to-text conversion of the raw data. To perform preprocessing comprising text-to-text conversion of raw data involves using normalization or vectorization techniques, which are mathematical concepts.
Regarding claim 5, specifying wherein the model parameters comprise weights and biases of each neuron in layers of the large language model in this manner does not overcome the rejection of claim 1 as modifying the parameters does not make the abstract ideas of claim 1 to not be mathematical concepts.
Regarding claim 6, specifying wherein the probabilistic model comprises a Gaussian Mixture Model, and the model selection criterion comprises a Bayesian information criterion in this manner does not overcome the rejection of claim 1 as modifying the probabilistic model and the model selection criterion does not make the abstract ideas of claim 1 to not be mathematical concepts.
Regarding claim 8:
Step 1 – Claim 8 is directed to a system: A system comprising…
Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mathematical concepts (see MPEP 2106.04(a)(2)(I)): see claim 1 above.
Step 2A, Prong 2 – The following limitations are additional elements that fail to integrate the abstract idea into a practical application:
a logs/text database comprising a dataset generated by a plurality of source devices… generating a “dataset” amounts to insignificant extra-solution activity of data outputting that does not add a meaningful limitation to the “system” (see MPEP 2106.05(g)). Furthermore, “a logs/text database” used as a mere tool to apply (i.e., storing the dataset required by the mathematical concepts) an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)).
a model parameter database storing a plurality of parameters of an embedding part of a model… a “model parameter database” used as a mere tool to apply (i.e., storing the parameters to be fine-tuned using mathematical concepts) an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)).
an electronic device executing a fine-tuning computer program that receives model parameters of an embedding part of a large language model… an electronic device executing a fine-tuning computer program used as a mere tool to apply (e.g., performing, generating, fitting, determining, or fine tuning) an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)). Furthermore, to receive model parameters amounts to insignificant extra-solution activity of data gathering that does not add a meaningful limitation to the “system” (see MPEP 2106.05(g)).
Claims 8-13 recite a system that parallels the method claims of 1-6, respectively. Therefore, the analysis discussed above with respect to claims 1-6 applies to claims 8-13, respectively. Accordingly, claims 8-13 are rejected based on substantially the same rationale as set forth above with respect to claims 1-6, respectively.
Regarding claim 15:
Step 1 – Claim 15 is directed to a product: A non-transitory computer readable storage medium…
Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mathematical concepts (see MPEP 2106.04(a)(2)(I)): see claim 1 above.
Step 2A, Prong 2 – The following limitations are additional elements that fail to integrate the abstract idea into a practical application:
non-transitory computer readable storage medium, including instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps… a non-transitory computer readable storage medium including instructions stored thereon and one or more processors used as mere tools to apply (e.g., performing, generating, fitting, determining, or fine tuning) an exception are generic elements for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)).
Claims 15-20 recite a non-transitory computer readable storage medium that parallels the method claims of 1-6, respectively. Therefore, the analysis discussed above with respect to claims 1-6 applies to claims 15-20, respectively. Accordingly, claims 15-20 are rejected based on substantially the same rationale as set forth above with respect to claims 1-6, respectively.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 5, 15, 17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 20210326751 A1, hereinafter Liu) in view of Jakubowski et al. (“Topology of Word Embeddings: Singularities Reflect Polysemy,” 2020, hereinafter Jakubowski), and further in view of Cunningham (US 20020129038 A1, hereinafter Cunningham).
Regarding claim 1:
Regarding the limitation a method for learning from unlabeled data and fine-tuning of models using persistent homology, comprising: receiving, by a fine-tuning computer program, a dataset from a dataset database, Liu teaches a method for learning from unlabeled data and fine-tuning of models (¶19 “The term “pretraining,” as used herein, refers to model training on a set of pretraining data to adjust model parameters in a manner that allows for subsequent tuning of those model parameters for one or more specific tasks… the pretraining can involve a self-supervised learning process on unlabeled training data”)… comprising: receiving, by a fine-tuning computer program (¶130 “computer-readable instructions which, when executed by the hardware processing unit, cause the hardware processing unit to: receive input data…,” Fig. 3 – 300-304, ¶34 “Training workflow 300 can include a pretraining stage 302 and a tuning stage 304”), a dataset… (Fig. 1 – 100-102, ¶24 “Natural language processing model 100 can receive pretraining examples 102, which can include documents, sentences, phrases, or other representations of language having various components, such as words and/or tokens”). However, Liu fails to teach using persistent homology and a dataset from a dataset database.
Jakubowski, in the same field of endeavor, teaches using persistent homology (Abstract: “We introduce a topological measure… based on persistent homology”) and a dataset from a dataset database (Page 108, Section 4.1, ¶1 “Assign a total of 8915 instances, extracted from various sources including CNN and ABC, of 100 different polysemous target words,” one of ordinary skill in the art would recognize that major news publications such as “CNN and ABC” utilize complex databases to store their texts and datasets).
Liu further teaches receiving, by the fine-tuning computer program, model parameters of an embedding part (Fig. 1 – 104(1)-(2), ¶24 “The components of the pretraining examples can be processed by embedding layers 104, which include a lexicon encoder 104(1) and a transformer encoder 104(2),” Fig. 3 – 300-304, 306, 326, ¶34 “Training workflow 300 can include a pretraining stage 302 and a tuning stage 304… the pretraining stage can be used to determine pretrained parameters for one or more layers of a machine learning model,” ¶35 “the pretraining stage 302 can utilize unlabeled training data 306… the unlabeled training data can provide an unlabeled corpus of documents in a given natural language. The embedding layers 104 can be pretrained by unsupervised learning to predict tokens in the corpus,” ¶38 “the embedding layers… can be tuned together in tuning stage 304… the pretrained parameters of the embedding layers can be provided in tuning model history 326,” wherein a “tuning stage” which is provided “pretrained parameters of the embedding layers” encompasses receiving… model parameters of an embedding part…) of a large language model (Fig. 1 – 100-108, Fig. 2 – 200, ¶¶24-26 “both the lexicon and transformer encoders operate to produce representations (e.g., vectors) that represent individual words or tokens in a vector space where semantically-similar and/or syntactically-similar words, tokens, sentences, phrases, documents, queries, etc., are relatively close to one another, and less semantically-similar or syntactically-similar words, sentences, tokens, phrases, documents, queries, etc., are relatively further apart. These vectors are also referred to herein as ‘embeddings.’ Lexicon encoder 104(1) can produce first embeddings 106, e.g., a sequence of embedding vectors for each word or token in the pretraining examples 102. An input to the lexicon encoder can be a sequence of tokens of length m… The lexicon encoder can map X into a sequence of one embedding vector for each token… these token embedding vectors are constructed by summing corresponding word, segment, and positional embeddings for each token in the pretraining examples 102… Transformer encoder 104(2) can obtain contextual information for each word or token, e.g., via self-attention, and generate second embeddings 108, e.g., a sequence of context embedding vectors. Self-attention is a mechanism relating positions of tokens within a sentence, paragraph, or document to compute the similarities between those tokens… the transformer encoder is a multilayer bidirectional transformer encoder that is configured to map the first embeddings 106 into the second embeddings 108… the second embeddings, or context embedding vectors, can be used as a shared representation of phrases or sentences across different tasks. The context embedding vectors represent the words or tokens as well as the context within which each word or token appears in an underlying document, query, or other input,” ¶42 “natural language processing models 100 and 200 can be neural networks with multiple layers… one or more mapping layers can include a lexicon encoder… one or more mapping layers can also include a transformer encoder,” wherein a “natural language processing model” which uses a “transformer encoder” with a “self-attention mechanism” encompasses a large language model; see “Documents Considered but Not Relied Upon” below).
Liu further teaches performing, by the fine-tuning computer program, text embedding on the dataset (Fig. 1 – 102-106, ¶25 “Lexicon encoder 104(1) can produce first embeddings 106, e.g., a sequence of embedding vectors for each word or token in the pretraining examples 102”).
Regarding the limitation generating, by the fine-tuning computer program, a persistent diagram of 0-Homology Group of an embedding manifold for the text embedding, wherein the embedding manifold is from the embedding part of the model, Liu teaches by the fine-tuning computer program (¶130, Fig. 3 – 300, 304, ¶34 “Training workflow 300 can include… a tuning stage 304”). However, Liu fails to teach generating… a persistent diagram of 0-Homology Group of an embedding manifold for the text embedding, wherein the embedding manifold is from the embedding part of the model…
Jakubowski teaches generating… a persistent diagram of 0-Homology Group of an embedding manifold for the text embedding (Page 108, Col. 2, Section 4.1, ¶2 “The training set provided comprises 65M occurrences of 127151 different words. We use this corpus to train our own vector representations,” Page 110, Section 4.4, Col. 1, ¶1 “Our hypothesis that the word space is a manifold pinched at polysemous words… The different clusters of the neighbourhood cloud obtained in this way are taken to represent different meanings of the target word,” Page 107, Fig. 5, Caption: “An idealized picture of the word space W near “mole”: four regions of the meaning manifold are glued together,” Page 107, Col. 1, ¶1 “the word space W… is at best a pinched manifold,” Page 107, Col. 2, Section 3.2, ¶1 “we can distinguish a singular point of a pinched manifold… by counting the connected components of a small punctured neighbourhood of the point,” wherein the measure of “connected components” in a topological space encompasses 0-Homology Group, ¶3 “Fix a word vector embedding, a target word w… The topological polysemy TPSn(w) of w with respect to our fixed word vector embedding… is computed as follows… Consider the punctured neighbourhood Nn(w) consisting of the n closest neighbours of w… Compute the degree zero persistence diagram of Nn`(w),” Page 105, Col. 1, Fig. 4, Caption: “An example of a persistence diagram, summarizing the persistent homology of some point cloud”), wherein the embedding manifold is from the embedding part of the model (Page 106, Col. 2, ¶1 “transformer based models that exploit massive data sets have been used to produce contextualised word embeddings,” Page 106, Col. 106, Section 3.1, ¶1 “the manifold hypothesis postulates that, in general, real world data tends to live on a small-dimensional submanifold of the vector space in which it is represented… For word vector embeddings, the ambient space… typically has dimension n in the range 50 ≤ n ≤ 300. The hypothesis states that word vectors in fact lie on, or are densely distributed around, a submanifold… of much smaller dimension,” wherein a “submanifold” of “word vector embeddings” produced by “transformer based models” encompasses wherein the embedding manifold is from the embedding part of the model).
Liu and Jakubowski are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the persistent homology, generating of the persistent diagram of 0-Homology Group, embedding manifold, and 0-homology group points of Jakubowski with the methodology of Liu. The motivation to do so is to better determine the meaning of text data (Jakubowski, Page 110, Col. 2, Section 5, ¶1 “it is more accurate and helpful to view the space of word embeddings as a pinched manifold. We introduce a topological measure of polysemy that correlates well with the number of meanings of a word”).
Regarding the limitation fitting, by the fine-tuning computer program, a probabilistic model comprising mixing components computed using a model selection criterion, Liu teaches by the fine-tuning computer program (¶130, Fig. 3 – 300, 304, ¶34 “Training workflow 300 can include… a tuning stage 304”). However, Liu fails to teach fitting… a probabilistic model comprising mixing components computed using a model selection criterion.
Cunningham, in the same field of endeavor, teaches fitting… a probabilistic model comprising mixing components computed using a model selection criterion (¶102 “Model selection involves deciding which of various possible Gaussian Mixture Models are suitable for use with a given data set,” ¶109 “Model selection using Akaike's Information Criteria,” ¶110 “It is necessary to select the optimum number of clusters for the model. Too few clusters, and the model is a poor fit to the data. Too many clusters, and the model does not perform well when generalized to new data,” Fig. 2B – 228, ¶111 “Block 228 of FIG. 2B performs the EM algorithm with different numbers of clusters… the highest [Akaike's Information Criteria] is the best model,” wherein “number of clusters for the model” encompasses mixing components).
Regarding the limitation generating, by the fine-tuning computer program, an unsupervised class separability metric from a log-likelihood of 0-homology group points conditioned on the probabilistic model, Liu teaches by the fine-tuning computer program (¶130, Fig. 3 – 300, 304, ¶34 “Training workflow 300 can include… a tuning stage 304”). However, Liu fails to teach generating… an unsupervised class separability metric from a log-likelihood of 0-homology group points conditioned on the probabilistic model.
Jakubowski teaches 0-homology group points (Page 106, Col. 1, ¶2 “we will concentrate on persistent homology in degree i = 0,” Page 105, Col. 1, Section 2.2, ¶1 “Topological data analysis (TDA) is an instrument for extracting topological information from a point cloud, that is a finite set of vectors W0,” Page 105, Col. 1, Fig. 3 depicts 0-homology group points, Col. 2, Fig. 4, Caption: “An example of a persistence diagram, summarizing the persistent homology of some point cloud”). However, Jakubowski fails to teach generating… an unsupervised class separability metric from a log-likelihood of 0-homology group points conditioned on the probabilistic model.
Cunningham teaches generating… an unsupervised class separability metric from a log-likelihood (¶47 “The EM algorithm… is based on distance computation. It can be seen as a generalization of clustering based on computing a mixture of probability distributions. It works by successively improving the solution found so far. The algorithm stops when the quality of the current solution becomes stable. The quality of the current solution is measured by a statistical quantity called log-likelihood (llh),” Fig. 2A – 202, 206, ¶56 “Block 202 is a decision… which is performed while the change in log-likelihood llh is greater than E… Upon completion of the loop, control transfers to Block 206 that produces the output… containing the updated mixture parameters with the highest log-likelihood,” wherein “the highest log-likelihood” encompasses an unsupervised class separability metric from a log-likelihood) of clusters conditioned on the probabilistic model (Fig. 2B – 228, ¶111 “Block 228 of FIG. 2B performs the EM algorithm with different numbers of clusters keeping track of log-likelihood and the total number of parameters. Akaike's Information Criteria combines these two parameters, wherein the highest AIC is the best model”).
Regarding the limitation determining, by the fine-tuning computer program, whether the unsupervised class separability metric meets a prespecified threshold or not, Liu teaches determining, by the fine-tuning computer program, whether an accuracy metric meets a prespecified threshold or not (Fig. 3 – 338, ¶39 “The next tuning iteration can proceed by retrieving the previous model 338 from the tuning model history and continuing with tuning iterations until a stopping condition is reached, e.g., the model… achieves a threshold accuracy on a test data set”). However, Liu fails to teach the unsupervised class separability metric.
Cunningham teaches the unsupervised class separability metric (¶56 “the highest log-likelihood”).
Liu further teaches and fine tuning, by the fine-tuning computer program, the model parameters in response to a specified threshold or maximum number of fine tuning shots not being met (¶39 “continuing with tuning iterations until a stopping condition is reached, e.g., the model converges, achieves a threshold accuracy on a test data set, a training budget is exhausted,” please note that a specified threshold or maximum number of fine tuning shots not being met is interpreted to be read disjunctively; Liu teaches the former specified threshold).
Regarding the limitation and applying, by the fine-tuning computer program, parameters from the fine tuning shot that achieved a highest unsupervised class separability metric to the large language model, Liu teaches and applying, by the fine-tuning computer program, parameters from the fine tuning shot that achieved a highest accuracy to the large language model (¶74 “The model was fine-tuned for up to 10 epochs with the provided task-specific training set and the most accurate task-specific model was picked based on its performance on the dev set”). However, Liu fails to teach a highest unsupervised class separability metric…
Cunningham teaches an unsupervised class separability metric (¶56 “the highest log-likelihood”).
Liu, Jakubowski, and Cunningham are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the persistent homology, persistent diagram of 0-Homology Group, embedding manifold, and 0-homology group points of Jakubowski and the probabilistic model comprising mixing components using a model selection criterion and unsupervised class separability metric from a log-likelihood of Cunningham with the methodology of Liu. The motivation to do so is to better determine the meaning of text data (Jakubowski, Page 110, Col. 2, Section 5, ¶1 “it is more accurate and helpful to view the space of word embeddings as a pinched manifold. We introduce a topological measure of polysemy that correlates well with the number of meanings of a word”) and to allow a training algorithm to “perform in a more robust and reproducible manner” (Cunningham, ¶22-23).
Regarding claim 3, Liu in view of Jakubowski and further in view of Cunningham teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated).
Regarding the limitation wherein the probabilistic model is trained on H0 points using an Expectation Maximization method, Jakubowski teaches H0 points (Page 106, Col. 1, ¶2 “we
will concentrate on persistent homology in degree i = 0,” Page 105, Col. 1, Section 2.2, ¶1 “Topological data analysis (TDA) is an instrument for extracting topological information from a point cloud, that is a finite set of vectors W0,” Page 105, Col. 1, Fig. 3 depicts H0 points). However, Jakubowski fails to teach wherein the probabilistic model is trained on H0 points using an Expectation Maximization method.
Cunningham teaches wherein the probabilistic model is trained on data using an Expectation Maximization method (¶175 “The data is accessed from a database, and then an Expectation-Maximization (EM) algorithm is performed in the computer-implemented data mining system to create the Gaussian Mixture Model for the accessed data”).
Liu, Jakubowski, and Cunningham are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the H0 points of Jakubowski and the Expectation Maximization method of Cunningham with the methodology of Liu. The motivation to do so is to better determine the meaning of text data (Jakubowski, Page 110, Col. 2, Section 5, ¶1 “it is more accurate and helpful to view the space of word embeddings as a pinched manifold. We introduce a topological measure of polysemy that correlates well with the number of meanings of a word”) and to allow a training algorithm to “perform in a more robust and reproducible manner” (Cunningham, ¶22-23).
Regarding claim 5, Liu in view of Jakubowski and further in view of Cunningham teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated).
Liu teaches wherein the model parameters comprise weights and biases of each neuron in layers of the large language model (¶15 “neural networks, use layers of nodes,” ¶16 “The inputs to a given node can be multiplied by a corresponding weight value for an edge between the input and the node. In addition, nodes can have individual bias values… The term “parameters” when used without a modifier is used herein to refer to learnable values such as edge weights and bias values that can be learned by training,” Fig. 1 – 104(1)-(2), ¶24 “embedding layers 104, which include a lexicon encoder 104(1) and a transformer encoder 104(2),” ¶27 “pretraining can be used to adjust the parameters of the… transformer encoder 104(2), and/or lexicon encoder 104(1)”).
Regarding claim 15:
Regarding the limitation a non-transitory computer readable storage medium, including instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving a dataset from a dataset database, Liu teaches a non-transitory computer readable storage medium, including instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising (Fig. 6 – 601-602, ¶92 “Generally, the devices… may have respective processing resources 601 and storage resources 602… The storage resources can include… memory devices,” ¶130 “a hardware processing unit and a storage resource storing computer-readable instructions which, when executed by the hardware processing unit, cause the hardware processing unit to: receive input data…”): receiving a dataset (Fig. 1 – 100-102, ¶24) … However, Liu fails to teach receiving a dataset from a dataset database.
Jakubowski teaches a dataset from a dataset database (Page 108, Section 4.1, ¶1 as explained above with respect to claim 1).
Claims 15, 17, and 19 recite a non-transitory computer readable storage medium that parallels the method claims of 1, 3, and 5, respectively. Therefore, the analysis discussed above with respect to claims 1, 3, and 5 applies to claims 15, 17, and 19, respectively. Accordingly, claims 15, 17, and 19 are rejected based on substantially the same rationale as set forth above with respect to claims 1, 3, and 5, respectively.
Claims 2 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Jakubowski and further in view of Cunningham, and further in view of Ward et al. (US 10303516 B1, hereinafter Ward).
Regarding claim 2, Liu in view of Jakubowski and further in view of Cunningham teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated).
Regarding the limitation wherein the dataset comprises logs that are generated by switches and/or routers in a network of electronic devices, Liu teaches a network of electronic devices (Fig. 6 - 600, ¶90 “system 600 includes a client device 610, a server 620, a server 630, and a client device 640, connected by one or more network(s) 650”). However, Liu fails to teach wherein the dataset comprises logs that are generated by switches and/or routers in a network of electronic devices.
Ward, in the same field of endeavor, teaches wherein the dataset comprises logs that are generated by switches and/or routers in a network (Fig. 1 – 102, Col. 8, Lines 2-5 “the network devices 102 can transmit electronic messages for use in managing computing resources… all at once or streaming over a period of time,” Col. 8, Lines 10-11 “network devices 102 may include local area network devices, such as routers… switches”).
Liu and Ward are analogous to the claimed invention as all are from the same field of endeavor of receiving and processing data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the switches and/or routers of Ward with the methodology of Liu. The motivation to do so is to increase efficiency and identify hidden relationships while analyzing data from many devices (Ward, Col. 10, Lines 54-56 “high value analytics can be applied to identify hidden relationships and drive increased efficiencies”).
Claim 16 recites a non-transitory computer readable storage medium that parallels the method claim of 2. Therefore, the analysis discussed above with respect to claim 2 applies to claim 16. Accordingly, claim 16 is rejected based on substantially the same rationale as set forth above with respect to claim 2.
Claims 4 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Jakubowski and further in view of Cunningham, and further in view of Gholizadeh et al. (“A Novel Method Of Extracting Topological Features From Word Embeddings,” 2020, hereinafter Gholizadeh).
Regarding claim 4, Liu in view of Jakubowski and further in view of Cunningham teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated).
Regarding the limitation further comprising: performing, by the fine-tuning computer program, preprocessing of raw data in the dataset, wherein the preprocessing comprises text-to-text conversion of the raw data, Liu teaches by the fine-tuning computer program (¶130, Fig. 3 – 300, 304, ¶34 “Training workflow 300 can include… a tuning stage 304”). However, Liu fails to teach further comprising: performing… preprocessing of raw data in the dataset, wherein the preprocessing comprises text-to-text conversion of the raw data.
Gholizadeh, in the same field of endeavor, teaches further comprising: performing… preprocessing of raw data in the dataset, wherein the preprocessing comprises text-to-text conversion of the raw data (Page 3, Section 2.1, ¶1 “Like any other text mining method, standard pre-processing possibly including lemmatization, removing stop words and if necessary lowercasing will be applied to the text”).
Liu and Gholizadeh are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the text-to-text conversion of the raw data of Gholizadeh with the methodology of Liu. The motivation to do so is to “outperform conventional text mining features” (Gholizadeh, Page 9, Section 5, ¶1).
Claim 18 recites a non-transitory computer readable storage medium that parallels the method claim of 4. Therefore, the analysis discussed above with respect to claim 4 applies to claim 18. Accordingly, claim 18 is rejected based on substantially the same rationale as set forth above with respect to claim 4.
Claims 6 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Jakubowski and further in view of Cunningham, and further in view of Gkortsas et al. (US 20220146705 A1, hereinafter Gkortsas).
Regarding claim 6, Liu in view of Jakubowski and further in view of Cunningham teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated).
Regarding the limitation wherein the probabilistic model comprises a Gaussian Mixture Model, and the model selection criterion comprises a Bayesian information criterion, Cunningham teaches wherein the probabilistic model comprises a Gaussian Mixture Model, and the model selection criterion comprises an Akaike’s Information Criterion (¶102 “Model selection involves deciding which of various possible Gaussian Mixture Models are suitable for use with a given data set,” ¶109 “Model selection using Akaike's Information Criteria,”). However, Cunningham fails to teach and the model selection criterion comprises a Bayesian information criterion.
Gkortsas, in the same field of endeavor, teaches a Bayesian information criterion (Fig. 5 – 230, ¶53 “the processing of 230 can employ a method where the optimal number of facies (clusters) is based on the repeatability or consistency of the clustering results,” ¶54 “the processing of 230 can employ Bayesian Information Criterion (BIC) to determine the quantity or number n of facies”).
Liu, Cunningham, and Gkortsas are analogous to the claimed invention as all are from the same field of endeavor of receiving and processing data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the Gaussian Mixture Model of Cunningham and the Bayesian information criterion of Gkortsas with the methodology of Liu. The motivation to do so is to allow a training algorithm to “perform in a more robust and reproducible manner” (Cunningham, ¶22-23) and to automatically classify data without human input (Gkortsas, ¶42 “the number or quantity n of facies used in the classification is determined automatically and the user does not have to give it as an input”).
Claim 20 recites a non-transitory computer readable storage medium that parallels the method claim of 6. Therefore, the analysis discussed above with respect to claim 6 applies to claim 20. Accordingly, claim 20 is rejected based on substantially the same rationale as set forth above with respect to claim 6.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Jakubowski and further in view of Cunningham, and further in view of Sharma et al. (US 20170032222 A1, hereinafter Sharma).
Regarding claim 8:
Regarding the limitation a system, comprising: a logs/text database comprising a dataset generated by a plurality of source devices, Liu teaches a system, comprising… a dataset (Fig. 6 – 600, ¶89 “FIG. 6 shows an example system 600 in which the present implementations can be employed,” Fig. 1 – 102, ¶24 “pretraining examples 102, which can include documents, sentences, phrases, or other representations of language having various components, such as words and/or tokens”) generated by a plurality of source devices (Fig. 6 – 610-611, 620-621, 650, ¶93 “Client device 610 can include a configuration module 611 that can interact with a model training module 621 on server 620… the configuration module can provide certain configuration parameters to the model training module,” Fig. 3 – 300, ¶94 “The configuration parameters can also include… unsupervised or self-supervised learning parameters and/or data sources… The model training module 621 uses these training configuration parameters to perform model training functionality… the model training module can perform training workflow 300 (FIG. 3) based on the training configuration parameters. As just one example, the unsupervised learning data sources can include one or more repositories of sentences,” ¶117 “the methods and functionality described herein can be performed on a single computing device and/or distributed across multiple computing devices that communicate over network(s) 650”). However, Liu fails to teach a system, comprising: a logs/text database comprising a dataset generated…
Jakubowski teaches a logs/text database comprising a dataset (Page 108, Col. 2, Section 4.1, ¶1 as described above with respect to claim 1).
Regarding the limitation a model parameter database storing a plurality of parameters of an embedding part of a large language model, Liu teaches a plurality of parameters of an embedding part (¶16 “The term “parameters” when used without a modifier is used herein to refer to learnable values such as edge weights and bias values,” Fig. 1 – 104(1)-(2), ¶24, ¶27 “Errors computed during pretraining can be used to adjust the parameters of the… transformer encoder 104(2), and/or lexicon encoder 104(1),” Fig. 3 – 300-304, 306, 326, ¶¶34-35, 38 as explained above with respect to claim 1) of a large language model (Fig. 1 – 100-108, Fig. 2 – 200, ¶¶24-26, 42 as explained above with respect to claim 1). However, Liu fails to teach a model parameter database storing a plurality of parameters…
Sharma, in the same field of endeavor, teaches a model parameter database storing a plurality of parameters (Fig. 2 – 202, Fig. 3 – 302, 316-318, ¶67 “input module 318 of the image data analysis device 302 may use the received first set of parameters to implement the pre-trained CNN 202. The first set of parameters and the training dataset 201 may be stored in the database 316”).
Liu further teaches an electronic device executing a fine-tuning computer program (¶111 “Processing capability can be provided by one or more hardware processors (e.g., hardware processing units/cores) that can execute computer-readable instructions to provide functionality,” ¶130, Fig. 3 – 300-304, ¶34 “Training workflow 300 can include a pretraining stage 302 and a tuning stage 304”) that receives model parameters of an embedding part of a model (Fig. 1 – 104(1)-(2), ¶24, Fig. 3 – 300-304, 306, 326, ¶34-35, ¶38 all as explained above with respect to claim 1).
Liu further teaches performs text embedding on the dataset (Fig. 1 – 102-106, ¶25).
Liu fails to teach generates a persistent diagram of 0-Homology Group of an embedding manifold for the text embedding, wherein the embedding manifold is from the embedding part of the model. However, Jakubowski teaches this limitation (Page 107, Fig. 5, Caption, Page 107, Col. 1, ¶1, Page 107, Col. 2, Section 3.2, ¶1-3, Page 105, Col. 1, Fig. 4, Caption, Page 108, Col. 2, Section 4.1, ¶2, Page 110, Section 4.4, Col. 1, ¶1, Page 106, Col. 2, ¶1, Page 106, Col. 2, Section 3.1, ¶1, all as explained above with respect to claim 1).
Liu fails to teach fits a probabilistic model comprising mixing components computed using a model selection criterion. However, Cunningham teaches this limitation (¶47, Fig. 2A – 202, 206, ¶56, Fig. 2B – 228, ¶111 all as explained above with respect to claim 1).
Regarding the limitation generates an unsupervised class separability metric from a log-likelihood of 0-homology group points conditioned on the probabilistic model, Jakubowski teaches 0-homology group points (Page 106, Col. 1, ¶2, Page 105, Col. 1, Section 2.2, ¶1, Page 105, Col. 1, Fig. 3 and Col. 2, Fig. 4, Caption, all as explained above with respect to claim 1). However, Jakubowski fails to teach generates an unsupervised class separability metric from a log-likelihood of 0-homology group points conditioned on the probabilistic model.
Cunningham teaches generates an unsupervised class separability metric from a log-likelihood of clusters conditioned on the probabilistic model (¶47, Fig. 2A – 202, 206, ¶56, Fig. 2B – 228, ¶111).
Regarding the limitation determines whether the unsupervised class separability metric meets a prespecified threshold or not, Liu teaches determines whether an accuracy metric meets a prespecified threshold or not (Fig. 3 – 338, ¶39). However, Liu fails to teach the unsupervised class separability metric.
Cunningham teaches the unsupervised class separability metric (¶56).
Liu further teaches fine tunes the model parameters in response to a specified threshold or maximum number of fine tuning shots not being met (¶39, as explained above with respect to claim 1).
Regarding the limitation and applying parameters from the fine tuning shot that achieved a highest unsupervised class separability metric to the large language model, Liu teaches and applying parameters from the fine tuning shot that achieved a highest accuracy to the large language model (¶74). However, Liu fails to teach a highest unsupervised class separability metric…
Cunningham teaches an unsupervised class separability metric (¶56).
Liu, Sharma, Jakubowski, and Cunningham are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the parameter database of Sharma, the persistent homology, persistent diagram of 0-Homology Group, embedding manifold, and 0-homology group points of Jakubowski, and the probabilistic model comprising mixing components using a model selection criterion and the unsupervised class separability metric from a log-likelihood of Cunningham with the system of Liu. The motivation to do so is to increase model accuracy after fine-tuning (Sharma, Fig. 4 – 400-408, ¶71 “method 400 also advantageously uses both color and depth modalities together during training that may lead to increased object recognition accuracy”), better determine the meaning of text data (Jakubowski, Page 110, Col. 2, Section 5, ¶1 “it is more accurate and helpful to view the space of word embeddings as a pinched manifold. We introduce a topological measure of polysemy that correlates well with the number of meanings of a word”), and to allow a training algorithm to “perform in a more robust and reproducible manner” (Cunningham, ¶22-23).
Claims 9-13 recite a system that parallels the method claims of 2-6, respectively. Therefore, the analysis discussed above with respect to claims 2-6 applies to claims 9-13, respectively. Accordingly, claims 9-13 are rejected based on substantially the same rationale as set forth above with respect to claims 2-6, respectively.
Documents Considered but Not Relied Upon
“Attention Is All You Need,” 2017, by Vaswani et al., hereinafter Vaswani. Vaswani is a frequently cited publication for introducing the Transformer architecture, which serves as the foundation for many modern LLMs known in the art. Vaswani’s architecture is described as having an encoder and decoder (see Page 2, Section 3) which use a multi-head self-attention mechanism (see Pages 4-5, Section 3.2.3), similar to the neural network architecture described in Liu.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM M LEE whose telephone number is (571)272-4761. The examiner can normally be reached Mon-Fri. 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WILLIAM M LEE/
Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145