Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This communication is responsive to the amendment filed on 03/16/2026.
Status of claims:
Claim 14 is canceled.
Claim 21 is newly added.
Claims 1-2, 5-6, 12-13, 15, 17-20 are amended.
claims 1-13 and 15-21 are presented for examination.
Response to Arguments
Applicant’s argument regarding to the amended claims have been considered but are moot because the new ground of rejection.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-13 and 15-21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claims 1, 18 and 20:
Step 1:
Claims are directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter.
Step 2A, Prong One:
The limitations “generating…; clustering…; assigning…weighting…”, are processes that under its broadest reasonable interpretation, covers a mental process as a form of evaluation or judgement, but for the recitation of generic computer components. That is nothing in the claim element precludes the steps from practically being performed in a human mind. For example, the limitations “generating…; clustering…; assigning…weighting…”, in the context of the claim encompasses one can manually or mentally with the aid of pen and paper determines a set of attributes from the weighted embeddings for the document.
If a claim limitation, under its broadest reasonable interpretation, covers
performance of the limitation in the mind but for the recitation of generic computer
components, then it falls within the "Mental Processes" grouping of abstract ideas.
Accordingly, the claim recites an abstract idea.
Step 2A, Prong Two: Integrated into a Practical Application
This judicial exception is not integrated into a practical application. The claim recites the additional elements “
“selecting… ”, amount to data gathering steps which is considered to be insignificant extra-solution activity. (See MPEP 2106.05(g).
“performing…;after selecting…., training…; outputting…” represent(s) an extra solution activity because it is a mere nominal or tangential addition to the claim, a mere generic transmission and presenting of collected and analyzed data. (See MPEP 2106.05(g)). Also is a mere implementation using a computer. It is at best generally linking the abstract idea to a particular field of use or technological environment of machine learning (see MPEP 2106.05(h).
Step 2B: Claim provides an Inventive Concept
“selecting… ”. These are identified as insignificant extra-solution activity above when re-evaluated this element is well-understood, routine, and conventional as evidenced by the court cases in MPEP 2106.05(d)(II), "i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); … OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network);" and thus remains insignificant extra-solution activity that does not provide significantly more.
“performing…;after selecting…., training…; outputting…”. This is identified as insignificant extra-solution activity above when re-evaluated this element is well-understood, routine, and conventional as evidenced by the court cases in MPEP 2106.05(d)(II), "iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334; i. … transmitting data over a network, …Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); … OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350,
1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)”.
“computer processor, computing device, non-transitory computer readable storage medium”, amount to elements that have been recognized as well-understood, routine, and conventional activity in particular fields, as demonstrate by: relevant court decision: the followings are example of the court decisions demonstrating well-understood, routine and conventional activities, See e.g., MPEP 2106.05(d)(II) and MPEP 2106.05(f)(2): computer readable storage media comprising instructions to implement a method, e.g., see versata Dev. Group, Inc. v SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015).
The conclusions for the mere implementation using a computer are carried over and does not provide significantly more.
The claims as a whole, does not amount to significantly more than the abstract
idea itself. This is because the claims do not affect an improvement to the functioning
of a computer itself; and the claims do not move beyond a general link of the use of an
abstract idea to a particular technological environment.
Accordingly, claims are directed to an abstract idea.
Claims 2-4, recites the limitations represent(s) an extra solution activity because it is a mere nominal or tangential addition to the claim, a mere generic transmission and presenting of collected and analyzed data. (See MPEP 2106.05(g)).
Claim 5, recites the limitation “comparing…’, which can manually or mentally with the aid of pen and paper determines a set of attributes from the weighted embeddings for the document. This is directed to an abstract idea;
“selecting…” represent(s) an extra solution activity because it is a mere nominal or tangential addition to the claim, a mere generic transmission and presenting of collected and analyzed data. (See MPEP 2106.05(g)).
Claims 6-7, recite the limitations, which can manually or mentally with the aid of pen and paper determines a set of clustering. This is directed to an abstract idea.
Claim 8, recites the limitations, which represent(s) an extra solution activity because it is a mere nominal or tangential addition to the claim, a mere generic transmission and presenting of collected and analyzed data. (See MPEP 2106.05(g)).
Claim 9, recites the limitation, which represent(s) an extra solution activity because it is a mere nominal or tangential addition to the claim, a mere generic transmission and presenting of collected and analyzed data. (See MPEP 2106.05(g)).
Claims 8-13, recites the additional limitations, which represent an extra solution activity because it is a mere nominal or tangential addition to the claim, a mere generic transmission and presenting of collected and analyzed data. (See MPEP 2106.05(g)).
Claims 15-17, recite the limitations represent(s) an extra solution activity because it is a mere nominal or tangential addition to the claim, a mere generic transmission and presenting of collected and analyzed data. (See MPEP 2106.05(g)).
Claim 19, recites the limitations represent(s) an extra solution activity because it is a mere nominal or tangential addition to the claim, a mere generic transmission and presenting of collected and analyzed data. (See MPEP 2106.05(g)).
Claim 21, recites the limitation a “determining…”, is a process that, under its broadest reasonable interpretation, covers a mental process as a form of evaluation or judgement, but for the recitation of generic computer components. That is nothing in this element precludes the steps from practically being performed in a human mind. The context of the claim encompasses one can manually or mentally with the aid of pen and paper.
the limitations “manipulating…; associating…”, represent(s) an extra solution activity because it is a mere nominal or tangential addition to the claim, a mere generic transmission and presenting of collected and analyzed data. (See MPEP 2106.05(g)).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 5-7, 9, 18 and 21 are rejected under 35 U.S.C. 103 as being unpatentable Over Keh et al., (US 2024/0086423), hereinafter “Keh”, in view Sastre et al., (US 2023/0274092), hereinafter “Sastre” and further in view of Zhang et al., (US 10,740,825), hereinafter “Zhang”.
Claim 1, Keh discloses a method comprising:
- selecting different cluster sizes for clustering data for a hyperparameter for a cluster size (abstract and par. [0005], selecting a subset of the clusters based on the cluster-specific binding metrics; performing a clustering-based process using the aggregate representations of the set of aptamers so as to generate a set of clusters, wherein at least two of the set of aptamers are assigned to each cluster of the set of clusters; identifying—for each cluster of the set of clusters and for each aptamer assigned to the cluster—an aptamer-specific binding metric corresponding to the aptamer and a specific target; determining—for each cluster of the set of clusters—a cluster-specific binding metric based on the aptamer-specific binding metrics corresponding to the aptamers assigned to the cluster and to specific target; selecting a subset of the set of clusters based on the cluster-specific binding metrics, where the subset is smaller than the set of clusters; and outputting an identification of aptamers corresponding to the selected subset of the at least two of the set of clusters);
- performing training of a model to select the hyperparameter for the cluster size (par. [0005], performing a clustering-based process using the aggregate representations of the set of aptamers so as to generate a set of clusters, wherein at least two of the set of aptamers are assigned to each cluster of the set of clusters; identifying—for each cluster of the set of clusters and for each aptamer assigned to the cluster—an aptamer-specific binding metric corresponding to the aptamer and a specific target; determining—for each cluster of the set of clusters—a cluster-specific binding metric based on the aptamer-specific binding metrics corresponding to the aptamers assigned to the cluster and to specific target; selecting a subset of the set of clusters based on the cluster-specific binding metrics, where the subset is smaller than the set of clusters; and outputting an identification of aptamers corresponding to the selected subset of the at least two of the set of clusters);
- after selecting the hyperparameter, training a set of model parameters for the model using the hyperparameter (par. [0005], performing a clustering-based process using the aggregate representations of the set of aptamers so as to generate a set of clusters, wherein at least two of the set of aptamers are assigned to each cluster of the set of clusters; identifying—for each cluster of the set of clusters and for each aptamer assigned to the cluster—an aptamer-specific binding metric corresponding to the aptamer and a specific target; determining—for each cluster of the set of clusters—a cluster-specific binding metric based on the aptamer-specific binding metrics corresponding to the aptamers assigned to the cluster and to specific target; selecting a subset of the set of clusters based on the cluster-specific binding metrics, where the subset is smaller than the set of clusters; and outputting an identification of aptamers corresponding to the selected subset of the at least two of the set of clusters).
- clustering respective segments in a set of clusters (par. [0005], performing a clustering-based process using the aggregate representations of the set of aptamers so as to generate a set of clusters, wherein at least two of the set of aptamers are assigned to each cluster of the set of clusters; identifying—for each cluster of the set of clusters and for each aptamer assigned to the cluster—an aptamer-specific binding metric corresponding to the aptamer and a specific target; determining—for each cluster of the set of clusters—a cluster-specific binding metric based on the aptamer-specific binding metrics corresponding to the aptamers assigned to the cluster and to specific target; selecting a subset of the set of clusters based on the cluster-specific binding metrics, where the subset is smaller than the set of clusters; and outputting an identification of aptamers corresponding to the selected subset of the at least two of the set of clusters; and par. [0006] for each cluster of the at least two of the set of clusters: identifying a binding-metric difference condition; detecting one or more aptamers for which the bind-metric difference condition is satisfied; and modifying, for each of the one or more aptamers, the aptamer-specific binding metric, wherein the output identifies the one or more aptamers).
However, Keh does not explicitly disclose “generating embeddings in an embedding space for segments of a plurality of documents, wherein the clustering is performed based on a position of respective embeddings in the embedding space and the cluster size for the hyperparameter; assigning a weight for respective clusters in the set of clusters based on how many documents include respective segments in a respective cluster”
Meanwhile, Sastre discloses generating embeddings in an embedding space for segments of a plurality of documents (par. [0006], [0021] and [0055], determine a list of dialogue topics, and each topic is described by a list of key utterances (text segments) to identify a set of utterance embeddings based on the set of utterances in an embedding space containing embeddings and relationships);
- wherein the clustering is performed based on a position of respective embeddings in the embedding space and the cluster size for the hyperparameter; assigning a weight for respective clusters in the set of clusters based on how many documents include respective segments in a respective cluster (par. [0006], clustering the set of utterance embeddings to obtain a plurality of clusters, each cluster comprising one or more utterance embeddings and encoding each document based on a number of times each utterance cluster identifier (ID) appears to obtain an encoded document);
- determining a weight for the cluster for respective embeddings (par. [0069], each unique cluster that appears in that document with the associated cluster weight for that document being the frequency of that utterance type in that document alone, not all the documents);
- weighting the respective embeddings for a document in the plurality of documents with the weight of the cluster (par. [0006] and [0069], each unique cluster that appears in that document with the associated cluster weight for that document being the frequency of that utterance type in that document alone, not all the documents).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective date of the claimed invention to have modified the system of Keh to embed document segments, cluster the embeddings across a corpus, and using cluster-based frequency weighting before classification, in order to improve retrieval precision, thereby enabling scalable organization, and supporting automated topic discovery, making it a powerful technique for both analytical and application-driven tasks.
Neither Keh and Sastre discloses “generate weighted embeddings" and determining a set of attributes from the weighted embeddings for the document.
On the other hand, Zhang discloses generate weighted embeddings (col.9, lines 45-65, generates a weighting factor for each word embedding used to generate the targeting embedding); and
- determining a set of attributes from the weighted embeddings for the document (col.10, lines 22-41, generating a cluster embedding from user embeddings in the latent space, generates a cluster embedding for each cluster. In one embodiment, the embedding module averages all user embeddings within a cluster to generate the cluster embedding, increase weighting of certain keyword embeddings which are common to all user embeddings within the cluster, the clustering module determines two clusters, and for each of the two clusters, the embedding module averages the user embeddings included in each cluster to determine the cluster embedding for cluster embedding for cluster).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh and Sastre to generate weighted embeddings, in order to prioritize more important information and improve model performance for specific tasks, thereby providing a more robust and accurate representation of data.
Claim 2, the combination of Keh, Sastre and Zhang discloses the invention as claimed. In addition, Sastre discloses generating the embeddings comprises: inputting a segment of the document into an encoder (par. [0051], encoder ingests a text segment of a plurality of words); and outputting an embedding in the embedding space based on the segment (par. [0051] and [0055], output a vector of fixed size (two or more words, three or more words, four or more words) in the embedding space, a first embedding vector corresponding to the first utterance and a second embedding vector corresponding to the second utterance may be very close to each other based on the segment).
Claim 5, the combination of Keh, Sastre and Zhang discloses the invention as claimed. In addition, Sastre discloses clustering respective segments comprises: comparing a position of an embedding in the embedding space to positions of one or more clusters (col.12, lines 53-67, generates a score for each cluster based on a comparison of the targeting embedding with each cluster embedding); and selecting a cluster based on the comparing (col.12, line 53-col.13, line 22, the content selection module quantifies the comparison by calculating a cosine similarity between two embeddings).
Claim 6, the combination of Keh, Sastre and Zhang discloses the invention as claimed. In addition, Sastre discloses clustering respective segments comprises: clustering embeddings for the plurality of documents to determine the set of clusters (col.9, lines 31-44, col.10, lines 1-38, cluster embedding is determined for each cluster of the multiple clusters).
Claim 7, the combination of Keh, Sastre and Zhang discloses the invention as claimed. In addition, Sastre discloses the weight of the cluster for respective segments is based on a frequency of occurrence of the respective embeddings in the plurality of documents compared to other clusters in the set of clusters (col.9, lines 12-36, generates a higher weight for words in a title of the content item compared to words in a body text of the content item).
Claim 9, the combination of Keh, Sastre and Zhang discloses the invention as claimed In addition, Sastre discloses wherein a number of clusters in the set of clusters is a setting (par. [0056], specify the number of clusters to generate).
Claim 18, is a non-transitory computer-readable storage medium having stored thereon computer executable instructions for executing the method of claim 1 above. It is rejected under the same rationale.
Claim 20, is an apparatus for performing the method of claim 1 above. It is rejected under the same rationale.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Keh , Sastre, in view of Zhang, and further in view of Gayatri et al., the article entitled “Hierarchical encoders for modeling and interpreting screenplays”, hereinafter “Gayatri”.
Claim 4, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “the document comprises a screenplay for content, and the screenplay includes text based on the content.
Meanwhile, Gayatri discloses the document comprises a screenplay for content, and the screenplay includes text based on the content (page 2, section , hierarchical scene encoders).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh, Sastre and Zhang to use a screenplay for content, in order to organize the plot, characters, and dialogue to guide the production team.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Keh , Sastre, in view of Zhang, and further in view of Church et al., (US 2020/0193093), hereinafter “Church”.
Claim 8, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “a first threshold of frequency that is used to ignore any clusters that occur less than the first threshold, and a second threshold of frequency that is used to ignore any clusters that occur more than the second threshold”.
Meanwhile, Church discloses a first threshold of frequency that is used to ignore any clusters that occur less than the first threshold, and a second threshold of frequency that is used to ignore any clusters that occur more than the second threshold (par. [0139] and claim 1, selecting a number of classes such that each class appears no less than a threshold frequency in the corpus; in n iterations, assigning the words in the block into the number of classes to obtain n assignment matrices, n is an integer larger than 1; given the n assignment matrices, obtaining n input corpuses by replacing words in the corpus with its class identifier).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh, Sastre and Zhang to ignore any clusters that occur more than the second threshold, in order to ensure the statistical inferences are more reliable.
Claims 10-13 are rejected under 35 U.S.C. 103 as being unpatentable over Keh, Sastre, in view of Zhang, and further in view of Lundgaard et al., (US 2023/0039734), hereinafter “Lundgaard”.
Claim 10, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “wherein weighting the respective embeddings using the weight for the cluster comprises: applying the weight for a respective cluster to the respective embedding”.
Meanwhile, Lundgaard discloses wherein weighting the respective embeddings using the weight for the cluster comprises: applying the weight for a respective cluster to the respective embedding (par. [0034], apply weighted averaging and/or average pooling to at least one of the textual embeddings and the image embeddings).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh, Sastre and Zhang to apply the weight for a respective cluster to the respective embedding, in order to allow machine learning models to assign varying levels of importance to different words or data points thereby leading to more accurate, efficient, and nuanced representations of the data.
Claim 11, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “different clusters are associated with different weights based on a frequency of occurrence of the cluster in the plurality of documents compared to other clusters in the set of clusters”.
Meanwhile, Lundgaard discloses different clusters are associated with different weights based on a frequency of occurrence of the cluster in the plurality of documents compared to other clusters in the set of clusters (par. [0037], determine a weighted average over the embedding inputs of two unique training examples, where lambda may be the weight of the average).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention modify the combined system of Keh, Sastre and Zhang to assign different weights to different clusters, in order to allow machine learning models to tune themselves, prioritize important information, and improve performance and efficiency.
Claim 12, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “outputting the set of attributes comprises: using a classifier that classifies the weighted embeddings for the document into the set of attributes for the document.
Meanwhile, Lundgaard discloses outputting the set of attributes comprises: using a classifier that classifies the weighted embeddings for the document into the set of attributes for the document (par. [0034], concatenate the textual embeddings and image embeddings with each other to create a single vector, and classify the combined embeddings).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh, Sastre and Zhang to classify the weighted embeddings, in order to allow models to better focus on the most relevant features for a specific classification task, leading to improved accuracy.
Claim 13, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “outputting the set of the attributes comprises: using a plurality of classifiers that are respectively trained to classify the weighted embeddings for the document into an attribute in a respective type of attribute, wherein the type of attribute is associated with one of the plurality of classifiers.
Meanwhile, Lundgaard discloses outputting the set of the attributes comprises: using a plurality of classifiers that are respectively trained to classify the weighted embeddings for the document into an attribute in a respective type of attribute, wherein the type of attribute is associated with one of the plurality of classifiers (par. [0037]-[0042], determine weighted average over the embedding inputs of two unique training examples, where lambda may be the weight of the average).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh, Sastre and Zhang to classify the weighted embeddings, in order to allow machine learning models to tune themselves, prioritize important information, and improve performance and efficiency.
Claims 3 and 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Keh, Sastre, in view of Zhang, and further in view of Rollings et al., (US 2021/0182328), hereinafter “Rollings”.
Claim 3, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “embedding vector that represents content of the segment a set of dimensions in the embedding space”.
Meanwhile, Rollings discloses the embedding comprises an embedding vector that represents content of the segment a set of dimensions in the embedding space (par. [0059], determines a cluster centroid which is a vector whose number of dimensions is equal to the number of dimensions in the portion embedding vectors of the portions being clustered).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system Keh, Sastre and Zhang to use an embedding vector, in order to capture meaning and relationships of data, thereby enabling more accurate inferences and reducing ambiguity.
Claim 15, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for the claimed performing training in which a model parameter of a classifier that classifies the weighted embeddings into the attributes for the document is adjusted.
Meanwhile, Rollings discloses performing training in which a model parameter of a classifier that classifies the weighted embeddings into the attributes for the document is adjusted (par. [0056] and [0073], generate a portion embedding vector by generating a pre-embedding vector for the portion and adjusting the pre-embedding portion for the vector using an embedding adjustment vector determined for the corpus of documents. Based on the adjusted score for each term, the terms can be ranked and a top-ranked number of terms are selected and stored as the names for the cluster).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh, Sastre and Zhang to perform training in which a parameter of a classifier, in order to optimize model performance on specific tasks or domains, ensuring that the generated word vectors accurately capture the relevant semantic and syntactic relationships of the data.
Claim 16, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “performing training in which a first threshold of frequency that is used to ignore any segments that occur less than the first threshold and a second threshold of frequency that is used to ignore any segments that occur more than the second threshold are adjusted”.
Meanwhile, Rollings discloses performing training in which a first threshold of frequency that is used to ignore any segments that occur less than the first threshold and a second threshold of frequency that is used to ignore any segments that occur more than the second threshold are adjusted (par. [0056] and [0073], generate a portion embedding vector by generating a pre-embedding vector for the portion and adjusting the pre-embedding portion for the vector using an embedding adjustment vector determined for the corpus of documents. Based on the adjusted score for each term, the terms can be ranked and a top-ranked number of terms are selected and stored as the names for the cluster).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh, Sastre and Zhang to perform training in which a first threshold of frequency, in order to optimize model performance on specific tasks or domains, ensuring that the generated word vectors accurately capture the relevant semantic and syntactic relationships of the data.
Claims 17 and 19, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “performing training in which a first parameter for a number of clusters is adjusted; performing training in which a second parameter of a classifier that classifies the weighted embeddings into the attributes for the document are adjusted; and performing training in which a first threshold of frequency that is used to ignore any segments that occur less than the first threshold and a second threshold of frequency that is used to ignore any segments that occur more than the second threshold are adjusted, wherein parameters of an encoder that determines the embeddings are not adjusted”.
Meanwhile, Rollings discloses performing training in which a model parameter of a classifier that classifies the weighted embeddings into the attributes for the document are adjusted (par. [0056] and [0073], generate a portion embedding vector by generating a pre-embedding vector for the portion and adjusting the pre-embedding portion for the vector using an embedding adjustment vector determined for the corpus of documents. Based on the adjusted score for each term, the terms can be ranked and a top-ranked number of terms are selected and stored as the names for the cluster) wherein the method further comprise: performing training in which a first threshold of frequency that is used to ignore any segments that occur less than the first threshold and a second threshold of frequency that is used to ignore any segments that occur more than the second threshold are adjusted, wherein parameters of an encoder that determines the embeddings are not adjusted (par. [0056] and [0073], generate a portion embedding vector by generating a pre-embedding vector for the portion and adjusting the pre-embedding portion for the vector using an embedding adjustment vector determined for the corpus of documents. Based on the adjusted score for each term, the terms can be ranked and a top-ranked number of terms are selected and stored as the names for the cluster).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh, Sastre and Zhang to perform training in which a first parameter for a number of clusters is adjusted, in order to optimize model performance on specific tasks or domains, ensuring that the generated word vectors accurately capture the relevant semantic and syntactic relationships of the data.
Claim 21, the combination of Keh, Sastre and Zhang discloses the invention as claimed, except for “determining a first embedding vector from weighted embeddings of a first document; manipulating the first embedding vector to a second embedding vector in the embedding space by applying a relationship to the first embedding vector; and - associating the second embedding vector with a third embedding vector for a second document to determine a semantic connection between the first document and the second document based on the relationship”.
Meanwhile, Rollings discloses determining a first embedding vector from weighted embeddings of a first document (par. [0056], [0081], [0092], generate a portion embedding vector for each portion. Embedding is a method of converting sets of discrete objects into points within a space and serves to quantify or categorize semantic similarities between linguistic items based on their distributional properties in large samples of language data. Accordingly, the portion embedding vector generated for a portion may represent the semantics of that portion regardless of the language or syntax utilized in the portion. The portion embedding vector generated for a portion may be stored in repository 105 in association with that portion. For example, the document repository system 101 may include or reference a repository of cross-lingual word embedding vectors such as the FastText embeddings provided by Project Muse (Multilingual Unsupervised and Supervised Embeddings). Other types of embeddings may also be utilized without loss of generality, including for example Word2Vec, GloVE, BERT, ELMO or the like). In this manner, regardless of the language of the portion that portion can be converted to a common representation of the semantics of the topics or concepts of that portion; and par. [0057], the topic clustering engine can thus map each portion (or the tokens thereof) to the list of word embedding vectors to produce the portion embedding vector for the portion. The word embedding vector may be of a dimension (e.g., number of integers) that may be user configured or empirically determined. This mapping may be done, for example, by mapping each token (or each of a determined subset of the tokens) of the portion to a vector in the word embedding vectors to determine a vector for each token of the portion, and utilizing that vector to generate the portion embedding vector according to the order the tokens occur in the portion. In one specific embodiment, the topic clustering engine 124 can utilize SIF to generate a portion embedding vector by generating a pre-embedding vector for the portion and adjusting the pre-embedding portion for the vector using an embedding adjustment vector determined for the corpus of documents);
- manipulating the first embedding vector to a second embedding vector in the embedding space by applying a relationship to the first embedding vector (par. [0081], [0092] and [0133], generate a portion embedding vector 315 for each portion 313. For example, the topic clustering engine 324 may include, or reference, a repository of cross-lingual word embedding vectors such as the FastText embeddings provided by Project Muse (Multilingual Unsupervised and Supervised Embeddings). Other types of embeddings may also be utilized without loss of generality. The portion embedding vector 315 generated for a portion 313 may be stored in repository 305 in association with that portion 313; snippet centroid for a cluster 711 can be determined by term ranker 706 based on a raw portion embedding vector for each of the snippets of the cluster 711. This raw portion embedding vector for a snippet may be stored in association with the snippet during the determination of the portion embedding vector for the snippet or, alternatively, may be determined for the snippet from a list of word embedding vectors such that each component of the raw embedding vector is equal to the unweighted average of the corresponding components in the list of the word embedding vectors for the snippet. Thus, based on the raw portion embedding vector for each snippet of the cluster 711, the snippet centroid for the cluster can be determined.); and
- associating the second embedding vector with a third embedding vector for a second document to determine a semantic connection between the first document and the second document based on the relationship (par. [0056], generate a portion embedding vector for each portion. Embedding is a method of converting sets of discrete objects into points within a space and serves to quantify or categorize semantic similarities between linguistic items based on their distributional properties in large samples of language data. Accordingly, the portion embedding vector generated for a portion may represent the semantics of that portion regardless of the language or syntax utilized in the portion…; par. [0057] The topic clustering engine can thus map each portion (or the tokens thereof) to the list of word embedding vectors to produce the portion embedding vector for the portion. The word embedding vector may be of a dimension (e.g., number of integers) that may be user configured or empirically determined. ... In one specific embodiment, the topic clustering engine 124 can utilize SIF to generate a portion embedding vector by generating a pre-embedding vector for the portion and adjusting the pre-embedding portion for the vector using an embedding adjustment vector determined for the corpus of documents; par. [0072] For each extracted term of a linguistic type, the topic clustering engine 124 may generate an embedding vector. …Based on the score for each term, the terms can be ranked and a top-ranked number of terms (e.g., of each linguistic category) may be selected and stored as the names 109 for the cluster (e.g., with a number of names stored not to exceed the storing name number)).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the combined system of Keh, Sastre and Zhang to determine a first embedding vector from weighted embeddings of a first document; manipulate the first embedding vector to a second embedding vector in the embedding space by applying a relationship to the first embedding vector; and associate the second embedding vector with a third embedding vector for a second document to determine a semantic connection between the first document and the second document based on the relationship, in order to align the model’s semantic understanding with your specific needs, thereby improving retrieval and reasoning accuracy, and optimizing performance for cost and speed.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure (see PTO-892).
The Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Loan T. Nguyen whose telephone number is (571) 270-3103. The examiner can normally be reached on Monday from 10:00 am - 6:00 pm, Thursday-Friday from 10:00 am - 2:00 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aleksandr Kerzhner can be reached on (571) 270-1760. The fax phone number for the organization where this application or proceeding is assigned is 571-270-4103. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LOAN T NGUYEN/Examiner, Art Unit 2165
/ALEKSANDR KERZHNER/Supervisory Patent Examiner, Art Unit 2165