Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgement is made of Applicant’s claim to foreign priority under 35 U.S.C. 119 (a) – (d).
Response to Amendment
Applicant’s Amendment, filed January 26, 2026, has been fully considered and entered. Accordingly, Claims 1-20 are pending in this application. Claims 11 and 17-20 have been cancelled. Claims 1 and 12-16 have been amended. Claims 1, 12, and 13 are Independent Claims.
In view of Applicant’s Amendment, the objection to Claim 17 has been withdrawn.
In view of Applicant’s Amendemnt, the Claims 13-16 no longer invoke 35 U.S.C. 112(f).
In view of Applicant’s Amendemnt, the rejections of Claims 11-16 and 20 under 35 U.S.C. 101 have been withdrawn.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 3, 4, 7, 12, 13, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Parakhin (PG Pub. No. 2018/0081992 A1) and further in view of Cheng (PG Pub. No. 2014/0258330 A1) and Manning (“Introduction to Information Retrieval”, Chapter 16.1 “Clustering in information retrieval”, https://nlp.stanford.edu/IR-book/html/htmledition/clustering-in-information-retrieval-1.html, Cambridge University Press, 2008).
Regarding Claim 1, Parakhin discloses a method for retrieving at least one specific set of data among a plurality of data using a specific query, the method being configured to be executed by at least one computer-implemented system, the method comprising:
a classification phase configured to classify the plurality of data into distinct indices, each index of the distinct indices being associated with at least one micro-index (see Parakhin, paragraph [0024], where system 200 receives training data in the form of query data 202 and document data 204; the query data 202 is used to create a query-level model 206, and the document data 204 is used to create a document-level model 208; see also paragraph [0022], where the relationship component 102 computes the relationships 104 between the query clusters and the document clusters), the classification phase being configured to be executed by a classification device of the computer-implemented system, the classification phase comprising the steps of:
acquiring the plurality of data (see Parakhin, paragraph [0024], where system 200 receives training data in the form of … document data 204);
generating the indices according to at least a predetermined organization, the predetermined organization being configured to generate at least one index for at least some of the data of the plurality of data (see Parakhin, paragraph [0025], where from the document-level model 208 are created document vectors 212 both in cluster-wise probabilistic spaces);
computing a first set of embeddings using a first computer-implemented mathematical pre-trained model, the first set of embeddings comprising embeddings representative of at least some of the data of the plurality of data (see Parakhin, paragraph [0025], where from the document-level model 208 are created document vectors 212 both in cluster-wise probabilistic spaces);
acquiring a dataset of queries (see Parakhin, paragraph [0024], where system 200 receives training data in the form of query data 202); and
computing a second set of embeddings using a second computer-implemented mathematical pre-trained model, the second set of embeddings comprising embeddings representative of at least one of the query of the dataset of queries (see Parakhin, paragraph [0025], where from the query-level model 206 are created query vectors 210 … both in cluster-wise probabilistic spaces).
Parakhin does not explicitly disclose:
generating a micro-index for each index of the distinct indices, wherein each micro-index is associated with an index of the distinct indices, and wherein each micro-index comprises at least one query from the dataset of queries;
the generating step comprising calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings; and
associating the queries of the dataset of queries with at least one index of the distinct indices based on the first distance metric to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index.
Parakhin in view of Cheng discloses generating a micro-index for each index of the distinct indices, wherein each micro-index is associated with an index of the distinct indices, and wherein each micro-index comprises at least one query from the dataset of queries (see Cheng, paragraph [0045], where the query enhancement module 310 uses clusters defined according to the clustering module 310 to modify a query and/or process search results of a query in order to better identify documents, such as product records, relevant to the received query; as will be described in greater detail below, the query enhancement module 310 may be operable to identify categories that are likely of interest to the author of the received query based on one or more categories associated with a cluster of queries related to the received query).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the inveniton to substitute the document/query co-clustering technique in Parakhin with the query cluster/cateogry association technique in Cheng as it amounts to simple substitution of one known element for another to obtain predictable results (see MPEP 2143(I)(B)).
Parakhin in view of Cheng does not disclose:
the generating step comprising calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings; and
associating the queries of the dataset of queries with at least one index of the distinct indices based on the first distance metric to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index.
Parakhin in view of Cheng and Manning discloses:
calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way);
associating the queries of the dataset of queries with at least one index of the distinct indices (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters; see also Claim 1, where the method includes computing a probability that the query belongs in a given query collection) based on the first distance metric (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way) to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin and Cheng with the distance based clustering techniques of Manning because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 3, Parakhin in view of Cheng and Manning discloses the method according to Claim 1, wherein:
Parakhin does not disclose each index of the distinct indices comprises a ranking of a set of data from the plurality of data, said ranking being related to the specific query, the set of data being related to the index, the ranking being generated by a computer-implemented mathematical ranking pre-trained model, based on features related to both the set of data and the specific query. Manning discloses each index of the distinct indices comprises a ranking of a set of data from the plurality of data, said ranking being related to the specific query, the set of data being related to the index, the ranking being generated by a computer-implemented mathematical ranking pre-trained model, based on features related to both the set of data and the specific query (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … cluster hypothesis offers an alternative: find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Manning because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 4, Parakhin in view of Cheng and Manning discloses the method according to Claim 1, wherein the dataset of queries has a predetermined statistical distribution being representative of the distinct indices (see Parakhin, paragraph [0019], where query-cluster-to-document-cluster transition parameters which describe the expected relevancy (relationship) of a document having an associated document cluster distribution to the query having an associated query cluster distribution).
Regarding Claim 7, Parakhin in view of Cheng and Manning discloses the method according to Claim 1, wherein the plurality of data comprises texts, and/or images, and/or videos, and/or multimedia files, and/or datasheets (see Parakhin, paragraph [0022], where the term ‘document’ includes one or more media types such as text, audio, video, image, and so on).
Regarding Claim 12, Parakhin discloses a non-transitory computer-readable medium (see Parakhin, paragraph [0056], where one or more programs and data can be stored in the memory subsystem 706, a machine readable and removable memory subsystem 718)) comprising executable instructions which, when executed by the at least one processor, cause the at least one processor to perform a method for retrieving at least one specific set of data among a plurality of data using a specific query, the method comprising:
a classification phase configured to classify the plurality of data into distinct indices, each index of the distinct indices being associated with at least one micro-index (see Parakhin, paragraph [0024], where system 200 receives training data in the form of query data 202 and document data 204; the query data 202 is used to create a query-level model 206, and the document data 204 is used to create a document-level model 208; see also paragraph [0022], where the relationship component 102 computes the relationships 104 between the query clusters and the document clusters), the classification phase being configured to be executed by a classification device of the computer-implemented system, the classification phase comprising the steps of:
acquiring the plurality of data (see Parakhin, paragraph [0024], where system 200 receives training data in the form of … document data 204);
generating the indices according to at least a predetermined organization, the predetermined organization being configured to generate at least one index for at least some of the data of the plurality of data (see Parakhin, paragraph [0025], where from the document-level model 208 are created document vectors 212 both in cluster-wise probabilistic spaces);
computing a first set of embeddings using a first computer-implemented mathematical pre-trained model, the first set of embeddings comprising embeddings representative of at least some of the data of the plurality of data (see Parakhin, paragraph [0025], where from the document-level model 208 are created document vectors 212 both in cluster-wise probabilistic spaces);
acquiring a dataset of queries (see Parakhin, paragraph [0024], where system 200 receives training data in the form of query data 202); and
computing a second set of embeddings using a second computer-implemented mathematical pre-trained model, the second set of embeddings comprising embeddings representative of at least one of the query of the dataset of queries (see Parakhin, paragraph [0025], where from the query-level model 206 are created query vectors 210 … both in cluster-wise probabilistic spaces).
Parakhin does not explicitly disclose:
generating a micro-index for each index of the distinct indices, wherein each micro-index is associated with an index of the distinct indices, and wherein each micro-index comprises at least one query from the dataset of queries;
the generating step comprising calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings; and
associating the queries of the dataset of queries with at least one index of the distinct indices based on the first distance metric to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index.
Parakhin in view of Cheng discloses generating a micro-index for each index of the distinct indices, wherein each micro-index is associated with an index of the distinct indices, and wherein each micro-index comprises at least one query from the dataset of queries (see Cheng, paragraph [0045], where the query enhancement module 310 uses clusters defined according to the clustering module 310 to modify a query and/or process search results of a query in order to better identify documents, such as product records, relevant to the received query; as will be described in greater detail below, the query enhancement module 310 may be operable to identify categories that are likely of interest to the author of the received query based on one or more categories associated with a cluster of queries related to the received query).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the inveniton to substitute the document/query co-clustering technique in Parakhin with the query cluster/cateogry association technique in Cheng as it amounts to simple substitution of one known element for another to obtain predictable results (see MPEP 2143(I)(B)).
Parakhin in view of Cheng does not disclose:
the generating step comprising calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings; and
associating the queries of the dataset of queries with at least one index of the distinct indices based on the first distance metric to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index.
Parakhin in view of Cheng and Manning discloses:
calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way);
associating the queries of the dataset of queries with at least one index of the distinct indices (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters; see also Claim 1, where the method includes computing a probability that the query belongs in a given query collection) based on the first distance metric (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way) to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin and Cheng with the distance based clustering techniques of Manning because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 13, Parakhin discloses a computer-implemented system for retrieving at least one specific set of data among a plurality of data using at least one specific query, the computer-implemented system comprising:
a classification device (see Parakhin, paragraph [0052], where computing system 700 for implementing various aspects includes the computer 702 having processing unit(s) 704), the classification device being configured to classify the plurality of data into distinct indices, each index of the distinct indices being associated with at least one micro-index (see Parakhin, paragraph [0024], where system 200 receives training data in the form of query data 202 and document data 204; the query data 202 is used to create a query-level model 206, and the document data 204 is used to create a document-level model 208; see also paragraph [0022], where the relationship component 102 computes the relationships 104 between the query clusters and the document clusters), the classification device comprising:
a first communication module, the first communication module being configured to acquire the plurality of data (see Parakhin, paragraph [0024], where system 200 receives training data in the form of … document data 204);
a second communication module, the second communication module being configured to acquire a dataset of queries (see Parakhin, paragraph [0024], where system 200 receives training data in the form of query data 202);
at least one processor (see Parakhin, paragraph [0052], where computing system 700 for implementing various aspects includes the computer 702 having processing unit(s) 704), the at least one processor of the classification device being configured to: being configured to generate the indices according to at least a predetermined organization (see Parakhin, paragraph [0025], where from the document-level model 208 are created document vectors 212 both in cluster-wise probabilistic spaces);
compute a first set of embeddings using a first computer-implemented mathematical pre-trained model, the first set of embeddings comprising embeddings representative of at least some of the data of the plurality of data (see Parakhin, paragraph [0025], where from the document-level model 208 are created document vectors 212 both in cluster-wise probabilistic spaces); and
computing a second set of embeddings using a second computer-implemented mathematical pre-trained model, the second set of embeddings comprising embeddings representative of at least one of the query of the dataset of queries (see Parakhin, paragraph [0025], where from the query-level model 206 are created query vectors 210 … both in cluster-wise probabilistic spaces).
Parakhin does not explicitly disclose:
generating a micro-index for each index of the distinct indices, wherein each micro-index is associated with an index of the distinct indices, and wherein each micro-index comprises at least one query from the dataset of queries;
the generating step comprising calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings; and
associating the queries of the dataset of queries with at least one index of the distinct indices based on the first distance metric to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index.
Parakhin in view of Cheng discloses generating a micro-index for each index of the distinct indices, wherein each micro-index is associated with an index of the distinct indices, and wherein each micro-index comprises at least one query from the dataset of queries (see Cheng, paragraph [0045], where the query enhancement module 310 uses clusters defined according to the clustering module 310 to modify a query and/or process search results of a query in order to better identify documents, such as product records, relevant to the received query; as will be described in greater detail below, the query enhancement module 310 may be operable to identify categories that are likely of interest to the author of the received query based on one or more categories associated with a cluster of queries related to the received query).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the inveniton to substitute the document/query co-clustering technique in Parakhin with the query cluster/cateogry association technique in Cheng as it amounts to simple substitution of one known element for another to obtain predictable results (see MPEP 2143(I)(B)).
Parakhin in view of Cheng does not disclose:
the generating step comprising calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings; and
associating the queries of the dataset of queries with at least one index of the distinct indices based on the first distance metric to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index.
Parakhin in view of Cheng and Manning discloses:
calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way);
associating the queries of the dataset of queries with at least one index of the distinct indices (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters; see also Claim 1, where the method includes computing a probability that the query belongs in a given query collection) based on the first distance metric (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way) to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin and Cheng with the distance based clustering techniques of Manning because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 15, Parakhin in view of Cheng and Manning discloses the computer-implemented system of Claim 13, wherein the at least one procsesor of the classification device is further configured to:
a first central processing unit (see Parakhin, paragraph [0052], where computing system 700 for implementing various aspects includes the computer 702 having processing unit(s) 704), the first central processing unit being configured to generate the indices according to at least a predetermined organization (see Parakhin, paragraph [0025], where from the document-level model 208 are created document vectors 212 both in cluster-wise probabilistic spaces);
generate a micro-index for each index of the distinct indices (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters); and
wherein the classification device further comprises a first graphic processing unit (see Parakhin, paragraph [0062], where one or more graphics interface(s) 736 (also commonly referred to as a graphics processing unit (GPU)), the first graphic processing unit being configured to:
compute a first set of embeddings using a first computer-implemented mathematical pre-trained model, the first set of embeddings comprising embeddings representative of at least some of the data of the plurality of data (see Parakhin, paragraph [0025], where from the document-level model 208 are created document vectors 212 both in cluster-wise probabilistic spaces);
compute a second set of embeddings using a second computer-implemented mathematical pre-trained model, the second set of embeddings comprising embeddings representative of at least one of the query of the dataset of queries (see Parakhin, paragraph [0025], where from the query-level model 206 are created query vectors 210 … both in cluster-wise probabilistic spaces)
Parakhin does not explicitly disclose:
calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings; and
associating the queries of the dataset of queries with at least one index of the distinct indices based on the first distance metric to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index.
Parakhin in view of Manning discloses:
calculating a first distance metric between at least one embedding of the first set of embeddings and at least one embedding of the second set of embeddings (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way);
associating the queries of the dataset of queries with at least one index of the distinct indices (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters; see also Claim 1, where the method includes computing a probability that the query belongs in a given query collection) based on the first distance metric (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way) to create at least one group of queries for each index, the group of queries defining at least one micro-index of the index (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Manning because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Claims 2, 8-10, 14, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Parakhin, Cheng, and Manning as applied to Claims 1, 3, 4, 7, 12, and 13 above, and further in view of Agarwal (PG Pub. No. 2021/0149963 A1).
Regarding Claim 2, Parakhin in view of Cheng and Manning discloses the method according to Claim 1, further comprising a retrieving phase configured to retrieve the specific set of data based on the specific query, the retrieving phase being configured to be executed by a retrieving device, the retrieving phase comprising steps of:
receiving the specific query (see Parakhin, Claims 1, where the system computes a total weight of each query in a query collection and computes the probability that the query belongs in a given query collection; see also Claim 2, where the relationships are automatically determined during an offline training phase and the query processing occurs online);
computing a specific embedding of the specific query using the second computer-implemented mathematical pre-trained model (see Parakhin, paragraph [0025], where from the query-level model 206 are created query vectors 210 … both in cluster-wise probabilistic spaces).
Parakhin does not explicitly disclose:
calculating a second distance metric between the specific embedding and at least one embedding of the second set of embeddings;
retrieving a specific index related to a specific micro-index based on the second distance metric; and
retrieving the specific set of data which are classified into the specific index.
Parakhin in view of Agarwal discloses:
calculating a second distance metric between the specific embedding and at least one embedding of the second set of embeddings (see Agarwal, paragraph [0039], where the predicate function may be based on measuring a distance value, e.g., assessing whether the candidate query and the representative query have at most a predefined threshold vector distance); and
retrieving a specific index related to a specific micro-index (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters) based on the second distance metric (see Agarwal, paragraph [0039], where the predicate function may be based on measuring a distance value, e.g., assessing whether the candidate query and the representative query have at most a predefined threshold vector distance).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Parakhin in view of Agarwal does not disclose retrieving the specific set of data which are classified into the specific index. Manning discloses retrieving the specific set of data which are classified into the specific index (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal and Manning because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 5, Parakhin in view of Cheng and Manning discloses the method according to Claim 1, wherein:
Parakhin does not disclose the step of associating each query of the dataset of queries with at least one index of the distinct indices is configured to use at least one classifier, the classifier being configured to use at least the first distance metric to predict at least one index of the distinct indices corresponding to at least one query of the dataset of queries. Parakhin in view of Agarwal discloses the step of associating each query of the dataset of queries with at least one index of the distinct indices is configured to use at least one classifier (see Agarwal, paragraph [0075], where non-limiting examples of training procedures for adjusting trainable parameters include … unsupervised learning methods (e.g., classification based on classes derived from unsupervised clustering methods), the classifier being configured to use at least the first distance metric to predict at least one index of the distinct indices corresponding to at least one query of the dataset of queries (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters; see also Claim 1, where the method includes computing a probability that the query belongs in a given query collection).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 8, Parakhin in view of Cheng and Manning discloses the method according to Claim 1, wherein:
Parakhin does not disclose the acquiring step of the dataset of queries comprises downloading the dataset of queries from at least one database and/or from at least one historic of queries of at least one user. Agarwal discloses the acquiring step of the dataset of queries comprises downloading the dataset of queries from at least one database and/or from at least one historic of queries of at least one user (see Agarwal, paragraph [0014], where the approach described herein may be used to provide domain-agnostic query exploration for any suitable collection of web content based on historical queries related to the web content).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 9, Parakhin in view of Cheng, Manning and Agarwal discloses the method according to Claim 2, wherein:
Parakhin does not explicitly disclose wherein the specific query is received from at least one input module, the input module comprising at least one of a keyboard, a microphone and/or a camera. Agarwal explicitly discloses wherein the specific query is received from at least one input module, the input module comprising at least one of a keyboard, a microphone and/or a camera (see Agarwal, paragraph [0016], where user query 110 may be provided via a search graphical user interface (GUI) 112 visually presented by the user device 104; user device 104 may present the search GUI based on suitable computer-executable code; see also paragraph [0081], where examples of user input devices include a keyboard).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 10, Parakhin in view of Cheng, Manning and Agarwal discloses the method according to Claim 2, wherein:
Parakhin does not explicitly disclose the specific query is received from a user, the user being associated with at least one of a smartphone, a computer, a self-driving vehicle, a drone, a robot, a vacuum and/or a person. Agarwal explicitly discloses the specific query is received from a user, the user being associated with at least one of a smartphone, a computer, a self-driving vehicle, a drone, a robot, a vacuum and/or a person (see Agarwal, paragraph [0016], where user query 110 may be provided via a search graphical user interface (GUI) 112 visually presented by the user device 104; user device 104 may present the search GUI based on suitable computer-executable code; see also paragraph [0081], where examples of user input devices include a keyboard).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 14, Parakhin in view of Cheng, Manning and Agarwal discloses the computer-implemented system of Claim 13, further comprising a retrieving device (see Parakhin, paragraph [0052], where computing system 700 for implementing various aspects includes the computer 702 having processing unit(s) 704), the retrieving device being configured to retrieve at least the specific set of data using at least the specific query, the retrieving device comprising:
a third communication module, the third communication module being configured to receive the specific query (see Parakhin, Claims 1, where the system computes a total weight of each query in a query collection and computes the probability that the query belongs in a given query collection; see also Claim 2, where the relationships are automatically determined during an offline training phase and the query processing occurs online); and
at least one processor (see Parakhin, paragraph [0052], where computing system 700 for implementing various aspects includes the computer 702 having processing unit(s) 704), the at least one processor of the retrieving device being configured to: compute a specific embedding of the specific query using the second computer-implemented mathematical pre-trained model (see Parakhin, paragraph [0025], where from the query-level model 206 are created query vectors 210 … both in cluster-wise probabilistic spaces); and
Parakhin does not explicitly disclose:
calculating a second distance metric between the specific embedding and at least one embedding of the second set of embeddings;
retrieving a specific index related to a specific micro-index based on the second distance metric; and
retrieving the specific set of data which are classified into the specific index.
Parakhin in view of Agarwal discloses:
calculating a second distance metric between the specific embedding and at least one embedding of the second set of embeddings (see Agarwal, paragraph [0039], where the predicate function may be based on measuring a distance value, e.g., assessing whether the candidate query and the representative query have at most a predefined threshold vector distance); and
retrieving a specific index related to a specific micro-index (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters) based on the second distance metric (see Agarwal, paragraph [0039], where the predicate function may be based on measuring a distance value, e.g., assessing whether the candidate query and the representative query have at most a predefined threshold vector distance).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Parakhin in view of Agarwal does not disclose retrieving the specific set of data which are classified into the specific index. Manning discloses retrieving the specific set of data which are classified into the specific index (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal and Manning because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Regarding Claim 16, Parakhin in view of Cheng, Manning and Agarwal discloses the computer-implemented system according to Claim 14, wherein the at least one processor of the retrieving device is further configured to:
wherein the retrieving device comprises a second graphic processing unit (see Parakhin, paragraph [0062], where one or more graphics interface(s) 736 (also commonly referred to as a graphics processing unit (GPU)), the second graphic processing unit being configured to compute a specific embedding of the specific query (see Parakhin, paragraph [0025], where from the query-level model 206 are created query vectors 210 … both in cluster-wise probabilistic spaces).
Parakhin does not explicitly disclose:
the second processing unit being configured to calculate a second distance metric between the specific embedding and at least one embedding of the second set of embeddings;
retrieving a specific index related to a specific micro-index based on the second distance metric; and
retrieving the specific set of data which are classified into the specific index.
Parakhin in view of Agarwal discloses:
the second processing unit being configured to calculate a second distance metric between the specific embedding and at least one embedding of the second set of embeddings (see Agarwal, paragraph [0039], where the predicate function may be based on measuring a distance value, e.g., assessing whether the candidate query and the representative query have at most a predefined threshold vector distance); and
retrieving a specific index related to a specific micro-index (see Parakhin, paragraph [0022], where the relationship component 102 computes the relationships 104 between query clusters and document clusters) based on the second distance metric (see Agarwal, paragraph [0039], where the predicate function may be based on measuring a distance value, e.g., assessing whether the candidate query and the representative query have at most a predefined threshold vector distance).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Parakhin in view of Agarwal does not disclose retrieving the specific set of data which are classified into the specific index. Manning discloses retrieving the specific set of data which are classified into the specific index (see Manning, paragraph 11, where search in the vector space model amounts to finding the nearest neighbors to the query … find the clusters that are closest to the query and only consider documents from these clusters; within this much smaller set, we can compute similarities exhaustively and rank documents in the usual way).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the distance based clustering techniques of Agarwal and Manning because they amount to combining prior art elements according to known methods to yield predictable results (see MPEP 2143(I)(A)).
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Parakhin, Cheng, and Manning as applied to Claims 1, 3, 4, 7, 12, 13, and 15 above, and further in view of Dutta (PG Pub. No. 2025/0130994 A1).
Regarding Claim 6, Parakhin in view of Cheng and Manning discloses the method according to Claim 1, comprising a step of:
Parakhin does not disclose adding additional queries to the dataset of queries, the additional queries having a first distance metric higher than a predetermined threshold and/or leading to a wrong specific set of data, and wherein the additional queries comprise additional data, the additional data being at least one of data about a duration of an interaction between a user and at least one data and/or data about a number of interactions between a user and at least one data. Parakhin in view of Dutta discloses adding additional queries to the dataset of queries, the additional queries having a first distance metric higher than a predetermined threshold and/or leading to a wrong specific set of data, and wherein the additional queries comprise additional data, the additional data being at least one of data about a duration of an interaction between a user and at least one data and/or data about a number of interactions between a user and at least one data (see Dutta, paragraph [0036], where parameters of a trained function can be adapted by means of training; in particular, … active learning can be used; see also paragraph [0115], where the training dataset 458 includes historical query data 452; the historical query data 452 can be obtained from historical transaction data stored by one or more databases, such as database 14; the historical query data 452 includes historical (e.g., past) queries and associated features; see also paragraph [0176], where processing of the received training dataset 752 includes outlier detection configured to remove data likely to skew training).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the cluster relationship techniques of Parakhin with the query clustering machine learning model refinement techniques of Dutta for the benefit of remediating outlier data (see Dutta, paragraph [0176]).
Response to Arguments
Applicant’s Arguments, filed January 26, 2026, have been fully considered, but they are moot in light of the new grounds of rejection.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FARHAD AGHARAHIMI whose telephone number is (571)272-9864. The examiner can normally be reached M-F 9am - 5pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Apu Mofiz can be reached at 571-272-4080. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/FARHAD AGHARAHIMI/Examiner, Art Unit 2161
/APU M MOFIZ/Supervisory Patent Examiner, Art Unit 2161