DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e)
or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. The present
application is recognized as a continuation of parent Application No. 18/518,281 filed
on 11/22/2023 and claiming priority to Provisional Application No. 65/427,714 filed 11/23/2022.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 1 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding independent claim 1,
Claim 1 recites the following limitation(s):
a primary data storage unit to store the index for the plurality of vectors;
and a quick-retrieval data storage unit to cache a segment of a data stored in the primary data storage unit, wherein the segment of the data includes a periodic queried data from the primary data.
While claim 1 provides antecedent basis for “a primary data storage unit” but does not provide antecedent basis for “primary data”. It is unclear what is being referred to as “primary data”, therefore claim 1 is indefinite.
The examiner believes the term “primary data” refers to the “index for the plurality of vectors” as in the first limitation listed above. The claim is being interpreted in that manner for the purposes of this Non-Final Rejection. The examiner suggests amending claim 1 to refer to the index stored in the primary data storage unit as described in the limitation above.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding independent claim 1,
Claim 1 recites the following limitation(s):
wherein the indexing processor is to: cluster the plurality of vectors into a first set of clusters, wherein each cluster among the first set of clusters includes a pre-determined number of vectors from among the plurality of vectors; Which comprises a step of clustering vectors recited at a high degree of generality. One of ordinary skill in the art may determine, either mentally or aided by pen & paper, a cluster of vectors based on metrics such as similarity, distance, etc. that are known in the art to determine a set of vectors.
determine a centroid of each cluster among the first set of clusters, wherein the centroid indicates a center of the cluster; Which recites a step of determining a centroid of a cluster recited at a high degree of generality. The term “determine” is considered to be an observation or evaluation which are considered concepts performed in the human mind. For example, one of ordinary skill in the art would be able to calculate a centroid using known methods such as by calculating an average of all values of a cluster either mentally or aided by pen & paper.
cluster the centroids of the first set of clusters to form a second set of clusters, Which comprises a step of clustering vectors recited at a high degree of generality. One of ordinary skill in the art may determine, either mentally or aided by pen & paper, a cluster of vectors based on metrics such as similarity, distance, etc. that are known in the art to determine a set of vectors.
determine a centroid of each cluster among the second set of clusters; Which recites a step of determining a centroid of a cluster recited at a high degree of generality. The term “determine” is considered to be an observation or evaluation which are considered concepts performed in the human mind. For example, one of ordinary skill in the art would be able to calculate a centroid using known methods such as by calculating an average of all values of a cluster either mentally or aided by pen & paper.
and cluster the centroids of the second set of clusters to obtain a single cluster of centroids, the single cluster of centroids includes a pre- determined number of centroids; Which recites a step of iteratively clustering, specifically the limitation describes a second step of clustering. As discussed above, a step of clustering may be performed, either mentally or aided by pen & paper, by one of ordinary skill in the art by applying cluster analysis tools such as similarity, distance, etc. that are known in the art to determine a set of vectors.
Claim 1 recites the following additional elements:
a vector data storage unit to be deployed in a multi-tenant environment having a plurality of tenants; which encompasses a step of selecting a type or source of data to be manipulated (e.g. a vector data storage unit is a source of data manipulated by the process of storing) which represents insignificant extra-solution activity as described in MPEP 2106.05(g).
the vector data storage unit comprising: an intermediate data storage unit to store metadata information for a plurality of vectors, metadata information for the plurality of tenants, and vector operations associated with the plurality of vectors, the intermediate data storage unit including an indexing processor for creating an index for the plurality of vectors, which encompasses a step of selecting a type or source of data to be manipulated (e.g. an intermediate data storage unit source of data manipulated by the process of storing. Metadata information and vector operations represent types of data being stored.) which represents insignificant extra-solution activity as described in MPEP 2106.05(g).
the second set of clusters including one or more centroids based on the plurality of vectors included in the first set of clusters; which encompasses a step of selecting a type or source of data to be manipulated (e.g. the second set of clusters including centroids describes a type of data to be manipulated, such as via the clustering.) which represents insignificant extra-solution activity as described in MPEP 2106.05(g).
wherein each cluster among the second set of clusters includes a pre-determined number of centroids; which encompasses a step of selecting a type or source of data to be manipulated (e.g. a cluster is a type of data. The limitation relates to an amount of a certain type of data.) which represents insignificant extra-solution activity as described in MPEP 2106.05(g).
a primary data storage unit to store the index for the plurality of vectors; which encompasses a step of selecting a type or source of data to be manipulated (e.g. a primary data storage unit is a source of data.) which represents insignificant extra-solution activity as described in MPEP 2106.05(g).
and a quick-retrieval data storage unit to cache a segment of a data stored in the primary data storage unit, wherein the segment of the data includes a periodic queried data from the primary data. which encompasses a step of selecting a type or source of data to be manipulated (e.g. a quick-retrieval data storage unit is a source of data.) which represents insignificant extra-solution activity as described in MPEP 2106.05(g).
The judicial exception is not integrated into a practical application because the computer system elements are recited at a high level of generality. Which amounts to no more than applying the steps to within a generic computer environment. Accordingly, these elements do not integrate the abstract idea into a practical application because the limitations do not impose any meaningful limits on practicing the abstract idea, see MPEP 2106.06(f). As such, the claim is directed to the abstract idea of a mental process.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements in the claim amount to no more than mere instructions applied to a generic computer environment. Mere instructions to apply a judicial exception using a generic computer environment cannot integrate a judicial exception into a practical application or provide an inventive concept.
The additional elements, taken either alone or in combination do not result in the claim, as a whole, amount to significantly more than the judicial exception.
The following limitations represent elements that have been recognized as well-understood, routine, conventional activity within the field of computer functions:
a vector data storage unit to be deployed in a multi-tenant environment having a plurality of tenants; Ma et al. (US PGPUB No. 2022/0350828; Pub. Date: Nov. 3, 2022) discloses the limitation at issue: See Paragraph [0024], (Disclosing a system for converting a structured text document stored in a database into one or more vectors and creating a similarity search index. FIG. 1 illustrates system 100 comprising a structured text document database 101 storing a plurality of structured text documents having associated authors and reviewers.) See Paragraph [0051], (The system may build a search index 103 using vectors representing structured text documents stored in the structured text document database 101, i.e. a vector data storage unit to be deployed in a multi-tenant environment having a plurality of tenants (e.g. search index 103 stores vectors associated with individual documents of a database, each document vector representing a tenant) ;)
the vector data storage unit comprising: an intermediate data storage unit to store metadata information for a plurality of vectors, metadata information for the plurality of tenants, and vector operations associated with the plurality of vectors, Escalona et al. (US PGPUB No. 2022/0164397; Pub. Date: May 26, 2022) discloses the limitation at issue: See Paragraph [0032], (Disclosing a system for relevance-based analysis and filtering of documents and media for one or more enterprises. The system comprises memory 112 comprising database 113 for storing received documents or text data, NLP data, one or more feature vectors, metadata, training data, ML model parameters, etc. Note [0042] metadata may be associated with feature vectors, Note [0049] metadata may include paragraph identifiers corresponding to paragraphs in an input document, i.e. the vector data storage unit comprising: an intermediate data storage unit to store metadata information for a plurality of vectors (e.g. feature vector metadata described in [0042]), metadata information for the plurality of tenants (e.g. metadata associated with documents),
the intermediate data storage unit including an indexing processor for creating an index for the plurality of vectors, Ma et al. (US PGPUB No. 2022/0350828; Pub. Date: Nov. 3, 2022) discloses the limitation at issue: See Paragraph [0061], (The system may generate an inverted file index (IVF) comprising a plurality of partitions. Partitions of the IVF search index start with centroids which define a partition consisting of all vectors closer to the centroid than any other centroid, i.e. the intermediate data storage unit including an indexing processor for creating an index for the plurality of vectors (e.g. system 100 may generate search index 103), wherein the indexing processor is to: cluster the plurality of vectors into a first set of clusters, wherein each cluster among the first set of clusters includes a pre-determined number of vectors from among the plurality of vectors;)
the second set of clusters including one or more centroids based on the plurality of vectors included in the first set of clusters; Ma et al. (US PGPUB No. 2022/0350828; Pub. Date: Nov. 3, 2022) discloses the limitation at issue: See Paragraph [0061], (Searching the IVF search index using a query vector comprises identifying a centroid closest to the search vector and applying a KNN algorithm to compute the distances between each vector in that centroid's partition and the search vector. A number "K" vectors in the partition associated with the K smallest distances are reported y the KNN algorithm as the K-nearest neighbors to the search vector, i.e. cluster the centroids of the first set of clusters to form a second set of clusters (e.g. the search process includes identifying partitions closest to a query vector wherein partitions are associated with centroids), the second set of clusters including one or more centroids based on the plurality of vectors included in the first set of clusters (e.g. the plurality of partitions are associated with centroids);)
wherein each cluster among the second set of clusters includes a pre-determined number of centroids; Ma et al. (US PGPUB No. 2022/0350828; Pub. Date: Nov. 3, 2022) discloses the limitation at issue: See Paragraph [0061], (The IVF search index comprises a plurality of partitions constructed as Dirichlet tesselation wherein each partition starts with a centroid placed into the search index as a dividing point. Each centroid defines a partition consisting of all vectors closer to the centroid that any other centroid, i.e. wherein each cluster among the second set of clusters includes a pre-determined number of centroids (e.g. the centroids associated with each of the plurality of partitions) ;)
a primary data storage unit to store the index for the plurality of vectors; Ma et al. (US PGPUB No. 2022/0350828; Pub. Date: Nov. 3, 2022) discloses the limitation at issue: See Paragraph [0067], (System 100 may construct similarity search index 406 for storing the plurality of generated vectors, i.e. a primary data storage unit to store the index for the plurality of vectors;)
and a quick-retrieval data storage unit to cache a segment of a data stored in the primary data storage unit, wherein the segment of the data includes a periodic queried data from the primary data. Wu et al. (US Patent No. 11,055,228; Date of Patent: Jul. 6, 2021) discloses the limitation at issue: See Col. 1, lines 44-51, (Disclosing a system for interacting with a multi-level memory comprising near memory 113 used to store more frequently accessed items of program code and/or data kept in system memory 112. Near memory 113 is utilized as a cache for a larger, slower far memory 114, i.e. a quick-retrieval data storage unit to cache a segment of a data stored in the primary data storage unit, wherein the segment of the data includes a periodic queried data from the primary data (e.g. near memory 113 is utilized as a cache for far memory 114 and stores frequently-accessed data items).)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ma et al. (US PGPUB No. 2022/0350828; Pub. Date: Nov. 3, 2022) in view of Escalona et al. (US PGPUB No. 2022/0164397; Pub. Date: May 26, 2022) and Wu et al. (US Patent No. 11,055,228; Date of Patent: Jul. 6, 2021).
Regarding independent claim 1,
Ma discloses a system comprising: a vector data storage unit to be deployed in a multi-tenant environment having a plurality of tenants; See Paragraph [0024], (Disclosing a system for converting a structured text document stored in a database into one or more vectors and creating a similarity search index. FIG. 1 illustrates system 100 comprising a structured text document database 101 storing a plurality of structured text documents having associated authors and reviewers.) See Paragraph [0051], (The system may build a search index 103 using vectors representing structured text documents stored in the structured text document database 101, i.e. a vector data storage unit to be deployed in a multi-tenant environment having a plurality of tenants (e.g. search index 103 stores vectors associated with individual documents of a database, each document vector representing a tenant) ;)
the intermediate data storage unit including an indexing processor for creating an index for the plurality of vectors, wherein the indexing processor is to: cluster the plurality of vectors into a first set of clusters, wherein each cluster among the first set of clusters includes a pre-determined number of vectors from among the plurality of vectors; See Paragraph [0061], (The system may generate an inverted file index (IVF) comprising a plurality of partitions. Partitions of the IVF search index start with centroids which define a partition consisting of all vectors closer to the centroid than any other centroid, i.e. the intermediate data storage unit including an indexing processor for creating an index for the plurality of vectors (e.g. system 100 may generate search index 103), wherein the indexing processor is to: cluster the plurality of vectors into a first set of clusters, wherein each cluster among the first set of clusters includes a pre-determined number of vectors from among the plurality of vectors;)
determine a centroid of each cluster among the first set of clusters, wherein the centroid indicates a center of the cluster; See Paragraph [0061], (Partitions of the IVF search index start with centroids which define a partition consisting of all vectors closer to the centroid than any other centroid, i.e. determine a centroid of each cluster among the first set of clusters, wherein the centroid indicates a center of the cluster (e.g. constructing a partition includes starting with centroids);)
cluster the centroids of the first set of clusters to form a second set of clusters, the second set of clusters including one or more centroids based on the plurality of vectors included in the first set of clusters; See Paragraph [0061], (Searching the IVF search index using a query vector comprises identifying a centroid closest to the search vector and applying a KNN algorithm to compute the distances between each vector in that centroid's partition and the search vector. A number "K" vectors in the partition associated with the K smallest distances are reported y the KNN algorithm as the K-nearest neighbors to the search vector, i.e. cluster the centroids of the first set of clusters to form a second set of clusters (e.g. the search process includes identifying partitions closest to a query vector wherein partitions are associated with centroids), the second set of clusters including one or more centroids based on the plurality of vectors included in the first set of clusters (e.g. the plurality of partitions are associated with centroids);)
wherein each cluster among the second set of clusters includes a pre-determined number of centroids; See Paragraph [0061], (The IVF search index comprises a plurality of partitions constructed as Dirichlet tesselation wherein each partition starts with a centroid placed into the search index as a dividing point. Each centroid defines a partition consisting of all vectors closer to the centroid that any other centroid, i.e. wherein each cluster among the second set of clusters includes a pre-determined number of centroids (e.g. the centroids associated with each of the plurality of partitions) ;)
determine a centroid of each cluster among the second set of clusters; See FIG. 4 & Paragraphs [0061]-[0062] & [0067], (System 400 comprises a vector calculation component 402a comprising converting text document data into one or more vectors for storage in the similarity search index. The IVF search index comprises centroids embodied as fictional vectors that are searched as part of executing a KNN algorithm to determine closest vectors to a query vector based on computing distances between each vector in a centroid's partition, i.e. determine a centroid of each cluster among the second set of clusters;)
and cluster the centroids of the second set of clusters to obtain a single cluster of centroids, the single cluster of centroids includes a pre- determined number of centroids; See Paragraph [0061], (Searching the IVF search index using a query vector comprises identifying a centroid closest to the search vector and applying a KNN algorithm to compute the distances between each vector in that centroid's partition and the search vector. A number "K" vectors in the partition associated with the K smallest distances are reported y the KNN algorithm as the K-nearest neighbors to the search vector, i.e. cluster the centroids of the second set of clusters to obtain a single cluster of centroids, the single cluster of centroids includes a pre- determined number of centroids (e.g. the reported number "K" of vectors identified by the KNN algorithm includes centroids which are represented as vectors);)
a primary data storage unit to store the index for the plurality of vectors; See Paragraph [0067], (System 100 may construct similarity search index 406 for storing the plurality of generated vectors, i.e. a primary data storage unit to store the index for the plurality of vectors;)
Ma does not disclose the vector data storage unit comprising: an intermediate data storage unit to store metadata information for a plurality of vectors, metadata information for the plurality of tenants, and vector operations associated with the plurality of vectors,
Escalona discloses the vector data storage unit comprising: an intermediate data storage unit to store metadata information for a plurality of vectors, metadata information for the plurality of tenants, See Paragraph [0032], (Disclosing a system for relevance-based analysis and filtering of documents and media for one or more enterprises. The system comprises memory 112 comprising database 113 for storing received documents or text data, NLP data, one or more feature vectors, metadata, training data, ML model parameters, etc. Note [0042] metadata may be associated with feature vectors, Note [0049] metadata may include paragraph identifiers corresponding to paragraphs in an input document, i.e. the vector data storage unit comprising: an intermediate data storage unit to store metadata information for a plurality of vectors (e.g. feature vector metadata described in [0042]), metadata information for the plurality of tenants (e.g. metadata associated with documents).)
The examiner notes that while Ma does not disclose an intermediate data storage unit, Ma discloses vector operations associated with the plurality of vectors, See Paragraph [0046], (System 100 comprises vector calculations 102a-102b comprising algorithms that generate vectors, i.e. and operations associated with the plurality of vectors.)
The examiner notes that one of ordinary skill in the art would be aware that algorithms such as vector calculations 102a-102b of Ma may be stored in memory such as in database 113 of Escalona for execution.
Ma and Escalona are analogous art because they are in the same field of endeavor, data storage techniques. It would have been obvious to anyone having ordinary skill in the art before the effective filing date to modify the system of Ma to include the method of storing metadata relating to vector and document information as disclosed by Escalona. Paragraph [0062] of Escalona discloses that the system allows for presentation of an enhanced GUI that allows a user to quickly receive the most relevant documents from a large quantity of input documents that would otherwise require a significant number of personnel to review.
Ma-Escalona does not disclose a quick-retrieval data storage unit to cache a segment of a data stored in the primary data storage unit, wherein the segment of the data includes a periodic queried data from the primary data.
Wu not discloses a quick-retrieval data storage unit to cache a segment of a data stored in the primary data storage unit, wherein the segment of the data includes a periodic queried data from the primary data. See Col. 1, lines 44-51, (Disclosing a system for interacting with a multi-level memory comprising near memory 113 used to store more frequently accessed items of program code and/or data kept in system memory 112. Near memory 113 is utilized as a cache for a larger, slower far memory 114, i.e. a quick-retrieval data storage unit to cache a segment of a data stored in the primary data storage unit, wherein the segment of the data includes a periodic queried data from the primary data (e.g. near memory 113 is utilized as a cache for far memory 114 and stores frequently-accessed data items).)
Ma, Escalona and Wu are analogous art because they are in the same field of endeavor, data storage techniques. It would have been obvious to anyone having ordinary skill in the art before the effective filing date to modify the system of Ma-Escalona to include the method of storing frequently accessed records in faster memory as disclosed by Wu. Col. 1, lines 52-65 of Wu disclose that near memory 113 has lower access times than far memory 114 by having a faster clock speed, which represents an improvement in data retrieval technology by providing more efficient access to stored data.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Fernando M Mari whose telephone number is (571)272-2498. The examiner can normally be reached Monday-Friday 7am-4pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ann J. Lo can be reached at (571) 272-9767. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/FMMV/Examiner, Art Unit 2159 /ANN J LO/Supervisory Patent Examiner, Art Unit 2159