DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Allowable Subject Matter
Claims 1-20 are allowable over the prior art. However, the claims remain rejected under 35 USC §§101 and 112(b).
Reasons For Allowance
The cited references do not disclose summing one or more inter-cluster vector distances between one or more pairs of the centroid vectors, summing one or more intra-cluster vector distances between one or more pairs of descriptive vectors, determining a plurality of scores of the hierarchy of vector clusters, selecting a subset of vector clusters in the hierarchy of vector clusters, and storing identifiers of centroid vectors of vector clusters in the selected subset, wherein the identifiers are stored with a timestamp in a history of media items.
Claim Rejections – 35 U.S.C. § 101
35 U.S.C. § 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to non-statutory subject matter.
These claims are rejected under 35 USC §101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites at a very level center points of data collections, distances between data/center points and selecting/storing data and data representations. Thus, the claims encompass the performance of the limitations in the mind, or alternatively the solving of a math problem (i.e., a series of mathematical steps, such as “summing” and “determining … scores”, determining a center point [“centroid”]) that are not tied to a practical application.
Regarding the independent claims:
Step 1: Yes, claim 1 recites a storage medium for a series of steps executed (therefore a process embodied in a product/machine), claim 11 is directed to a system (therefore a product/machine), and claim 20 is directed to a method (therefore a process). Thus, each of these claims is directed to a statutory category.
Step 2A, Prong 1 (Judicial Exception Recited?): Yes. Claims 1, 11 and 20 recite limitations directed to an abstract idea: “determining centroid vectors of one or more vector clusters in a hierarchy of vector clusters, wherein each centroid vector of the centroid vectors corresponds to a respective vector cluster; summing one or more inter-cluster vector distances between one or more pairs of the centroid vectors; summing one or more intra-cluster vector distances between one or more pairs of descriptive vectors; determining a plurality of scores of the hierarchy of vector clusters; selecting a subset of vector clusters in the hierarchy of vector clusters;”. As drafted, each of these limitations recites a mentally performable process (or alternatively a mathematical step) as one can determine a center location of data collection points, determine/add distances between points, formulate scores and choose data collections/objects via a mental process or using paper and pencil (or in the alternative as a series of mathematical steps).
Step 2A, Prong 2 (Integrated into a Practical Application?): No. Claim 1 recites the following additional elements, "non-transitory computer readable storage medium” and execution by “one or more processors”. Claim 11 recites “a computing device”, “one or more processors”, "non-transitory computer readable storage medium”. Claim 20 recites “a computer”. Each of these are merely high-level recitations of generic computer components and represent mere instructions to apply on a computer as in MPEP 2106.05(f), which does not provide integration into a practical application.
Additionally, claims 1, 11 and 20 each recites “storing …” data. It is noted that: Use of a computer or other machinery in its ordinary capacity for economic or other tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., a fundamental economic practice or mathematical equation) does not integrate a judicial exception into a practical application or provide significantly more. See Affinity Labs v. DirecTV, 838 F.3d 1253, 1262, 120 USPQ2d 1201, 1207 (Fed. Cir. 2016) (cellular telephone); TLI Communications LLC v. AV Auto, LLC, 823 F.3d 607, 613, 118 USPQ2d 1744, 1748 (Fed. Cir. 2016) (computer server and telephone unit).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose meaningful limits on practicing the abstract idea. Viewing the additional limitations together and the claims as a whole, nothing provides integration into a practical application. Therefore, each claim is directed to an abstract idea.
Step 2B (Inventive Concept Provided?): No. As discussed with respect to Step 2A, the elements (i.e., steps of “storing”) in the claim amount to no more than mere instructions to apply the exception. Mere instructions to apply an exception using generic computer components (e.g., storage media, processors) cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
With respect to the “storing” limitation discussed above, and when re-evaluated, this element is well-understood, routine, and conventional as evidenced by the court cases in MPEP 2106.05(d)(II), "i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); … OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93;" and thus remains insignificant extra-solution activity that does not provide significantly more.
Therefore, each of the independent claims, taken as a whole, does not change this conclusion and the claims are ineligible.
Claims 2-10 depend upon claim 1, and do not correct the issues set forth above. These claims essentially further involve elements such as: determining scores (mental process / pencil and paper), selecting/choosing using scores (mental process / pencil and paper), data description (insignificant extra-solution activity of selection of a particular source/type of data to manipulate - MPEP 2106.05(g)), data/identifier description (insignificant extra-solution activity of selection of a particular source/type of data to manipulate - MPEP 2106.05(g)), transmitting information (receiving/transmitting over a network), clustering/grouping based upon a determined distance(mental process / pencil and paper, or alternatively a mathematical concept), and data description (insignificant extra-solution activity of selection of a particular source/type of data to manipulate - MPEP 2106.05(g)).
Therefore, these claims are likewise rejected.
Claims 12-19 depend upon claim 11, and do not correct the issues set forth above. These claims are substantially similar to claims 2-7 and 9-10, and are therefore likewise rejected.
Therefore, the claims (1-20) are reasonably rejected under 35 USC §101.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. § 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Regarding independent claims 1, 11 and 20:
The independent claims encompass the determining of only one centroid vector cluster (e.g., claim 1 line 4 “determining centroid vectors of one or more vector clusters …” - only requires that one centroid vector be determined). However, the next step/limitation (“summing”) requires at least one pair of centroid vectors (i.e., two centroid vectors) in order to have a distance value element (i.e., a distance between … pairs of centroid vectors). Therefore, in the case where only one centroid vector is “determined” at line 4 of the claim, the “summing” step cannot be performed, rendering the meaning/scope of the independent claims unclear.
Therefore, the scope of each claim is ambiguous.
Claims 2-10 and 12-19 depend upon claims 1 and 11, respectively, and do not correct the issues set forth above. Therefore, these claims are likewise rejected.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Relevance is provided in at least the Abstract of each cited document.
Non-Patent Literature
Ozkarahan, Esen, “Multimedia Document Retrieval”, Information Processing and Management, Vol. 31, No. 1, © 1994 Elsevier Science Ltd, pp. 113-131.
This study develops an integrated conceptual representation media documents that are viewed to comprise an object-oriented scheme for multi- database. the necessary abstractions for the conceptual It develops model and extensions to the RM/T relational model used as the search structure. It then develops a retrieval model in which the database search space is first narrowed down, based on user query, by an associative search. The associative search is followed by semantic and media-specific searches. (page 113, Abstract). A fast-access path structure of such a database can be constructed by first clustering m documents into n, number of clusters, and then representing each cluster by a single vector called the centroid. The centroid vector is formed by computing the average weights of index terms contained in the document vectors of the cluster. If we use a binary weighting scheme, then a term’s weight in a document is 1 if the term appears in that document, or it is 0 otherwise. (page 119, 3rd paragraph). This study presents an information system methodology that represents multimedia documents and enables querying and searching of the multimedia database in an integrated manner. The study is based on a conceptual representation of the integrated multimedia database. The documents are represented as objects in the conceptual model, which allows abstractions such as class and aggregation hierarchies, nonoptional and spatial relation- ships. A specific model is developed to represent these abstractions conceptually. After representing documents conceptually, this study develops a retrieval model. The retrieval model encompasses search phases referred to as associative, semantic, and media searches. The associative search phase utilizes a document vector-clustering scheme. The document-clustering scheme used here is based on intercoupling among documents. This is a semantic clustering technique, and a summary of its methodology is presented in the appendix. (page 127, 2nd and 3rd paragraphs of section “6. Conclusion”).
Saad, Fathi H., et al., “Comparison of Hierarchical Agglomerative Algorithms for Clustering Medical Documents”, International Journal of Software Engineering & Applications (IJSEA), Vol. 3, No. 3, May 2012, pp. 1-15.
what they are looking for effectively by organizing large amounts of information into a small number of meaningful clusters. The produced clusters contain groups of objects which are more similar to each other than to the members of any other group. Thus, the aim of high-quality document clustering algorithms is to determine a set of clusters in which the inter-cluster similarity is minimized and intra-cluster similarity is maximized. The most important feature in many clustering algorithms is treating the clustering problem as an optimization process, that is, maximizing or minimizing a particular clustering criterion function defined over the whole clustering solution. The only real difference between agglomerative algorithms is how they choose which clusters to merge. The main purpose of this paper is to compare different agglomerative algorithms based on the evaluation of the clusters quality produced by different hierarchical agglomerative clustering algorithms using different criterion functions for the problem of clustering medical documents. Our experimental results showed that the agglomerative algorithm that uses I1 as its criterion function for choosing which clusters to merge produced better clusters quality than the other criterion functions in term of entropy and purity as external measures. (page 1, Abstract). The centroid vector CS is the vector obtained by averaging the weights of different terms in the set of documents S [4, 7, 19] (page 3, last paragraph of section “3. Similarity Measures”). The second criterion function attempts to find the clustering solution that maximizes the similarity between each document and the cluster’s centroid. Using the cosine function to measure the similarity between a document and a centroid (page 4, last paragraph of section “4.1. Internal Criterion Functions”).
US Patent Application Publications
Freire 2009/0210406
A method is provided for organizing a plurality of documents that include forms. An initial set of clusters is defined for the plurality of documents. The initial set of clusters is reclustered based on similarity values calculated in multiple feature spaces. For example, a first feature space may be associated with a content of a document while a second feature space may be associated with a content of a form associated with the document. Each cluster has an associated centroid vector in each feature space that is used to represent the cluster. The similarity between the document and each cluster is calculated in both feature spaces. Each document is assigned to the cluster whose centroid is most similar. The cluster centroids may be recalculated and the process repeated until the cluster assignments become stable. (Abstract). Thus, in an exemplary embodiment, a first centroid vector is calculated in a first feature space (i.e., document-feature space) for a cluster C selected from the initialized clusters based on the relevant documents assigned to the cluster C. A second centroid vector is calculated in a second feature space (i.e., form-feature space) for the cluster C based on a form in the relevant documents assigned to the cluster C. A first similarity value (i.e., cosine distance) between a first vector and the first centroid vector is calculated where the first vector is calculated in the first feature space for a relevant document. A second similarity value (i.e., cosine distance) between a second vector and the second centroid vector is calculated where the second vector is calculated in the second feature space based the identified first form. A similarity value between the relevant document and the cluster C is calculated based on the calculated first similarity value and the calculated second similarity value, for example, using equation (2). The process is repeated for each of the initialized clusters to assign the relevant document to the closest cluster based on the calculated similarity value for each of the plurality of clusters. The process is repeated to assign each relevant document to the closest cluster based on the calculated similarity value for each of the plurality of clusters. (para 0066).
US Patents
Caid 5,619,709
A system and method for generating context vectors for use in storage and retrieval of documents and other information items. Context vectors represent conceptual relationships among information items by quantitative means. A neural network operates on a training corpus of records to develop relationship-based context vectors based on word proximity and co-importance using a technique of "windowed co-occurrence". Relationships among context vectors are deterministic, so that a context vector set has one logical solution, although it may have a plurality of physical solutions. No human knowledge, thesaurus, synonym list, knowledge base, or conceptual hierarchy, is required. Summary vectors of records may be clustered to reduce searching time, by forming a tree of clustered nodes. Once the context vectors are determined, records may be retrieved using a query interface that allows a user to specify content terms, Boolean terms, and/or document feedback. The present invention further facilitates visualization of textual information by translating context vectors into visual and graphical representations. Thus, a user can explore visual representations of meaning, and can apply human visual pattern recognition skills to document searches. (Abstract). The computer-implemented process of claim 17, wherein the step of storing using convergent k-means clustering comprises the substeps of:
(e.1) initially partitioning the summary vectors into a plurality of clusters, each cluster containing at least one summary vector;
(e.2) determining a centroid vector for each cluster;
(e.3) for each summary vector, performing the steps of:
(e.3.1) determining the distance between the summary vector and each centroid vector;
(e.3.2) determining which centroid vector has the smallest distance to the summary vector;
(e.3.3) responsive to the determined centroid vector not being associated with the cluster containing the summary vector, adjusting the contents of the clusters so that the cluster associated with the determined centroid contains the summary vector;
(e.4) adjusting the centroid vectors for the clusters that have been adjusted; and
(e.5) responsive to a stopping criterion not having been achieved, repeating steps (e.1) through (e.5). (dependent claim 18).
Jiang 7,296,009
The automatic determination of the threshold T is done as follows. First, many different partitions of the document set are generated by varying the threshold T. Large threshold values result in a small number of general clusters while small threshold values produce a large number of more specific clusters. Next, each partition is assigned a value that indicates the quality of the partition. This value takes into account cohesion, i.e. the closeness of the documents within the same cluster as well as the isolation of different clusters. This value is the sum of the inter-cluster distances and the intra-cluster distances. The inter-cluster distance is the distance of each document from its cluster centroid and the intra-cluster distance is the distance of each cluster centroid from the global centroid (the average of the word frequency vectors of all the documents in the document set). When there is one document per cluster or when all documents are grouped into a cluster, this value takes on the maximum value, which is the sum of the distances of the documents from the global centroid. The best partition is when this value is minimised and a few compact clusters are obtained. By this process, the distance threshold T that generates clusters that reflect the natural structure of the document set is determined. Once the clusters are generated, the feature extractor 12 is used to choose a unique topic name based on the documents that make up the clusters. (col. 14 lines 17-42).
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to examiner ROBERT STEVENS whose telephone number is (571) 272-4102. The examiner can normally be reached Mon - Fri 6:00 - 2:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amy Ng can be reached on (571) 270-1698. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ROBERT STEVENS/Primary Examiner, Art Unit 2164
May 1, 2026