Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner’s Note
Providing supporting paragraph(s) for each limitation of amended/new claim(s) in Remarks is strongly requested for clear and definite claim interpretations by Examiner (e.g., to avoid rejections under 35 U.S.C § 112(a) “Lack of written description”)
Applicant can schedule interviews (via Automated Interview Request (AIR)) at any stage of the prosecution (e.g., Non-Final, Final, and After-Final) to discuss any issues related to, for example, rejections under 35 U.S.C § 101 and § 102/103, for moving toward allowance.
If a limitation has bold brackets (i.e. [·]) around claim languages, the bracketed claim languages indicate that they have not been taught yet by the current prior art reference but they will be taught by another prior art reference afterwards.
If a limitation has one or more bold underlines, the one or more bold underlined claim languages indicate that they are taught by the current prior art reference, while the one or more non-underlined claim languages indicate that they have been taught already by one or more previous art references.
Priority
Acknowledgment is made of applicant's claim for the present application filed on 09/18/2023.
Response to Arguments
Applicant's arguments filed on 06/05/2026 have been fully considered but they are not persuasive.
In Remarks, regarding 35 USC § 101, Applicant contends:
“In this case Applicant's specification explicitly mentions that "Negative training data 212 helps the categorization system 200 learn alternate uses of similar terminology that are not related to the topic of interest and thereby reduce false positives." (See Applicant's Specification as filed, at paragraph [0027]-[0028] ). In addition, Applicant's Specification points that "the negative training data helps the machine learning process learn what to ignore." (See Applicant's Specification as filed, at paragraph [0040]). The claimed invention therefore improves machine learning classification performance through a particular training methodology rather than merely applying a generic machine-learning model to existing data.”
“Rather, the claims recite a specific technique for constructing training data that improves the machine learning model's ability to distinguish relevant industry-specific terminology from similar terminology appearing in unrelated industries.”
“Instead, the specification discloses a specific machine-learning training technique that improves classification performance.”
Examiner’s response:
The examiner understands the applicant’s assertion.
However, it appears that each processing step is just applying the abstract idea to a general field of endeavor with additional elements. In addition, improvements to technology or technical field are not necessarily reflected in the claims. Thus, the claim does not integrate the judicial exception into a practical application, and the claim does not amount to significantly more than the judicial exception.
The examiner understands the applicant’s assertion “In this case Applicant's specification explicitly mentions that "Negative training data 212 helps the categorization system 200 learn alternate uses of similar terminology that are not related to the topic of interest and thereby reduce false positives." (See Applicant's Specification as filed, at paragraph [0027]-[0028] ). In addition, Applicant's Specification points that "the negative training data helps the machine learning process learn what to ignore." (See Applicant's Specification as filed, at paragraph [0040]). The claimed invention therefore improves machine learning classification performance through a particular training methodology rather than merely applying a generic machine-learning model to existing data.”
However, it is not clear why/how “learning alternate uses of similar terminology that are not related to the topic of interest” could be an improvement. It appears that using positive and negative training data to reduce false positives is known, and it seems that the present invention just uses “alternate uses of similar terminology that are not related to the topic of interest” as negative training data. But it is not clear why/how using that kind of negative training data provides improvements. Providing more details may help overcome the present rejections.
The examiner understands the applicant’s assertion “Rather, the claims recite a specific technique for constructing training data that improves the machine learning model's ability to distinguish relevant industry-specific terminology from similar terminology appearing in unrelated industries.”
It seems that “constructing training data” is the claimed key inventive concept of the presentation application. But it is not clear why/how “distinguishing relevant industry-specific terminology from similar terminology appearing in unrelated industries” is an improvement based on the recited claims. Rather, it appears that it is just an improvement to the abstract ideas. Providing more details may help overcome the present rejections.
The examiner understands the applicant’s assertion “Instead, the specification discloses a specific machine-learning training technique that improves classification performance.”
However, improving classification performance does not always provide improvements. Rather, for now, it appears that it is just an improvement to the abstract ideas. In addition, it appears that in general, reducing false positives is obtained during machine learning. Providing more details may help overcome the present rejections.
Currently, the limitations do not clearly show e.g., improvements in computer technology and improvements to other technical fields. Rather, the improvements in Remarks are about just improving the abstract ideas of the independent claims. It doesn’t seem that the specification and/or the independent claims clearly show how the inventive concept of the claims enables improvements and how they are tied together. The applicant may need to amend the claims to show how the claim languages and improvements are tied together.
To find a valid improvement to a technology, MPEP 2106.04(d)(1) says the specification must explain the improvement and that the claim must reflect the disclosed improvement. Furthermore, the improvement should not be merely a consequence of the abstract idea. See MPEP 2106.05(a). An improvement in the abstract idea itself is not an improvement to technology.
For at least these reasons, Applicant's arguments are not convincing.
Applicant’s arguments regarding 35 USC § 102/103 with respect to the independent claims have been considered but are moot because the arguments are directed to amended limitation(s) that has/have not been previously examined.
Claim Objections
Claim(s) 1, 13, 25 is/are objected to because of the following informalities.
Claim(s) 1 is/are objected to because of the following informalities: it appears that “vectorized positive training data and negative training data” (5th last line) needs to read “vectorized lemmatized positive training data and negative training data” or something else. Appropriate correction is required. In addition, claim(s) 13, 25 is/are objected to for the same reason.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 1-2, 4-14, 16-26, 28-36 is/are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “similar” (claim 1, line 7) is a relative term which renders the claim indefinite. The term “similar” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. In addition, claim(s) 13, 25 is/are rejected for the same reason.
Claim(s) 1, 13, 25 each recite(s) limitations that raise issues of indefiniteness as set forth above, and their dependent claims are rejected at least based on their direct and/or indirect dependency from the claims listed above. Appropriate explanation and/or amendment is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-2, 4-14, 16-26, 28-36 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1
Step 1: “Is the claim to a process, machine, manufacture, or composition of matter?”
The claim is directed to a method. Therefore, yes.
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?”
lemmatizing the positive training data and negative training data, wherein the lemmatization reduces a number of unique words in the positive training data and negative training data; (i.e., mental process)
vectorizing the lemmatized positive training data and negative training data into a vector space to form a vectorized positive training data and negative training data; and (i.e., mental process)
The claim is directed to an abstract idea. Therefore, yes.
Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements:
receiving labeled positive training data; (insignificant extra-solution activity of receiving data, see MPEP 2106.05(g))
wherein the labeled positive training data comprises search terms associated with an industry basket, wherein the industry basket is a grouping of companies that operate within same industry (a particular type or source of model/data, Field of Use and Technological Environment, see MPEP 2106.05(h))
wherein the negative training data comprises a set of search terms that are similar to the search terms in the positive training data and are used in a different context, and wherein the search terms are associated with another industry basket that is different from the industry basket; (a particular type or source of model/data, Field of Use and Technological Environment, see MPEP 2106.05(h))
receiving negative training data; (insignificant extra-solution activity of receiving data, see MPEP 2106.05(g))
training a machine learning model, with the vectorized positive training data and negative training data, to identify a new category of data from existing categories in the positive training data and negative training data (adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, see MPEP 2106.05(f))
wherein the vectorized negative training data is used for the machine learning model to learn alternate uses of terminology that are not related to topic of interest to reduce false positives (adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, see MPEP 2106.05(f))
Therefore, no.
Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?”
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Specifically, the claimed inventions simply append well-understood, routine and conventional activities previously known to the industry, both when viewed independently and as an ordered combination, specified at a high level of generality, to the judicial exception, (e.g., a claim to an abstract idea requiring no more than a generic computer to perform generic computer functions that are well-understood, routine and conventional activities previously known to the industry). Therefore, no.
Regarding claim 2
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?”
generate weights for the positive training data and negative training data; and (i.e., mental process)
The claim is directed to an abstract idea. Therefore, yes.
Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements:
applying a term frequency-inverse document frequency (TF-IDF) algorithm to the vector space to (well-understood, routine, and conventional generic computer and/or model, see MPEP 2106.05(f))
feeding the weighted positive training data and negative training data into a machine learning classifier (insignificant extra-solution activity of mere data gathering, see MPEP 2106.05(g))
Therefore, no.
Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?”
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, no.
Regarding claim 4
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?” The claim recites the abstract idea identified above regarding claim 1. Therefore, yes.
Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements:
applying sentence transformer fine-tuning (SetFit) embeddings to a key term search according to proximity; and (adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, see MPEP 2106.05(f))
feeding results of the key term search into a machine learning classifier (insignificant extra-solution activity of mere data gathering, see MPEP 2106.05(g))
Therefore, no.
Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?”
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, no.
Regarding claim 5
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?”
removing ambiguous labels from the positive training data and negative training data (i.e., mental process)
The claim is directed to an abstract idea. Therefore, yes.
The claim does not add any additional elements (Step 2A Prong 2) or significantly more (Step 2B).
Regarding claim 6
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?”
reviewing positive training data and negative training data for specified key terms (i.e., mental process)
The claim is directed to an abstract idea. Therefore, yes.
The claim does not add any additional elements (Step 2A Prong 2) or significantly more (Step 2B).
Regarding claim 7
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?”
responsive to a determination any of the specified key terms are missing from the positive training data and negative training data, adding the missing specified key terms to the positive training data and negative training data (i.e., mental process)
The claim is directed to an abstract idea. Therefore, yes.
The claim does not add any additional elements (Step 2A Prong 2) or significantly more (Step 2B).
Regarding claim 8
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?” The claim recites the abstract idea identified above regarding claim 1. Therefore, yes.
Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements:
the machine learning model comprises an XGBoost classifier (a particular type or source of model/data, Field of Use and Technological Environment, see MPEP 2106.05(h))
Therefore, no.
Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?”
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, no.
Regarding claim 9
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?” The claim recites the abstract idea identified above regarding claim 1. Therefore, yes.
Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements:
the positive training data comprises search terms from global filings and an industry basket (a particular type or source of model/data, Field of Use and Technological Environment, see MPEP 2106.05(h))
Therefore, no.
Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?”
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, no.
Regarding claim 10
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?” The claim recites the abstract idea identified above regarding claim 1. Therefore, yes.
Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements:
the negative training data comprises search terms from global filings and other industry baskets (a particular type or source of model/data, Field of Use and Technological Environment, see MPEP 2106.05(h))
Therefore, no.
Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?”
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, no.
Regarding claim 11
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?”
identifies a number of new terms related to the new category of data (i.e., mental process)
The claim is directed to an abstract idea. Therefore, yes.
Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements:
the machine learning model (well-understood, routine, and conventional generic computer and/or model, see MPEP 2106.05(f))
Therefore, no.
Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?”
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, no.
Regarding claim 12
Step 2A Prong 1: “Does the claim recite an abstract idea, law of nature, or natural phenomenon?” The claim recites the abstract idea identified above regarding claim 1. Therefore, yes.
Step 2A Prong 2: “Does the claim recite additional elements that integrate the judicial exception into a practical application?” The following elements are directed to additional elements:
the new terms are associated with a number of companies in an emerging industry (a particular type or source of model/data, Field of Use and Technological Environment, see MPEP 2106.05(h))
Therefore, no.
Step 2B: “Does the claim recite additional elements that amount to significantly more than the judicial exception?”
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, no.
Regarding claim 13
The claim is rejected for the reasons set forth in the rejection of Claim 1 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 14
The claim is rejected for the reasons set forth in the rejection of Claim 2 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 16
The claim is rejected for the reasons set forth in the rejection of Claim 4 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 17
The claim is rejected for the reasons set forth in the rejection of Claim 5 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 18
The claim is rejected for the reasons set forth in the rejection of Claim 6 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 19
The claim is rejected for the reasons set forth in the rejection of Claim 7 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 20
The claim is rejected for the reasons set forth in the rejection of Claim 8 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 21
The claim is rejected for the reasons set forth in the rejection of Claim 9 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 22
The claim is rejected for the reasons set forth in the rejection of Claim 10 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 23
The claim is rejected for the reasons set forth in the rejection of Claim 11 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 24
The claim is rejected for the reasons set forth in the rejection of Claim 12 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 25
The claim is rejected for the reasons set forth in the rejection of Claim 1 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 26
The claim is rejected for the reasons set forth in the rejection of Claim 2 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 28
The claim is rejected for the reasons set forth in the rejection of Claim 4 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 29
The claim is rejected for the reasons set forth in the rejection of Claim 5 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 30
The claim is rejected for the reasons set forth in the rejection of Claim 6 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 31
The claim is rejected for the reasons set forth in the rejection of Claim 7 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 32
The claim is rejected for the reasons set forth in the rejection of Claim 8 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 33
The claim is rejected for the reasons set forth in the rejection of Claim 9 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 34
The claim is rejected for the reasons set forth in the rejection of Claim 10 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 35
The claim is rejected for the reasons set forth in the rejection of Claim 11 under 35 U.S.C. 101, mutatis mutandis.
Regarding claim 36
The claim is rejected for the reasons set forth in the rejection of Claim 12 under 35 U.S.C. 101, mutatis mutandis.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 5-8, 11-14, 17-20, 23-26, 29-32, 35-36 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhopale et al. (A Review-and-Reviewer based approach for Fake Review Detection) in view of An et al. (Fine-grained Category Discovery under Coarse-grained supervision with Hierarchical Weighted Self-contrastive Learning)
Regarding claim 1
Bhopale teaches
A computer-implemented method for training a machine learning model for categorizing data, the method comprising:
(Bhopale [sec(s) Abs] “This paper presents an approach to fake review detection, essentially for online hotel reviews, by combining the review-based approach and reviewer-based approach. Different Natural Language Processing techniques such as tokenization, lemmatization, vectorization, etc. are used to extract insightful features from the review text data. After text mining, the data is used to train different classification models using machine learning algorithms that detect fake reviews.” [sec(s) I] “With the advent of big data, statistical learning and artificial intelligence, “natural language processing” has emerged as a research area where a lot of focus is put on studying this phenomenon of natural language and also enabling the computers to understand it. Describing the origination of Natural Language Processing (NLP), the first basis for the study of language was called ‘linguistics’, which includes the detailed study of grammar, semantics, and phonetics, which was turned to computational linguistics, which made use of computer tools to study language. Further, computational linguistics became known as natural language process. NLP is majorly based on statistics, which often leads to it being described as Statistical Natural Language Processing and involves dealing with string data types, text data, documents, etc. [1] NLP helps computers understand unstructured text and extract meaningful data from it.”;)
receiving labeled positive training data, wherein the labeled positive training data comprises search terms associated with an industry basket, wherein the industry basket is a grouping of companies that operate within same industry;
(Bhopale [sec(s) III.B] “The data used is corpus consisting of truthful and deceptive hotel reviews for a total of 1600 reviews for 20 popular Chicago hotels. In the dataset, each of the 20 hotels has 20 reviews. This data is obtained from and described in [10] and [11]. The data is perfectly balanced with no skew as there are exactly 800 reviews for truthful and deceptive reviews each. Each type of review is again distinguished as positive or negative. The positive reviews are obtained from [10] via sources like TripAdvisor and Mechanical Turk. The negative reviews are obtained from [11] via sources like Expedia, Hotels.com, Orbitz, Priceline, Yelp, TripAdvisor and Mechanical Turk.” [sec(s) III.C] “After reading in the data, the target labels ‘truthful’ and ‘deceptive’ have been encoded and binarized to form the attribute ‘fake’, wherein a 0 means the review is genuine and 1 means that the review is fake.” [sec(s) I] “By studying the current methodologies of review classification and their shortcomings, the major objective of this research is to apply various machine learning algorithms to develop a model that determines whether a review is genuine or fake by creating a web user interface and testing the model against testing dataset and tweak its accuracy. It is also majorly aimed at aiding the customer in the decision making process by providing honest and genuine reviews, thereby enhancing their experience.”;)
receiving negative training data, wherein the negative training data comprises a set of search terms that are similar to the search terms in the positive training data and are used in a different context, and wherein the search terms are associated with another industry basket that is different from the industry basket;
(Bhopale [sec(s) III.B] “The data used is corpus consisting of truthful and deceptive hotel reviews for a total of 1600 reviews for 20 popular Chicago hotels. In the dataset, each of the 20 hotels has 20 reviews. This data is obtained from and described in [10] and [11]. The data is perfectly balanced with no skew as there are exactly 800 reviews for truthful and deceptive reviews each. Each type of review is again distinguished as positive or negative. The positive reviews are obtained from [10] via sources like TripAdvisor and Mechanical Turk. The negative reviews are obtained from [11] via sources like Expedia, Hotels.com, Orbitz, Priceline, Yelp, TripAdvisor and Mechanical Turk.” [sec(s) III.C] “After reading in the data, the target labels ‘truthful’ and ‘deceptive’ have been encoded and binarized to form the attribute ‘fake’, wherein a 0 means the review is genuine and 1 means that the review is fake. … Both of these columns have been dropped from the dataset and only the review ‘text’ column of the dataset has been taken as the data to be used for training the model and the newly created column ‘fake’ has been taken as the column with the target values.” [sec(s) I] “By studying the current methodologies of review classification and their shortcomings, the major objective of this research is to apply various machine learning algorithms to develop a model that determines whether a review is genuine or fake by creating a web user interface and testing the model against testing dataset and tweak its accuracy. It is also majorly aimed at aiding the customer in the decision making process by providing honest and genuine reviews, thereby enhancing their experience.” [sec(s) V] “The model can also be generalised for further online reviews for hospitals, restaurants, theatres, shops, and other such services.”;)
lemmatizing the positive training data and negative training data, wherein the lemmatization reduces a number of unique words in the positive training data and negative training data;
(Bhopale [sec(s) Abs] “This paper presents an approach to fake review detection, essentially for online hotel reviews, by combining the review-based approach and reviewer-based approach. Different Natural Language Processing techniques such as tokenization, lemmatization, vectorization, etc. are used to extract insightful features from the review text data. After text mining, the data is used to train different classification models using machine learning algorithms that detect fake reviews.” [sec(s) III.C] “For data preprocessing and text cleaning, regular expressions have been used. Using regular expressions, all special characters, single characters and extra spaces have been removed. Thereafter, lemmatization has been performed on the review text to extract the root word for all the tokenized words in the review text [19]. Followed by lemmatization, vectorization has been carried out using Term Frequency & Inverse Document Frequency (TF-IDF) vectorization.”; Note that for more details on “lemmatization”, please refer to [19], Uma et al. (Formation of SQL from Natural Language Query using NLP) – sec III.)
vectorizing the lemmatized positive training data and negative training data into a vector space to form a vectorized positive training data and negative training data; and
(Bhopale [sec(s) Abs] “This paper presents an approach to fake review detection, essentially for online hotel reviews, by combining the review-based approach and reviewer-based approach. Different Natural Language Processing techniques such as tokenization, lemmatization, vectorization, etc. are used to extract insightful features from the review text data. After text mining, the data is used to train different classification models using machine learning algorithms that detect fake reviews.” [sec(s) III.C] “For data preprocessing and text cleaning, regular expressions have been used. Using regular expressions, all special characters, single characters and extra spaces have been removed. Thereafter, lemmatization has been performed on the review text to extract the root word for all the tokenized words in the review text [19]. Followed by lemmatization, vectorization has been carried out using Term Frequency & Inverse Document Frequency (TF-IDF) vectorization. This has been done to essentially obtain a bag-of-words and represent the text data in a mathematical format suitable for feeding into the machine learning models. n-grams is a sequence of n items collected from a given sample of text or speech from a corpus.”;)
(Note: Hereinafter, if a limitation has bold brackets (i.e. [·]) around claim languages, the bracketed claim languages indicate that they have not been taught yet by the current prior art reference but they will be taught by another prior art reference afterwards.)
training a machine learning model, with the vectorized positive training data and negative training data, to identify a [new] category of data from existing categories in the positive training data and negative training data, wherein the vectorized negative training data is used for the machine learning model to learn alternate uses of terminology that are [not] related to topic of interest to reduce false positives.
(Bhopale [sec(s) Abs] “This paper presents an approach to fake review detection, essentially for online hotel reviews, by combining the review-based approach and reviewer-based approach. Different Natural Language Processing techniques such as tokenization, lemmatization, vectorization, etc. are used to extract insightful features from the review text data. After text mining, the data is used to train different classification models using machine learning algorithms that detect fake reviews.” [sec(s) III.C] “For data preprocessing and text cleaning, regular expressions have been used. Using regular expressions, all special characters, single characters and extra spaces have been removed. Thereafter, lemmatization has been performed on the review text to extract the root word for all the tokenized words in the review text [19]. Followed by lemmatization, vectorization has been carried out using Term Frequency & Inverse Document Frequency (TF-IDF) vectorization. This has been done to essentially obtain a bag-of-words and represent the text data in a mathematical format suitable for feeding into the machine learning models. n-grams is a sequence of n items collected from a given sample of text or speech from a corpus.” [sec(s) III.D] “For detection of fake reviews, five different machine learning techniques have been used for classification. The classification models created are Random Forest classifier, Multinomial Na¨ ıve Bayes classifier, Singular Vector Classifier, eXtreme Gradient (XG) Boost Classifier and XG Boost Random Forest Classifier” [sec(s) III.E] “Thereafter, the input data from the web page’s input text dialog box is passed through the same data preprocessing and text mining pipeline and later fed as input to the classification model to predict the output label. The label is then displayed onto the Flask web page.”;)
However, the combination of Bhopale does not appear to explicitly teach:
training a machine learning model, with the vectorized positive training data and negative training data, to identify a [new] category of data from existing categories in the positive training data and negative training data, wherein the vectorized negative training data is used for the machine learning model to learn alternate uses of terminology that are [not] related to topic of interest to reduce false positives.
(Note: Hereinafter, if a limitation has one or more bold underlines, the one or more underlined claim languages indicate that they are taught by the current prior art reference, while the one or more non-underlined claim languages indicate that they have been taught already by one or more previous art references.)
An teaches
training a machine learning model, with the vectorized positive training data and negative training data, to identify a new category of data from existing categories in the positive training data and negative training data, wherein the vectorized negative training data is used for the machine learning model to learn alternate uses of terminology that are not related to topic of interest to reduce false positives.
(An [fig(s) 1] [sec(s) Abs] “Novel category discovery aims at adapting models trained on known categories to novel categories. Previous works only focus on the scenario where known and novel categories are of the same granularity. In this paper, we investigate a new practical scenario called Fine-grained Category Discovery under Coarse grained supervision (FCDC). FCDC aims at discovering fine-grained categories with only coarse-grained labeled data, which can adapt models to categories of different granularity from known ones and reduce significant labeling cost.” [sec(s) 1] “Discovering novel categories based on some known categories has attracted much attention in both Natural Language Processing (Zhang et al., 2021; Zhao et al., 2021) and Computer Vision (Zhong et al., 2021; Han et al., 2019). Previous works assume that novel categories are of the same granularity (or of the same class hierarchy level) as known categories. However, in real-world scenarios, novel categories can be more fine-grained sub-categories of known ones (e.g., sports and tennis). … To meet this requirement, we investigate a new scenario named Fine-grained Category Discovery under Coarse-grained supervision (FCDC). As shown in Figure 1, FCDC needs models to discover fine-grained categories (e.g., tennis and music) based only on coarse-grained (e.g., sports and arts) labeled data which are easier and cheaper to obtain”;)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bhopale with the new category of An.
One of ordinary skill in the art would have been motived to combine in order to reduce significant labeling cost and adapt models to novel categories of different granularity from known ones.
(An [sec(s) 6] “In this paper, we investigate a novel task named Fine-grained Category Discovery under Coarse grained supervision (FCDC), which can reduce significant labeling cost and adapt models to novel categories of different granularity from known ones. We further propose a hierarchical weighted self contrastive model to approach the FCDC task by better controlling intra-class and inter-class distance. By performing supervised and contrastive learning on shallow and deep layers of pre-trained models, our model can learn fine-grained knowledge from shallow to deep with only coarse-grained supervision. Extensive experiments on public datasets show that our approach is more effective and efficient than compared methods.”)
Regarding claim 2
The combination of Bhopale, An teaches claim 1.
wherein training the machine learning model comprises: (See claim 1)
Bhopale further teaches
applying a term frequency-inverse document frequency (TF-IDF) algorithm to the vector space to generate weights for the positive training data and negative training data; and
(Bhopale [sec(s) Abs] “After text mining, the data is used to train different classification models using machine learning algorithms that detect fake reviews.” [sec(s) III.C] “The dataset also contains a column that states the information of the hotel name for which the review was given along with a column that mentions the source of the review. Both of these columns have been dropped from the dataset and only the review ‘text’ column of the dataset has been taken as the data to be used for training the model and the newly created column ‘fake’ has been taken as the column with the target values. … For data preprocessing and text cleaning, regular expressions have been used. Using regular expressions, all special characters, single characters and extra spaces have been removed. Thereafter, lemmatization has been performed on the review text to extract the root word for all the tokenized words in the review text [19]. Followed by lemmatization, vectorization has been carried out using Term Frequency & Inverse Document Frequency (TF-IDF) vectorization. This has been done to essentially obtain a bag-of-words and represent the text data in a mathematical format suitable for feeding into the machine learning models. n-grams is a sequence of n items collected from a given sample of text or speech from a corpus.”;)
feeding the weighted positive training data and negative training data into a machine learning classifier.
(Bhopale [sec(s) III.C] “Followed by lemmatization, vectorization has been carried out using Term Frequency & Inverse Document Frequency (TF-IDF) vectorization. This has been done to essentially obtain a bag-of-words and represent the text data in a mathematical format suitable for feeding into the machine learning models. n-grams is a sequence of n items collected from a given sample of text or speech from a corpus.” [sec(s) III.D] “For detection of fake reviews, five different machine learning techniques have been used for classification. The classification models created are Random Forest classifier, Multinomial Na¨ ıve Bayes classifier, Singular Vector Classifier, eXtreme Gradient (XG) Boost Classifier and XG Boost Random Forest Classifier” [sec(s) III.E] “Thereafter, the input data from the web page’s input text dialog box is passed through the same data preprocessing and text mining pipeline and later fed as input to the classification model to predict the output label. The label is then displayed onto the Flask web page.”;)
Regarding claim 5
The combination of Bhopale, An teaches claim 1.
Bhopale further teaches
removing ambiguous labels from the positive training data and negative training data.
(Bhopale [sec(s) III.C] “Additionally, the dataset also contains a column named ‘polarity’ which provides the sentiment of the review by stating whether the review given was positive or negative. This again has been written in the column in string data types and is thus binarized to integer values of 0 and 1; where, 0 is for a negative review and 1 is for a positive review. The dataset also contains a column that states the information of the hotel name for which the review was given along with a column that mentions the source of the review. Both of these columns have been dropped from the dataset and only the review ‘text’ column of the dataset has been taken as the data to be used for training the model and the newly created column ‘fake’ has been taken as the column with the target values. … For data preprocessing and text cleaning, regular expressions have been used. Using regular expressions, all special characters, single characters and extra spaces have been removed. Thereafter, lemmatization has been performed on the review text to extract the root word for all the tokenized words in the review text [19]. Followed by lemmatization, vectorization has been carried out using Term Frequency & Inverse Document Frequency (TF-IDF) vectorization. This has been done to essentially obtain a bag-of-words and represent the text data in a mathematical format suitable for feeding into the machine learning models. n-grams is a sequence of n items collected from a given sample of text or speech from a corpus.”;)
Regarding claim 6
The combination of Bhopale, An teaches claim 1.
wherein training the machine learning model further comprises (See claim 1)
Bhopale further teaches
reviewing positive training data and negative training data for specified key terms.
(Bhopale [sec(s) III.C] “Additionally, the dataset also contains a column named ‘polarity’ which provides the sentiment of the review by stating whether the review given was positive or negative. This again has been written in the column in string data types and is thus binarized to integer values of 0 and 1; where, 0 is for a negative review and 1 is for a positive review. The dataset also contains a column that states the information of the hotel name for which the review was given along with a column that mentions the source of the review. Both of these columns have been dropped from the dataset and only the review ‘text’ column of the dataset has been taken as the data to be used for training the model and the newly created column ‘fake’ has been taken as the column with the target values. … For data preprocessing and text cleaning, regular expressions have been used. Using regular expressions, all special characters, single characters and extra spaces have been removed. Thereafter, lemmatization has been performed on the review text to extract the root word for all the tokenized words in the review text [19]. Followed by lemmatization, vectorization has been carried out using Term Frequency & Inverse Document Frequency (TF-IDF) vectorization. This has been done to essentially obtain a bag-of-words and represent the text data in a mathematical format suitable for feeding into the machine learning models. n-grams is a sequence of n items collected from a given sample of text or speech from a corpus. Mainly three types of n-grams have been collected on the tokined words: unigrams, bigrams and trigrams. These are defined based on the number of tokenized words that are taken together while creating n-grams. Various combinations of unigrams, bigrams and trigrams have also been used during the vectorization process to crunch the data into a mathematical format for different grams”;)
Regarding claim 7
The combination of Bhopale, An teaches claim 6.
Bhopale further teaches
responsive to a determination any of the specified key terms are missing from the positive training data and negative training data, adding the missing specified key terms to the positive training data and negative training data.
(Bhopal [fig(s) 1] “Data Preprocessing and Text Mining using NLP”, “Training Data” [sec(s) III.A] “It illustrates the end-to-end process of the preparation of the classification models for fake review detection.” [sec(s) III.C] “Text data is not in the traditional form of data that includes numeric values that can be easily fed into various machine learning and statistical models. Rather, the text data needs to be preprocessed in order to convert it into a suitable form that can be used as input for various machine learning algorithms. The processes of feature extraction and data manipulation have been carried out on the review text data in this data preprocessing stage which is discussed in this section. After reading in the data, the target labels ‘truthful’ and ‘deceptive’ have been encoded and binarized to form the attribute ‘fake’, wherein a 0 means the review is genuine and 1 means that the review is fake. Additionally, the dataset also contains a column named ‘polarity’ which provides the sentiment of the review by stating whether the review given was positive or negative. This again has been written in the column in string data types and is thus binarized to integer values of 0 and 1; where, 0 is for a negative review and 1 is for a positive review. The dataset also contains a column that states the information of the hotel name for which the review was given along with a column that mentions the source of the review. Both of these columns have been dropped from the dataset and only the review ‘text’ column of the dataset has been taken as the data to be used for training the model and the newly created column ‘fake’ has been taken as the column with the target values. … This has been done to essentially obtain a bag-of-words and represent the text data in a mathematical format suitable for feeding into the machine learning models.”;)
Regarding claim 8
The combination of Bhopale, An teaches claim 1.
Bhopale further teaches
wherein the machine learning model comprises an XGBoost classifier.
(Bhopale [sec(s) III.D] “For detection of fake reviews, five different machine learning techniques have been used for classification. The classification models created are Random Forest classifier, Multinomial Na¨ ıve Bayes classifier, Singular Vector Classifier, eXtreme Gradient (XG) Boost Classifier and XG Boost Random Forest Classifier” [sec(s) III.E] “Thereafter, the input data from the web page’s input text dialog box is passed through the same data preprocessing and text mining pipeline and later fed as input to the classification model to predict the output label. The label is then displayed onto the Flask web page.”;)
Regarding claim 11
The combination of Bhopale, An teaches claim 1.
An further teaches
wherein the machine learning model identifies a number of new terms related to the new category of data.
(An [fig(s) 1] [sec(s) Abs] “Novel category discovery aims at adapting models trained on known categories to novel categories. Previous works only focus on the scenario where known and novel categories are of the same granularity. In this paper, we investigate a new practical scenario called Fine-grained Category Discovery under Coarse grained supervision (FCDC). FCDC aims at discovering fine-grained categories with only coarse-grained labeled data, which can adapt models to categories of different granularity from known ones and reduce significant labeling cost.” [sec(s) 1] “Discovering novel categories based on some known categories has attracted much attention in both Natural Language Processing (Zhang et al., 2021; Zhao et al., 2021) and Computer Vision (Zhong et al., 2021; Han et al., 2019). Previous works assume that novel categories are of the same granularity (or of the same class hierarchy level) as known categories. However, in real-world scenarios, novel categories can be more fine-grained sub-categories of known ones (e.g., sports and tennis). … To meet this requirement, we investigate a new scenario named Fine-grained Category Discovery under Coarse-grained supervision (FCDC). As shown in Figure 1, FCDC needs models to discover fine-grained categories (e.g., tennis and music) based only on coarse-grained (e.g., sports and arts) labeled data which are easier and cheaper to obtain. … Inspired by this phenomenon, the core motivation of our model is to learn coarse-grained knowledge by shallow layers of BERT and learn more fine-grained knowledge by the rest of deep layers hierarchically. This motivation is not only consistent with the feature extraction process of BERT, but also corresponding with the shallow-to-deep learning process of humans. Specifically, we use given coarse-grained labels to train shallow layers of BERT to learn some surface knowledge, then we propose a weighted self-contrastive module to train deep layers of BERT to learn more fine-grained knowledge based on the learned surface knowledge.”;)
The combination of Bhopale, An is combinable with An for the same rationale as set forth above with respect to claim 1.
Regarding claim 12
The combination of Bhopale, An teaches claim 11.
An further teaches
wherein the new terms are associated with a number of companies in an emerging industry.
(An [sec(s) 4] “To evaluate effectiveness of our model, we conduct experiments on three public datasets. Statistics of three datasets can be found in Table 1. CLINC is an intent classification dataset released by Larson et al. (2019). Web of Science (WOS) is a paper classification dataset released by Kowsari et al. (2017). HWU64is a personal assistant query classification dataset released by Liu et al. (2021).”;)
The combination of Bhopale, An is combinable with An for the same rationale as set forth above with respect to claim 1.
Regarding claim 13
The claim is rejected for the reasons set forth in the rejection of Claim 1.
Regarding claim 14
The claim is rejected for the reasons set forth in the rejection of Claim 2.
Regarding claim 17
The claim is rejected for the reasons set forth in the rejection of Claim 5.
Regarding claim 18
The claim is rejected for the reasons set forth in the rejection of Claim 6.
Regarding claim 19
The claim is rejected for the reasons set forth in the rejection of Claim 7.
Regarding claim 20
The claim is rejected for the reasons set forth in the rejection of Claim 8.
Regarding claim 23
The claim is rejected for the reasons set forth in the rejection of Claim 11.
Regarding claim 24
The claim is rejected for the reasons set forth in the rejection of Claim 12.
Regarding claim 25
The claim is rejected for the reasons set forth in the rejection of Claim 1.
Regarding claim 26
The claim is rejected for the reasons set forth in the rejection of Claim 2.
Regarding claim 29
The claim is rejected for the reasons set forth in the rejection of Claim 5.
Regarding claim 30
The claim is rejected for the reasons set forth in the rejection of Claim 6.
Regarding claim 31
The claim is rejected for the reasons set forth in the rejection of Claim 7.
Regarding claim 32
The claim is rejected for the reasons set forth in the rejection of Claim 8.
Regarding claim 35
The claim is rejected for the reasons set forth in the rejection of Claim 11.
Regarding claim 36
The claim is rejected for the reasons set forth in the rejection of Claim 12.
Claim(s) 4, 16, 28 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhopale et al. (A Review-and-Reviewer based approach for Fake Review Detection) in view of An et al. (Fine-grained Category Discovery under Coarse-grained supervision with Hierarchical Weighted Self-contrastive Learning) in view of Tunstall et al. (Efficient Few-Shot Learning Without Prompts)
Regarding claim 4
The combination of Bhopale, An teaches claim 1.
wherein training the machine learning model comprises: (See claim 1)
Bhopale further teaches
feeding results of the key term search into a machine learning classifier.
(Bhopale [fig(s) 1] “Data Preprocessing and Text Mining using NLP”, “Training Data” [sec(s) Abs] “This paper presents an approach to fake review detection, essentially for online hotel reviews, by combining the review-based approach and reviewer-based approach. Different Natural Language Processing techniques such as tokenization, lemmatization, vectorization, etc. are used to extract insightful features from the review text data. After text mining, the data is used to train different classification models using machine learning algorithms that detect fake reviews.” [sec(s) III.C] “Rather, the text data needs to be preprocessed in order to convert it into a suitable form that can be used as input for various machine learning algorithms. The processes of feature extraction and data manipulation have been carried out on the review text data in this data preprocessing stage which is discussed in this section. … For data preprocessing and text cleaning, regular expressions have been used. Using regular expressions, all special characters, single characters and extra spaces have been removed. Thereafter, lemmatization has been performed on the review text to extract the root word for all the tokenized words in the review text [19]. Followed by lemmatization, vectorization has been carried out using Term Frequency & Inverse Document Frequency (TF-IDF) vectorization. This has been done to essentially obtain a bag-of-words and represent the text data in a mathematical format suitable for feeding into the machine learning models. n-grams is a sequence of n items collected from a given sample of text or speech from a corpus.” [sec(s) III.A] “It illustrates the end-to-end process of the preparation of the classification models for fake review detection.”;)
However, the combination of Bhopale, An does not appear to explicitly teach:
applying sentence transformer fine-tuning (SetFit) embeddings to a key term search according to proximity; and
Tunstall teaches
applying sentence transformer fine-tuning (SetFit) embeddings to a key term search according to proximity; and
(Tunstall [sec(s) 1] “In this paper, we propose SETFIT, an approach based on Sentence Transformers (ST) (Reimers and Gurevych, 2019) that dispenses with prompts altogether and does not require large-scale PLMs to achieve high accuracy. For example, with only 8 labeled examples in the Customer Reviews (CR) sentiment dataset, SETFIT is competitive with fine tuning on the full training set, despite the fine-tuned model being three times larger (see Figure 1).” [sec(s) 3] “SETFIT is based on Sentence Transformers (Reimers and Gurevych, 2019) which are modifications of pretrained transformer models that use Siamese and triplet network structures to derive semantically meaningful sentence embeddings. The goal of these models is to minimize the distance between pairs of semantically similar sentences and maximize the distance between sentence pairs that are semantically distant. Standard STs output a fixed, dense vector that is meant to represent textual data and can then be used by machine learning algorithms” [sec(s) 3.1] “SETFIT uses a two-step training approach in which we first fine-tune an ST and then train a classifier head. In the first step, an ST is fine-tuned on the input data in a contrastive, Siamese manner on sentence pairs. In the second step, a text classification head is trained using the encoded training data generated by the fine-tuned ST from the first step. Figure 2 illustrates this process, and we discuss these two steps in the following sections”;)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bhopale, An with the SetFit embedding of Tunstall.
One of ordinary skill in the art would have been motived to combine in order to outperform the state-of-the art prompt-free method and rank alongside much larger prompt-based, few-shot models.
(Tunstall [sec(s) 1] “We summarize our contributions as follows: 1. We propose SETFIT– a simple and prompt free method– and provide a comprehensive guide for applying it in practical few-shot set tings. 2. We evaluate SETFIT’s performance on a number of few-shot text classifications tasks and show that it outperforms the state-of-the art prompt-free method and ranks alongside much larger prompt-based, few-shot models. 3. We make the code and data used in our work publicly available.”;)
Regarding claim 16
The claim is rejected for the reasons set forth in the rejection of Claim 4.
Regarding claim 28
The claim is rejected for the reasons set forth in the rejection of Claim 4.
Claim(s) 9-10, 21-22, 33-34 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhopale et al. (A Review-and-Reviewer based approach for Fake Review Detection) in view of An et al. (Fine-grained Category Discovery under Coarse-grained supervision with Hierarchical Weighted Self-contrastive Learning) in view of Rollins et al. (US 20170193619 A1)
Regarding claim 9
The combination of Bhopale, An teaches claim 1.
Bhopale further teaches
wherein the positive training data comprises search terms from global [filings] and an industry basket.
(Bhopale [sec(s) III.B] “The data used is corpus consisting of truthful and deceptive hotel reviews for a total of 1600 reviews for 20 popular Chicago hotels. In the dataset, each of the 20 hotels has 20 reviews. This data is obtained from and described in [10] and [11]. The data is perfectly balanced with no skew as there are exactly 800 reviews for truthful and deceptive reviews each. Each type of review is again distinguished as positive or negative. The positive reviews are obtained from [10] via sources like TripAdvisor and Mechanical Turk. The negative reviews are obtained from [11] via sources like Expedia, Hotels.com, Orbitz, Priceline, Yelp, TripAdvisor and Mechanical Turk.”;)
However, the combination of Bhopale, An does not appear to explicitly teach:
wherein the positive training data comprises search terms from global [filings] and an industry basket.
Rollins teaches
wherein the positive training data comprises search terms from global filings and an industry basket.
(Rollins [par(s) 44] “Matching engine 104 may use supervised machine learning via a naïve Bayesian classifier based rules-engine to continuously learn and improve over time the optimal inclusion and/or weightings of the above mentioned attributes and/or other characteristics of the parties to predict successful R&D collaboration among parties.” [par(s) 4] “The global system of patent offices and registries, which includes over 150 different entities may not alleviate these difficulties, as global ownership rights may be difficult to understand because there is no single entity with the ability to verify global ownership rights.” [par(s) 45] “Measures of successful/unsuccessful R&D collaboration may include some or all of annual patent fee payments, global patent filing breadth, litigation records, citations, derivative patent applications (e.g., continuation, divisional, and/or continuation-in-part), licensing information mergers and acquisition activity, and investment activity” [par(s) 49] “Matching engine 104 may use various public and private information sources to perform the above described predictions and projections, including, for example, measures of stock performance, measures of global patent fee payments per patent per portfolio, measures of R&D funding budgets, measures of R&D licensing revenues (e.g., as determined from United States SEC Filings and global equivalent regulatory documents), measures of patents used as loan collateral (e.g., as determined by available patent assignment information.)” [par(s) 71] “Adjacent technologies may be identified by nearest neighbors clustering of patent classification(s) clusters that define a technology in global patent databases of documents and/or nearest neighbors of text phrases in global patent databases of documents that define a technology. In R&D Technology areas where R&D collaboration in the system has been demonstrated to be more difficult, for example, a higher risk may associated.” [par(s) 35] “Such information may be derived from a party's volunteering the information and/or from public sources, such as SEC filings, and academic and Government Public Funding Disclosures.”;)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Bhopale, An with the global filings of Rollins.
One of ordinary skill in the art would have been motived to combine in order to continuously learn and improve over time the optimal inclusion and/or weightings of the mentioned attributes and/or other characteristics of the parties to predict successful R&D collaboration among parties.
(Rollins [par(s) 44] “Matching engine 104 may use supervised machine learning via a naïve Bayesian classifier based rules-engine to continuously learn and improve over time the optimal inclusion and/or weightings of the above mentioned attributes and/or other characteristics of the parties to predict successful R&D collaboration among parties.”)
Regarding claim 10
The combination of Bhopale, An teaches claim 1.
Bhopale further teaches
wherein the negative training data comprises search terms from global [filings] and other industry baskets.
(Bhopale [sec(s) III.B] “The data used is corpus consisting of truthful and deceptive hotel reviews for a total of 1600 reviews for 20 popular Chicago hotels. In the dataset, each of the 20 hotels has 20 reviews. This data is obtained from and described in [10] and [11]. The data is perfectly balanced with no skew as there are exactly 800 reviews for truthful and deceptive reviews each. Each type of review is again distinguished as positive or negative. The positive reviews are obtained from [10] via sources like TripAdvisor and Mechanical Turk. The negative reviews are obtained from [11] via sources like Expedia, Hotels.com, Orbitz, Priceline, Yelp, TripAdvisor and Mechanical Turk.”;)
However, the combination of Bhopale, An does not appear to explicitly teach:
wherein the negative training data comprises search terms from global [filings] and other industry baskets.
Rollins teaches
wherein the negative training data comprises search terms from global filings and other industry baskets.
(Rollins [par(s) 44] “Matching engine 104 may use supervised machine learning via a naïve Bayesian classifier based rules-engine to continuously learn and improve over time the optimal inclusion and/or weightings of the above mentioned attributes and/or other characteristics of the parties to predict successful R&D collaboration among parties.” [par(s) 4] “The global system of patent offices and registries, which includes over 150 different entities may not alleviate these difficulties, as global ownership rights may be difficult to understand because there is no single entity with the ability to verify global ownership rights.” [par(s) 45] “Measures of successful/unsuccessful R&D collaboration may include some or all of annual patent fee payments, global patent filing breadth, litigation records, citations, derivative patent applications (e.g., continuation, divisional, and/or continuation-in-part), licensing information mergers and acquisition activity, and investment activity” [par(s) 49] “Matching engine 104 may use various public and private information sources to perform the above described predictions and projections, including, for example, measures of stock performance, measures of global patent fee payments per patent per portfolio, measures of R&D funding budgets, measures of R&D licensing revenues (e.g., as determined from United States SEC Filings and global equivalent regulatory documents), measures of patents used as loan collateral (e.g., as determined by available patent assignment information.)” [par(s) 71] “Adjacent technologies may be identified by nearest neighbors clustering of patent classification(s) clusters that define a technology in global patent databases of documents and/or nearest neighbors of text phrases in global patent databases of documents that define a technology. In R&D Technology areas where R&D collaboration in the system has been demonstrated to be more difficult, for example, a higher risk may associated.” [par(s) 35] “Such information may be derived from a party's volunteering the information and/or from public sources, such as SEC filings, and academic and Government Public Funding Disclosures.”;)
The combination of Bhopale, An is combinable with Rollins for the same rationale as set forth above with respect to claim 9.
Regarding claim 21
The claim is rejected for the reasons set forth in the rejection of Claim 9.
Regarding claim 22
The claim is rejected for the reasons set forth in the rejection of Claim 10.
Regarding claim 33
The claim is rejected for the reasons set forth in the rejection of Claim 9.
Regarding claim 34
The claim is rejected for the reasons set forth in the rejection of Claim 10.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEHWAN KIM whose telephone number is (571)270-7409. The examiner can normally be reached Mon - Thu 7:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached on (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SEHWAN KIM/Examiner, Art Unit 2129